A cybersecurity professional from Goa used ChatGPT's image recognition to locate his missing clogs outside the Mahakaleshwar Temple in Ujjain. Shubhang Borkar photographed rows of footwear, fed the images to the chatbot alongside a reference picture of his shoes, and let the AI scan for a match. It worked.

Borkar shared the experiment in an Instagram Reel titled "Used ChatGPT to find my lost clogs" about four days ago. The video opens with a pan across chaotic shoe racks, hundreds of pairs left by pilgrims before entering one of Hinduism's 12 sacred Jyotirlingas. He describes himself as a "hacker on weekdays (cybersec), explorer on weekends."
How the search worked
Borkar opened ChatGPT and uploaded two things: a snapshot of what his clogs looked like and a wider shot of the footwear scattered outside the temple. He told the chatbot he had lost his shoes and asked it to scan the images for a match.
As he moved along different sections of the racks, he sent additional photographs. ChatGPT analyzed each batch. Eventually, the tool flagged a specific pair. The video shows Borkar retrieving them.
The approach relies on GPT-4's multimodal vision capabilities, which let the model interpret images and compare visual details. OpenAI introduced these features in late 2023, but real-world demonstrations beyond document parsing and accessibility remain uncommon.
Why the temple setting matters
Indian temples typically require visitors to remove footwear before entry. At major pilgrimage sites like Mahakaleshwar, thousands of pairs pile up in racks or on the ground. Losing shoes is routine. Most visitors rely on memory, shoe-minders, or branded bags to keep track of their belongings.
Borkar's workaround is not scalable, you still need phone access, a reference image, and patience, but it demonstrates that the vision model can handle cluttered, real-world scenes better than many users assume.
Practical limits of the hack
The method depends on having a clear reference image. Without one, ChatGPT has no baseline to match against. Lighting, camera angle, and crowd density all affect accuracy. Borkar had to send multiple photographs before the model found a plausible match.
It also assumes the shoes are visible in the first place. If someone moved or took them, image recognition offers no help.
Still, the technique points to broader applications: identifying luggage on a crowded carousel, spotting a parked car in a packed lot, or verifying items in cluttered inventory. None of these are ChatGPT's intended use case, but the vision model can handle them when prompted correctly.
Logicity's Take
This is a gimmick, but a useful one. Most ChatGPT users never touch image input; they treat it as a text tool. Borkar's experiment shows the vision layer is underutilized. For teams building internal tools, the same capability could power visual QA, inventory checks, or asset tagging without training a custom model. OpenAI's API pricing for GPT-4 vision starts at roughly $0.01 per image. Google's Gemini and Anthropic's Claude offer similar multimodal features at comparable rates.
Related guide on securing AI tool accounts
What this says about AI adoption
The viral reaction to Borkar's video underscores a gap between what multimodal AI can do and what most people use it for. ChatGPT's weekly active user base crossed 400 million earlier this year, but the majority interact through text prompts alone.
Demonstrations like this, quirky as they are, push adoption further. They also raise the question of what other everyday problems could be solved by pointing a phone camera at them and asking a model to look.
Need Help Implementing This?
If you're exploring image recognition or multimodal AI for your product, reach out to the Logicity team. We can connect you with implementation partners or review your use case.
Source: mint
Huma Shazia
Senior AI & Tech Writer
Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.
Related Articles
More in Trending Tech
Humanity Just Went Farther Into Space Than Ever Before — And Made It Back Alive
Four astronauts splashed down in the Pacific Ocean on April 10, 2026, after traveling farther from Earth than any human beings in history. The Artemis II crew shattered a 56-year-old distance record set by Apollo 13, journeying nearly 253,000 miles from our planet during their 10-day lunar flyby mission. This marks the first time humans have ventured beyond low Earth orbit since 1972.

Amflow's Electric Bikes Are Blowing The Competition Away
Amflow, the e-bike brand spun out of DJI, has just released two impressive new electric mountain bikes that are breaking the mold with unprecedented power, range, and lightness. The flagship bikes are powered by the innovative Avinox motors and come with features like onboard navigation and heart rate control.

Canva Just Made a Power Play: Here's What It Means for the Future of Design and Marketing
Canva has made a bold move by acquiring two companies, Simtheory and Ortto, to boost its AI and marketing automation capabilities. This strategic move is set to revolutionize the way teams work on design and marketing projects. With these acquisitions, Canva is poised to become an all-in-one platform for businesses and individuals alike.


