All posts

Goa man finds lost shoes at Ujjain temple using ChatGPT

Huma ShaziaAugust 16, 2026 at 1:16 PM3 min read
Goa man finds lost shoes at Ujjain temple using ChatGPT

A cybersecurity professional from Goa used ChatGPT's image recognition to locate his missing clogs outside the Mahakaleshwar Temple in Ujjain. Shubhang Borkar photographed rows of footwear, fed the images to the chatbot alongside a reference picture of his shoes, and let the AI scan for a match. It worked.

Goa man finds lost shoes at Ujjain temple using ChatGPT
Source: mint

Borkar shared the experiment in an Instagram Reel titled "Used ChatGPT to find my lost clogs" about four days ago. The video opens with a pan across chaotic shoe racks, hundreds of pairs left by pilgrims before entering one of Hinduism's 12 sacred Jyotirlingas. He describes himself as a "hacker on weekdays (cybersec), explorer on weekends."

Advertisements

How the search worked

Borkar opened ChatGPT and uploaded two things: a snapshot of what his clogs looked like and a wider shot of the footwear scattered outside the temple. He told the chatbot he had lost his shoes and asked it to scan the images for a match.

As he moved along different sections of the racks, he sent additional photographs. ChatGPT analyzed each batch. Eventually, the tool flagged a specific pair. The video shows Borkar retrieving them.

The approach relies on GPT-4's multimodal vision capabilities, which let the model interpret images and compare visual details. OpenAI introduced these features in late 2023, but real-world demonstrations beyond document parsing and accessibility remain uncommon.

Why the temple setting matters

Indian temples typically require visitors to remove footwear before entry. At major pilgrimage sites like Mahakaleshwar, thousands of pairs pile up in racks or on the ground. Losing shoes is routine. Most visitors rely on memory, shoe-minders, or branded bags to keep track of their belongings.

Borkar's workaround is not scalable, you still need phone access, a reference image, and patience, but it demonstrates that the vision model can handle cluttered, real-world scenes better than many users assume.

400 million+
Weekly active ChatGPT users as of February 2025, per OpenAI CEO Sam Altman

Practical limits of the hack

The method depends on having a clear reference image. Without one, ChatGPT has no baseline to match against. Lighting, camera angle, and crowd density all affect accuracy. Borkar had to send multiple photographs before the model found a plausible match.

It also assumes the shoes are visible in the first place. If someone moved or took them, image recognition offers no help.

Still, the technique points to broader applications: identifying luggage on a crowded carousel, spotting a parked car in a packed lot, or verifying items in cluttered inventory. None of these are ChatGPT's intended use case, but the vision model can handle them when prompted correctly.

ℹ️

Logicity's Take

This is a gimmick, but a useful one. Most ChatGPT users never touch image input; they treat it as a text tool. Borkar's experiment shows the vision layer is underutilized. For teams building internal tools, the same capability could power visual QA, inventory checks, or asset tagging without training a custom model. OpenAI's API pricing for GPT-4 vision starts at roughly $0.01 per image. Google's Gemini and Anthropic's Claude offer similar multimodal features at comparable rates.

Also Read
Check if your ChatGPT, Claude, or Perplexity account is hacked

Related guide on securing AI tool accounts

What this says about AI adoption

The viral reaction to Borkar's video underscores a gap between what multimodal AI can do and what most people use it for. ChatGPT's weekly active user base crossed 400 million earlier this year, but the majority interact through text prompts alone.

Demonstrations like this, quirky as they are, push adoption further. They also raise the question of what other everyday problems could be solved by pointing a phone camera at them and asking a model to look.

ℹ️

Need Help Implementing This?

If you're exploring image recognition or multimodal AI for your product, reach out to the Logicity team. We can connect you with implementation partners or review your use case.

Source: mint

H

Huma Shazia

Senior AI & Tech Writer

Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.

Related Articles