Localisation
Teaching an agent Urdu, Roman Urdu, and ٹماٹر
A produce agent for Pakistani farmers that only understands formal English is a demo. Getting to something usable meant handling three scripts, local crop names, and right-to-left rendering as first-class requirements.
The first messages real users sent us did not look like our test fixtures. They looked like this:
800 kg tamatar multan, kal ready
ٹماٹر ۸۰۰ کلو، ملتان
have aloo 2 ton faisalabad ready friday
Three scripts, mixed numerals, crop names in whichever form came to hand. If the parser only accepts the third line, the product serves the smallest and least underserved slice of its own market.
Aliases before intelligence
Before any model call, a crop alias table resolves local names to canonical crops: ٹماٹر and tamatar → tomato, آلو and aloo → potato, پیاز and pyaaz → onion, mirch → green chilli. Nine crops, dozens of surface forms. It's unglamorous string work, and it fixes more real failures than any prompt change we made.
The model handles what the table cannot: dates expressed as "kal" or "after two days", quantities in maunds, locations written loosely. Extraction returns crop, quantity in kilos, city, and ready date — and when confidence is low, the agent asks one clarifying question instead of guessing.
Alias resolved
- tamatar · ٹماٹر
- → Tomato
- kal
- → 1 Sep 2026
- Confidence
- High
Urdu is not a translation layer
Adding Urdu properly meant more than a dictionary of UI strings. It meant right-to-left chat bubbles, Nastaliq rendering with a font that survives long lines, and model prompts that switch language so replies come back in Urdu rather than translated English. It also meant the quick-reply buttons, approval cards and alerts all switching together — a half-translated interface reads as broken, not bilingual.
One tap in the chat header flips the entire conversation, including everything the agent says next.
Voice matters more than we expected
Typing Urdu on a phone in a field with wet hands is a poor interface. Speech input, transcribed and then parsed by the same extraction path, became the fastest way to file a lot. It works in Chromium-based browsers today, which is a browser-support compromise we accepted to ship.
Grading has to agree with the gate
The vision model returns grade, ripeness, defect rate, and quality notes from a photo. Its real test is not whether the label sounds right — it's whether the buyer's inspector at the receiving mandi reaches the same conclusion. So the output stays conservative, always shows a confidence level, and never overwrites what the seller says about their own crop. When there are no photos, the agent estimates from description and says so.