ai in agriculture

Voice-first, not chat-first

Every farmer-facing build at KissanAI ships with voice on by default, in every language we can serve well. Why access, not features, drives that call, and what it means for an agribusiness scoping a deployment.

· 5 min read
A voice waveform carrying language, weather, pest, and crop context while a chat box remains secondary.

When an agribusiness team sits down with us to scope a farmer-facing deployment, the first number I show them is the ratio of voice queries to text queries on a live system. In what we have run so far, voice queries outnumber text somewhere between two and four to one depending on the region. Hold the exact ratio loosely, the sample behind it is still small at our stage. The direction has not flipped once, and each time it has come in higher than the partner expected when they wrote the spec.

That ratio is the premise of this post. For the farmer we are building for, voice is not a nice feature. It is the entry point. A product that treats voice as a “v2 enhancement” or a “premium tier” or a “long-term roadmap item” is excluding the user whose access depends on it.

That changes how we build, and how we ask partners to scope.

The text-only spec rarely survives production

When a new deployment goes live, voice is on. Text is also available for users who prefer it. The choice is the user’s, not the deploying business’s. More than one partner has come to us scoping a “WhatsApp text bot” because text felt simpler to build and certify. Each time, the voice share in production was high enough that the text-only framing was gone within the first quarter.

The channel matters as much as the modality. Farmers already live on WhatsApp voice notes and ordinary phone calls, and a deployment that meets them there, voice in and voice out, gets used. One that asks them to type into a new app mostly does not. So when a partner asks where to start, our answer is consistent to the point of being boring: voice and text together from the first release, on the channel your farmers already use. Scoping “text now, voice later” creates a deployment that does not match its own usage pattern and has to be rebuilt mid-cycle, for roughly the same effort as building both up front.

The dialect is the expensive part

Indian languages have substantial intra-language variation. Marathi as spoken in Vidarbha is not Marathi as spoken in Pune. A Konkani-influenced Marathi from coastal Maharashtra is different again. A model that recognises “standard” Marathi well will still fail on these regional dialects.

We have built a regionalisation layer in our voice pipeline. The same Marathi-language deployment, depending on the farmer’s location and the dialect features the speech model recognises, will respond with appropriately localised vocabulary and idiom. The model knows that “phool” in a cotton-belt context refers to early bloom, and that “phool” in a different agricultural context might mean something else.

This is dialect-aware language modelling rather than translation, a different engineering problem and roughly an order of magnitude more involved than running outputs through a translator. When a partner asks why the language line in our proposal costs what it costs, this layer is usually the answer.

Voice has to work on the device and connection the farmer actually has

A farmer in rural India is typically on a budget Android handset two to four years old, on a 3G or weak-4G connection, in conditions that include wind, ambient farm noise, and occasionally a tractor running ten metres away. The microphone is not studio-grade. The audio that arrives at our server is degraded in specific ways.

We trained our speech recognition models on degraded audio, not on clean studio recordings. The public speech-recognition models commonly available perform worse on the audio we actually receive than they do on the benchmarks they advertise against. In our testing, training for the audio we actually get has been the difference between transcription accuracy in the mid-70s and in the low 90s under production conditions. Early numbers from a small deployment base, but the gap is not subtle.

The same logic applies to the response generation. Synthesised voice has to be intelligible over a low-bandwidth audio call. It has to be slower than a typical news-broadcasting voice (farmers are doing other things while listening), with explicit emphasis on key words like dosages and rates. It has to retain regional accent characteristics that make the farmer feel they are talking to someone from their part of the country, not to a generic Indian English newsreader.

Three live examples

A cotton farmer in Adilabad, in northern Telangana, asks about the pink bollworm threshold for spray application. The query comes in Telangana Telugu, which does not sound much like the coastal Telugu most speech models are tuned on. The assistant recognises the dialect features, retrieves the regionally appropriate threshold (about 15% away from the generic extension publication’s figure), and responds with the application rate and product class permissible in Telangana, all in voice, in the farmer’s dialect, in under three seconds.

A vegetable grower in Tamil Nadu asks about a fungicide for blight on tomatoes. The query comes in Tamil with a regional pest term. The assistant recognises the term, maps it to the standardised disease entity, retrieves the Tamil Nadu state’s specific advisory (which differs from neighbouring Karnataka’s by the product class recommended), and delivers the answer with brand-specific product mentions because the farmer is interacting through a specific agribusiness’s deployment.

A wheat farmer in Punjab asks about irrigation timing during a fluctuating winter weather pattern. The query is in Punjabi. The assistant retrieves the local weather forecast, the soil-moisture state from the regional aggregate (since the individual farmer’s soil sensor data is not available in this deployment), and a Punjab-specific irrigation schedule. The response is delivered in Punjabi voice, with a specific recommendation tied to the next 72 hours.

None of these farmers typed anything, which is the point. A text-only deployment never gets to the first answer.

How this changes the scope

The phrase we use internally is “voice-first, not chat-first,” and every product specification at KissanAI runs through it as a filter. Is this experience workable for a user who cannot read? Is the answer’s pacing right for a user who is listening, not skimming? Does the dialect coverage match the geographic deployment? When one of those answers comes back no, the spec goes back for another pass.

We do not claim voice-first is the only correct paradigm. Different geographies and different user bases have different defaults. Mid-tier urban users in metros often prefer text, and enterprise users on the agribusiness side want web dashboards. The framing applies specifically to the farmer-facing layer of a smallholder agricultural deployment, and for that layer we treat voice as a baseline requirement, budgeted the same way the agronomy content itself is budgeted.

Most of the scoping debates we sit in are really a quieter question: which users is the business willing to lose? Put voice in the first release and that question mostly stops coming up.

Sources

  1. KissanAI · KissanAI KissanAI