Founding the AI AgriBench consortium
Why we joined the Center for Digital Agriculture and the Extension Foundation in launching AI AgriBench, a benchmarking consortium for generative AI in agriculture. What it is trying to do, what we are contributing, and why agricultural AI needs this work.
Today we are joining the Center for Digital Agriculture and the Extension Foundation to launch AI AgriBench, a generative AI benchmarking consortium for agriculture. The announcement is going up across our partners’ channels and ours; this post is the longer version of what we are committing to and why.
Why a consortium
In the past 18 months a wave of generative AI products has hit agriculture. Advisory chatbots, vision diagnosis apps, irrigation planners, market forecasters, voice copilots for farmers. Some of them are well-built. Many are not. Sitting at the table with agribusinesses and policymakers who are trying to buy or fund these tools, I can report that there is no shared way to tell which is which.
In every other domain that has navigated this kind of transition (medicine, finance, engineering) the answer was an institutional layer that defined what “good” meant and made the definition public. Clinical trial standards, financial reporting standards, engineering codes. These bodies started as coordinations among practitioners who realised they could not let the market sort it out without expensive failures first.
Agricultural AI is at exactly that moment. AI AgriBench is an attempt to start that layer.
Who is at the table
The Center for Digital Agriculture at the University of Illinois at Urbana-Champaign brings the academic anchor. Vikram Adve and the CDA team have been building research-grade evaluation methodology for digital agriculture for years, and the methodology layer of AI AgriBench will draw heavily on their work, with public peer review at the centre.
The Extension Foundation brings the deployment-side institutional credibility. Extension services are the layer of agriculture that historically certified what was safe to recommend to farmers; they are the legitimate body to ask “is this AI tool’s advice on a par with what an extension officer would have given?” Their involvement is what gives the consortium standing to be taken seriously by buyers.
KissanAI brings the technical contribution. The same evaluation methodologies that gate our internal Dhenu releases (accuracy, source grounding, regional appropriateness, language fidelity) are being contributed to the consortium as starting material. We are also contributing a portion of our curated Indian-agriculture evaluation corpus, anonymised and structured into evaluation pairs that other teams can use.
We are not the only technical contributors; the consortium is intentionally multi-vendor, with pilot members across the agribusiness, extension, and AI infrastructure spaces. A layer like this cannot be owned by any one vendor, including us, without losing the legitimacy that makes it useful.
The first year
The founding year starts with methodology. What does a “good answer” mean for an agricultural AI? An LLM-judged response is one signal. A grounded response that cites an extension publication is a stronger one. A response that matches what a credentialed agronomist would have given is the strongest. The methodology working group is layering these into a composite score and publishing the methodology in the open.
Benchmarks also need data. The consortium is curating publicly shareable evaluation datasets across major crop-region combinations: corn-soy-wheat in the US Midwest, cotton-rice-pulses in India, viticulture in southern Europe. The datasets include the questions farmers actually ask and ground-truth answers validated by extension-affiliated agronomists. Public datasets are how the field moves together rather than each vendor claiming its own private number.
Eventually, models evaluated under the consortium methodology will report into a public leaderboard. The leaderboard itself matters less than what it forces: a vendor can no longer claim accuracy without showing what it was evaluated on, a buyer can read the results and ask follow-up questions, and a vendor who refuses to participate will increasingly look like it has something to hide.
Boundaries
A word on boundaries. AI AgriBench is not a certification body: a high score means a model performed well on the evaluation, while deployment safety depends on operational factors the evaluation cannot reach (escalation paths to humans, monitoring in production, the deploying operator’s commercial alignment with farmer interest). Membership is open too. Any team building agricultural AI can evaluate against the public set and report results, and we expect competitors to do so, since a benchmark serves the field as a floor rather than serving one vendor as a moat. Nor should anyone expect quick wins here. Methodology development takes time, reference datasets take longer, and the first published evaluations are months out. Moving slowly on standards work is precisely what makes the standards trustworthy when they land.
Why KissanAI is in
There is a self-interested reason and an ecosystem reason. Both are true.
The self-interested reason: we believe Dhenu evaluates well against rigorous methodology. The cost of participating in a public benchmark is that the world finds out where Dhenu is weak. The benefit is that the world finds out where Dhenu is strong, in a way that does not depend on us asserting it. For a product like ours, built domain-deep and voice-first in a farmer’s own languages, a credible third-party evaluation moves more procurement conversations than any marketing claim.
The ecosystem reason: agriculture cannot afford to wait out the bad-products phase the way some industries did. The user is a farmer betting a season. A wrong recommendation at scale has real cost. A shared evaluation layer that helps the field separate good products from confident-sounding bad ones is, structurally, the way agriculture protects its own users from the rough edges of a new technology curve.
The consortium will publish the methodology draft and the initial reference dataset summaries through 2025. We will write about the parts of the methodology where we have specific technical input (the multilingual evaluation, the regional appropriateness scoring, the source-grounding metric) as those drafts harden.
If you are building agricultural AI, the consortium is open and the early-collaborator door is wide. Find us on the AI Alliance channels or directly through the Center for Digital Agriculture.
Sources
- Center for Digital Agriculture, University of Illinois · Center for Digital Agriculture Research
- Extension Foundation · Extension Foundation Research