Trust in AI for agriculture starts with benchmarks
An essay on why public benchmarks, not glossy demos, are how AI advisory services earn the trust of farmers and policymakers. The premise behind AI AgriBench, what evaluation we are doing, and what we are deliberately not.
There is no shortage of AI products being sold to agriculture this year. Voice assistants, vision diagnostic apps, irrigation planners, mandi forecasters. Some of them work. A lot of them do not. The frustrating part, sitting eighteen months into the commercial wave of agricultural AI, is that a buyer still has no shared way to separate them.
AI AgriBench was created for that gap. The work is slow and heavy on methodology, which is exactly what public evaluation looks like when it cannot be reduced to a demo.
A farmer is not running an A/B test
A farmer trying a new AI tool is not running an A/B test. They are betting a season. If the tool tells them to plant on a particular date or treat a particular pest with a particular product, and the recommendation is wrong, the cost is not “we’ll know next quarter.” The cost is a yield hit they will carry into the next planting cycle.
The way other domains earned trust at this stage was through standards bodies and public evaluation. Medicine has the FDA and a pharmacopeia of clinical-trial reporting standards; finance has GAAP and audit. Each of those bodies started as a coordination among practitioners who realised they could not let the market sort it out without a generation of expensive failures first.
Agricultural AI has no such body yet. AI AgriBench is the early version of one.
How the evaluation is being built
AI AgriBench is led by the Center for Digital Agriculture at the University of Illinois at Urbana-Champaign, with the Extension Foundation as institutional partner, KissanAI as a technical contributor, and a working group of agronomy researchers, AI engineers, and agribusiness practitioners. The structure is deliberate; the academic anchor and the extension network constrain the technical providers, and the reverse holds too.
The work starts with evaluation methodology. What does “accurate” mean for an agricultural AI? The working group is defining a layered evaluation rather than a single score: factual accuracy, source grounding, agronomic appropriateness, compliance with regional regulations, and language quality in the language of delivery, each scored independently and reported as layers. A model that is factually correct but recommends a product banned in the user’s state cannot hide behind an aggregate number.
Underneath the methodology sit reference datasets. AI AgriBench is curating publicly shareable evaluation sets across major crop and region combinations: corn-soy-wheat in the US, cotton-rice-pulses in India, viticulture in southern Europe. The datasets include both questions (the kind farmers actually ask) and ground-truth answers (validated by extension-affiliated agronomists). Public datasets are how the field moves together rather than each vendor claiming its own private number.
Eventually, models tested against the AI AgriBench evaluation will report into a public leaderboard. A leaderboard sounds like marketing until you notice what it forces: a vendor can no longer claim accuracy without showing what it was evaluated on, and a buyer (an agribusiness, a state extension service, a farmer cooperative) can read the results and ask the obvious follow-up questions. The dynamic this creates is the same one that publication norms create in medicine: the conversation moves from “trust us” to “show your work.”
KissanAI’s part
Our contribution starts with data. We have spent years curating Indian-agriculture content: extension publications, product labels, agronomy guides, real farmer queries in regional languages. A portion of that corpus, anonymised and structured into evaluation pairs, is being contributed to AI AgriBench as the India-specific reference set.
The other piece is technical methodology. The same models we evaluate internally before shipping a Dhenu checkpoint are being evaluated under the AI AgriBench methodology. The methodology is being refined against our results and others’. The intent is that by 2026 the same evaluation suite that gates a Dhenu release is also the suite a third party can run against any competitor model on the same questions.
We are not the only ones doing this, which is rather the idea; a consortium whose methodology depended on a single team would not deserve the name.
Where a benchmark stops
Some deliberate non-goals are worth stating.
AI AgriBench is not a certification. A high score does not mean the model is safe to deploy; it means the model performed well on the evaluation, while deployment safety depends on many things the evaluation does not cover (escalation paths to humans, monitoring in production, the operator’s commercial incentives). Certification is a longer conversation, and the benchmark is a precondition for that conversation rather than a substitute for it.
Nor will the benchmark cover everything. Some agricultural AI tasks resist clean evaluation: long-running advisory relationships with a single grower, where the value is in the cumulative trust, are not benchmark-shaped, and vision diagnosis of unusual symptoms on rare crops is hard to evaluate at scale. We are starting with the questions that have clean ground truth and expanding the perimeter over time.
Anyone building agricultural AI can evaluate against the public set and report results, including our competitors. A benchmark owned in practice by one vendor would not solve the problem.
The procurement problem
The trust problem is sharpest at the policy layer, where a state government procuring AI for extension services, or an agribusiness deploying it to a retailer network, or a donor funding a farmer-facing pilot, is being asked to make commitments without the kind of evaluation rigor any other domain would demand. A benchmark they can read and reason about is the artifact that lets these procurement and funding decisions become rigorous.
We will be at the 2025 Farm Progress Show in Decatur, Illinois, on August 27 with the rest of the consortium for a session on the work. For everyone else, the methodology drafts and reference dataset summaries are being published as they mature; the Center for Digital Agriculture is the right place to follow them.
A procurement team should be able to ask which questions a model was tested on, who reviewed the answers, and where it failed. The benchmark is meant to make those ordinary questions answerable.
Sources
- AI AgriBench Consortium: Center for Digital Agriculture · Center for Digital Agriculture, University of Illinois Research
- Extension Foundation · Extension Foundation Research