Augere · An Auctus Agri working paper AA · 2026 · 005

Volume I, Issue 5 · 23 August 2026 · 14 min read · Applied Research · doi:10.5281/zenodo.22070535

Should you buy that agricultural AI? Eight questions before you sign.

Before you buy an artificial intelligence product for your farm or agribusiness, ask what decisions it changes, how often you make those decisions, what happens when it is wrong, and whose farms its data came from. If the vendor cannot answer those four questions clearly, the software's sophistication will not save the investment. Fit, not intelligence, determines whether agricultural artificial intelligence pays off.

Jackson Mambozoukuni · Auctus Agri · ORCID 0009-0009-9343-1910

Keywordsartificial intelligence · technology adoption · decision making · agribusiness management · extension · farm management

By artificial intelligence this paper means software that learns patterns from data rather than following fixed rules. The claim above is not a dismissal of that technology. It is already doing real work in agriculture: spotting disease before a scout would, guiding sprayers to treat weeds and skip crop, and forecasting yields with useful accuracy. Innovation in the sector has concentrated heavily in systems that replicate perception and judgment rather than muscle, which is precisely why so much of it lands on decisions rather than tasks. The difficulty is that these successes and the failures sitting beside them look nearly identical in a sales demonstration.

"The model can be right and the investment can still be wrong."

Most agricultural AI works. So why do so many projects fail?

The adoption picture is sharply uneven, and it has been for longer than the current excitement. Guidance and auto-steer systems now cover well over half of the acreage of major United States row crops, while information-intensive tools such as variable rate application have stalled far below expectations for two decades. The same gradient runs through livestock: large dairies adopt advanced technology at several times the rate of small ones. Measured use of artificial intelligence across agriculture as a whole remains low relative to other industries, and a federal assessment of precision agriculture identified high up-front costs and uncertain payoffs as the top reasons.

The usual explanations are all real. Rural connectivity is patchy. Capital is scarce. Farmers reasonably worry about who ends up owning their data, and surveys find that most cannot say what their technology agreements permit. Managers distrust software that produces an answer without explaining itself, and uncertain returns are hard to finance. There is also a timing problem: venture investment in agrifood technology peaked at $51.7 billion in 2021 on the promise of transformation, then fell below $16 billion by 2024. The sector has entered a colder, more demanding phase in which buyers want evidence of return rather than evidence of capability. Vendors who thrived by demonstrating what their software could do are now asked what it earned. Many cannot say.

Agriculture is not the outlier here; it is the rule. Industry surveys of enterprise artificial intelligence keep reporting the same shape: the large majority of corporate pilots produce no measurable profit-and-loss impact, not because the models are weak but because the tools are never embedded in real workflows. The pilot trap, a cycle in which a technology is repeatedly shown to work in demonstrations but never becomes an ordinary, profitable part of the operation, is a general disease of buying intelligence without buying fit. None of this should surprise an economist: nearly seventy years ago, the classic study of hybrid corn showed that a plainly superior technology diffused at a speed set almost entirely by its profitability in each district, not by its ingenuity. Farmers adopted where and when it paid, and not before.

The question that matters is fit, not sophistication

Consider what a vendor actually tells you. The statement "Our model detects leaf disease with 98% accuracy" is about the software in isolation. It is probably true. It is also nearly useless on its own, because it says nothing about whether that accuracy will turn into money on your farm.

To turn accuracy into money, a chain of events has to be completed. The software has to change a decision you actually make. Changing that decision has to improve the outcome. Moreover, the improvement has to be worth more than the cost of producing it. Break any link, and the software can be flawless while the investment returns nothing. A disease detector on a farm that already sprays preventively on a fixed schedule makes no decision. A yield forecast that arrives after the marketing decision has been made improves no outcome. A model that costs more to feed and maintain than the losses it prevents fails the third test. The useful consequence is that those conditions are largely visible in advance: a manager can move the evaluation upstream of the cheque, from "is this impressive?" to "will this work here?"

Four conditions that predict whether AI pays

Frequency. Machine learning repays its fixed setup cost through repetition. It is no accident that the clearest commercial success in agricultural AI to date, camera-guided targeted spraying, sits on a decision made millions of times per pass: one manufacturer reports average herbicide savings of 59% across more than a million acres. Stated properly, the economics is that the expected value across the decision opportunities a model informs must recover its fixed and recurrent costs. Frequency is the commonest way that happens; it is not the only way. A once-a-year cultivar, insurance or marketing decision can carry enough value in a single use, but a low-frequency proposition must clear the same cost bar from far fewer chances.

Reversibility. No model is right every time, so what matters is the cost of its mistakes. Where a wrong call is cheap to correct, automation is attractive and moderate accuracy suffices. Where a wrong call destroys a crop, denies a farmer credit, or puts a machine among people, the burden of proof rises steeply, and a human must stay in the loop. Reversibility, not accuracy alone, should set the bar a system must clear.

Representative data. A model is only as transferable as the data behind it. A disease classifier built on large, well-instrumented farms in one region frequently underperforms on different soils, cultivars, and management. This is the most under-examined condition, precisely because a demonstration on the vendor's data reveals nothing about performance on yours. Ask whose farms the training data came from, and ask where yours will go.

Total cost and ownership. The purchase price is a fraction of the total cost: integration, connectivity, sensor upkeep, data cleaning, staff training, and periodic retraining as conditions drift. Engineers call the accumulating upkeep hidden technical debt, a bill that grows quietly until someone has to pay it. And the most common failure of a deployed model is not technical but organisational: nobody owns keeping it running after the pilot. Ask who runs it in year two, and name them.

Notice what is absent from that list. Nothing about the algorithm, accuracy percentages or model architecture. Those things matter, but algorithmic sophistication is not the buyer's first question. The four conditions are the buyer's problem from the start, and the buyer is the only person who can assess them.

Diagram: four checkable conditions (frequency, reversibility, representative data, cost and ownership) feed eight questions asked in one vendor conversation, producing one of three verdicts: green light, pilot carefully, or walk away.
Figure 1  Four checkable conditions determine whether agricultural artificial intelligence pays off. A failure on the first or third question caps the verdict regardless of the others.

Eight questions a manager can actually ask

To be useful at the moment of decision, those conditions must become questions answerable from a single vendor conversation. Each is scored, and the total maps to one of three verdicts: green light, pilot carefully, or walk away for now.

#QuestionWhat it testsCondition it servesScore
1What decision does this product actually change?Is a real decision affected?The decision premise the chain starts from−1 to 2
2How often is that decision made?Frequency: does the benefit compound?Frequency0 to 2
3If the product is wrong, how bad is it?Reversibility; stakes of errorReversibility−1 to 2
4Whose farms was the model trained on?Representative training dataRepresentative data−1 to 2
5What evidence of performance can they show?Proof beyond the demonstrationVerification of the other answers0 to 2
6Do you know the true total cost, not the price?Total cost of ownershipCost and ownership: the money−1 to 2
7Who runs it after the pilot ends?Operational ownershipCost and ownership: the people0 to 2
8Does it need connectivity and the records you have?Fit to the actual infrastructureCost and ownership: the infrastructure−1 to 2

Table 1  The eight questions. Totals run from minus five to sixteen; the maximum is 16.

Using the test takes one conversation: ask the eight questions, score each answer 2, 1, 0 or −1, apply the caps, and read the verdict. A proposition earns a green light if it scores 11 or more and no answer scores minus one. It earns a pilot carefully if it scores six or more and neither question 1 nor question 3 scores minus one. Everything else is a walk away.

The scoring is deliberately asymmetric, through three distinct mechanisms. Two hard caps: a minus one on question 1 (no identifiable decision) or question 3 (irreversible errors) makes the verdict walk away regardless of the total. A green-light veto: a minus one on any question blocks a green light but not a pilot, because resolving a single identified dealbreaker is exactly what a pilot is for. And the ordinary numerical thresholds, which rank everything the caps and the veto let through. The cap on question 3 is deliberately unconditional: strong evidence lowers an error rate, but it does not undo an irreversible error. A strong average cannot rescue a proposition that fails a fundamental test; one broken link nullifies the rest of the chain. The full paper's sensitivity appendix shows that moving the thresholds changes almost nothing, while removing the caps and the veto changes the verdicts, which is exactly where an instrument of this kind should be robust and where it should not.

The test scores fit, not merit. A low score does not mean the product is bad; it means the product is poorly matched to this operation, now. The same tool can be a walk-away for one farm and a green light for its neighbour.

Testing the filter on four common propositions

An instrument that returns the same answer for everything is worthless. The test was applied to four archetypal vendor propositions constructed from publicly described product categories; the examples are illustrative rather than empirical, but they show the test discriminating, and doing so for different reasons. One caution: on the definition this paper uses, two of the four may not be artificial intelligence at all. They are scored anyway, because they are sold as AI, and because the test is deliberately label-indifferent.

Vendor propositionTotalIndicative cost (relative)Verdict
Camera-guided selective sprayer10High: hardware retrofit plus subscriptionPilot carefully
Satellite imagery dashboard subscription6Lowest of the four: small annual subscriptionWalk away
Credit-scoring tool for input finance4Moderate: fee per assessmentWalk away
Autonomous weeder, specialty crops13Highest of the four: major capital purchaseGreen light

Table 2  Summary of the four worked examples; per-question scores are in the full paper. Two propositions fail for opposite reasons; the most expensive passes; the cheapest is refused.

The sprayer scores well on frequency and ownership, but the evidence is the vendor's own and the training data comes from other regions: a pilot, with a local trial before scale. The satellite dashboard looks harmless, is cheap, and is refused, because the vendor cannot name the decision it changes; managers subscribe, look at the maps for a season, and change nothing. The credit-scoring tool fails differently: denying a farmer credit is irreversible and the model was trained on other farmers in other places. High stakes and unrepresentative training data are the combination the test is designed to catch. The autonomous weeder earns its green light on frequent, recoverable errors, independent evidence, known running cost and a named operator, and it is also the most expensive of the four. Price is not the variable that matters.

For advisers, lenders and vendors

The eight questions were written for a manager, but they may be most valuable in the hands of the people managers turn to for advice. Extension officers and farm advisors are routinely asked some version of "Should I buy this?" and are rarely equipped to answer, because evaluating a machine learning product appears to require technical expertise. It does not: every one of the eight questions can be answered without knowing what a neural network is. An advisor who asks what decision the product changes, who trained it, and who will run it in year two is conducting a more rigorous evaluation than a technical review of the algorithm would. Lenders and insurers face the same problem from the other side: a proposition that changes no decision will not generate the cash flow to service the loan taken out to buy it. And the questions are a useful discipline for vendors. A serious vendor welcomes them and has answers ready. A weak one deflects toward accuracy statistics. That asymmetry is itself diagnostic, and costs nothing to observe.

What this test cannot do

An instrument built to resist overclaiming should not overclaim. This is a structured heuristic, not a predictive model. Its weights are reasoned rather than calibrated against outcome data. It scores fit; it does not verify a vendor's technical claims, which still requires scrutiny and, where the stakes are high, an independent trial. It assumes a decision maker with the discretion to buy or decline; smallholders operating through cooperatives face collective and institutional questions it does not reach. And it is a snapshot: a walk away today may be a green light in two seasons, which is why the instrument is built to be rerun.

So what?

The debate about artificial intelligence in agriculture is conducted mostly between people selling it and people dismissing it. Both positions ask the wrong question: whether the technology works. It does, sometimes, under conditions that can be named. The right question is narrower and far more useful: will it work here, on this decision, at this price, with this data, run by this person? That question can be answered before the money moves. Answering it well is ordinary management discipline, applied to a technology that has been unusually successful at avoiding it. If a vendor cannot tell you what decision their product changes, there is not yet one worth paying for. Everything else in the paper is an elaboration of that sentence.

The full working paper, permanently archived

The complete paper, with the scoring rule stated in full, per-question scores for all four worked examples, the sensitivity analysis and the reference list, is deposited on Zenodo under a CC BY 4.0 licence: doi:10.5281/zenodo.22070535.

The companion instruments are live, free, and run entirely offline; nothing leaves your device. The Reality Filter is the eight-question test as an interactive browser instrument: answer the questions, read the verdict, copy it to your clipboard. The Reality Filter Scorecard is the spreadsheet edition, scoring up to three competing propositions side by side with a transparent scoring key. Both implement this paper's published scoring rule exactly, verified against its four worked examples. The five-question pre-purchase quick screen joins them at its release.

Evaluating a specific proposition? Talk to the practice

↓ Download the full paper as PDF → Run the Reality Filter in your browser ↓ Download the Excel Scorecard

All free · no email required · the instruments run offline and collect nothing

Recommended citationMambozoukuni, J. (2026). "Should You Buy That Agricultural AI? Eight Questions Before You Sign." Auctus Agri working paper, Augere AA · 2026 · 005. https://doi.org/10.5281/zenodo.22070535

Augere is the applied-research imprint of Auctus Agri, a South African agribusiness practice. This paper is part of the firm's technology-appraisal line, which tests whether agricultural technology propositions are supported by the decisions they change, the evidence behind them and the economics of their use.

Book a 30-minute scoping call →