Skip to content
Anora

Learn · Foundations

Scaling a method

Everything in this series so far produces a number. This piece is about the harder question underneath all of them: how far does a result travel before it quietly stops being true?

Loading ground…
Three districts the same crop, different ground
 
 
 
Three districts

The same crop, on ground that is not the same

Three neighbouring districts of farmland. The crop is the same, the sensor is the same and the date is the same. What differs is the ground: field sizes, how much of each plot the canopy has actually closed over, and how bright the soil underneath is.

None of that is exotic. It is the ordinary variation between one district and the next, and it is enough.

Ground truth

A number is not a finding until somebody has stood in it

At ten metres a pixel is a hundred square metres of the world, and it very rarely contains one thing. It holds part of a plot, a strip of track, a hedge and some bare ground, and the sensor returns a single value for all of it.

So the only way to know whether a method works is to go and look. Somebody walks the ground with a GPS and records what is actually there, and those records become the standard the satellite answer is scored against.

Validated here

Tuned against the survey, it works almost perfectly

Here a threshold has been fitted against the surveyed ground in the first district alone, exactly as it would be in a real programme. Every value above it is called crop and everything below it is not.

It performs extremely well. Both figures beside the picture come from comparing the classified result against the survey, and on this district there is almost nothing to argue with. This is the moment where a report gets written and an accuracy gets quoted.

Scaled next door

Carry it to the next district and it starts to slip

The same rule, unchanged, applied to the neighbouring districts. Red marks every pixel where the satellite answer and the ground disagree. In the second district the errors are scattered and few. In the third they cluster, and they cluster in a particular kind of field.

Nothing was done wrong. The threshold was fitted honestly and applied consistently. It simply encodes an assumption about how much soil shows through a canopy, and that assumption travelled to a district where it is not true.

The figure that matters

The headline accuracy holds while the answer rots underneath it

This is the part worth slowing down for. Overall accuracy in the third district still looks respectable, because most of the landscape is bare ground and bare ground is still being called correctly. That single number would pass a review.

Now look at how much of the crop was actually found. Roughly one field in six has gone missing, and it is not a random sixth: it is the sparse and stressed crop, which is precisely the crop anyone commissioning the work wanted to know about. A headline accuracy figure hid the only failure that mattered.

In practice

So we tell you where the numbers stop being trustworthy

There are only really three defences and none of them is clever. Stratify before sampling, so that survey effort lands in each distinct kind of ground rather than clustering where access is easy. Re-validate at the boundaries, where conditions change, instead of assuming the middle is representative. And quote accuracy per class, because that is the number that moves first.

Scaling satellite work across a region is genuinely cheap, and that is exactly why it deserves the scepticism. The method costs almost nothing to run over the next district. Knowing whether it still means anything there is the part you are actually paying for.

Why this is the piece that matters commercially

Anyone can run an index over a region and hand back a map. The map will look authoritative and it will be partly wrong, in ways that are invisible unless somebody has checked. The work that separates a measurement from a picture is the sampling design, the field validation and the honesty about where the result degrades.

That is also why we quote accuracy per class and per stratum rather than as a single headline. It is a less impressive number to put on a slide, and it is the one that tells you whether you can act on the result.

Where we use this

Everywhere, because it is not a technique so much as the condition on all the others. District and catchment scale agricultural work, where a method fitted in one agro-ecological zone is asked to cover several. Smallholder supply chain verification, where plots are small, mixed and highly variable between regions. Any programme that starts as a pilot in one area and is then expected to cover a country.

The threshold above is fitted in your browser against the first district only, and every accuracy figure is measured by comparing the result against the modelled ground. The landscape is modelled rather than observed.

Put this to work
Bring us a question about your ground
If you need a result that holds up across a region rather than in the one place it was tested, this is the part of the work we take most seriously.
Get in touch →