Three districts
The same crop, on ground that is not the same
Three neighbouring districts of farmland. The crop is the same, the sensor is the same and the date is the same. What differs is the ground: field sizes, how much of each plot the canopy has actually closed over, and how bright the soil underneath is.
None of that is exotic. It is the ordinary variation between one district and the next, and it is enough.
Ground truth
A number is not a finding until somebody has stood in it
At ten metres a pixel is a hundred square metres of the world, and it very rarely contains one thing. It holds part of a plot, a strip of track, a hedge and some bare ground, and the sensor returns a single value for all of it.
So the only way to know whether a method works is to go and look. Somebody walks the ground with a GPS and records what is actually there, and those records become the standard the satellite answer is scored against.
Validated here
Tuned against the survey, it works almost perfectly
Here a threshold has been fitted against the surveyed ground in the first district alone, exactly as it would be in a real programme. Every value above it is called crop and everything below it is not.
It performs extremely well. Both figures beside the picture come from comparing the classified result against the survey, and on this district there is almost nothing to argue with. This is the moment where a report gets written and an accuracy gets quoted.
Scaled next door
Carry it to the next district and it starts to slip
The same rule, unchanged, applied to the neighbouring districts. Red marks every pixel where the satellite answer and the ground disagree. In the second district the errors are scattered and few. In the third they cluster, and they cluster in a particular kind of field.
Nothing was done wrong. The threshold was fitted honestly and applied consistently. It simply encodes an assumption about how much soil shows through a canopy, and that assumption travelled to a district where it is not true.
The figure that matters
The headline accuracy holds while the answer rots underneath it
This is the part worth slowing down for. Overall accuracy in the third district still looks respectable, because most of the landscape is bare ground and bare ground is still being called correctly. That single number would pass a review.
Now look at how much of the crop was actually found. Roughly one field in six has gone missing, and it is not a random sixth: it is the sparse and stressed crop, which is precisely the crop anyone commissioning the work wanted to know about. A headline accuracy figure hid the only failure that mattered.
In practice
So we tell you where the numbers stop being trustworthy
There are only really three defences and none of them is clever. Stratify before sampling, so that survey effort lands in each distinct kind of ground rather than clustering where access is easy. Re-validate at the boundaries, where conditions change, instead of assuming the middle is representative. And quote accuracy per class, because that is the number that moves first.
Scaling satellite work across a region is genuinely cheap, and that is exactly why it deserves the scepticism. The method costs almost nothing to run over the next district. Knowing whether it still means anything there is the part you are actually paying for.