
Key Takeaways
Sekargas Hamilton, an independent GAFTA and FOSFA accredited inspection company in the Baltic region, ran a structured head-to-head test: GrainODM against four of its own experienced technicians, on 10 wheat samples across 21 impurity categories.
GrainODM returned each analysis in about 3 seconds, against 8 minutes for a rushed manual test and roughly 30 minutes for a full EN 15587 analysis.
The AI was the most self-consistent participant in the test. Its impurity totals stayed within a 1.14pp band across all samples, tighter than any of the four technicians (1.43pp to 4.18pp).
In 7 of 21 impurity categories, the AI already measures inside the noise band of the human reference.
The test also measured the reference itself: the four technicians disagreed on quality class in a third of cases, and differed by up to 5.69pp on sound grain content in the same sample.
Six categories, mostly visually similar damaged kernels, still carry a measurable bias. Until targeted calibration closes them, GrainODM is positioned as a pre-screening and audit-trail layer, not an autonomous replacement for the lab.
Nikolaj Stankevič, Head of Inspections at Sekargas Hamilton, reviews the pilot on the record: data transparency, traceability, and faster access to results are where he sees the value for their clients.
Most AI case studies are written by the vendor. This one starts with a company whose entire business is being the neutral party.
Sekargas Hamilton, UAB is one of the leading independent inspection and quality control companies in the Baltic region. Based in Klaipėda and part of the international J.S. Hamilton Group, it provides sampling, laboratory testing, cargo inspection and certification for agricultural and other commodity sectors, accredited to GAFTA, FOSFA International and ISO/IEC 17025. When an exporter and a buyer disagree about a cargo, Sekargas Hamilton is often the party whose numbers settle it.
Our clients are exporters, importers, traders, processors and logistics companies, for whom a reliable and independent quality assessment is critical. We work every day in an environment where data accuracy, traceability and trust are essential.
- Nikolaj Stankevič, Head of Inspections, Sekargas Hamilton
That is a demanding place to test an AI system. So we did the test their way: same samples, same standard, their technicians, their laboratory. This post covers both halves of the story: the review from the inspection side, in Nikolaj’s own words, and the measured results of the head-to-head test, including the parts that did not go our way.
Why an inspection company started looking at automation
Nothing about the pressure Nikolaj describes is unique to the Baltics. It is the same pressure every quality function in the grain trade is under: more data expected, faster, with a shorter tolerance for error.
Cargo and quality control processes are becoming increasingly complex, and clients expect faster access to information and greater data transparency. A large share of these processes has historically relied on manual data collection and administration, so we saw an opportunity to increase efficiency, reduce the risk of errors and ensure even better data traceability throughout the inspection process.
- Nikolaj Stankevič, Head of Inspections, Sekargas Hamilton
The important word there is traceability, not speed. In an accredited inspection environment, a result that cannot be reconstructed later is worth very little. That framing shaped what we were asked to prove.
What they looked for in a partner
We were looking for a partner who understands not only technology but also the specifics of grain trading and quality control. It was important to us that the solution matched real inspection processes, ensured data integrity and traceability, and could be developed together with business needs.
- Nikolaj Stankevič, Head of Inspections, Sekargas Hamilton
This is the part vendors tend to underestimate. An inspection lab does not buy a camera. It buys a method it will have to defend to a client, an auditor, and occasionally an arbitration panel. Our job in the pilot was less about demonstrating a model and more about fitting inside a procedure that already works.
The test: 4 technicians, 10 samples, 21 categories
The pilot ended with a structured head-to-head comparison, designed by the laboratory rather than by us.
- 10 wheat samples, each analysed across 21 impurity categories
- 4 experienced Sekargas Hamilton technicians, plus GrainODM
- Roughly 840 individual measurements
- Method: EN 15587, in two modes
- Full analysis, ~30 minutes per sample
- Rushed analysis, ~8 minutes per sample, simulating peak-season conditions
- GrainODM: ~3 seconds per sample
- Each technician performed an equal number of tests in both modes, so mode effects could be separated from individual skill
- One sample was given twice to the same technician, as a repeatability check
- The reference value for every measurement was the average of the four technicians


A wheat sample from the Sekargas Hamilton test: the raw plate versus every kernel segmented and classified by GrainODM.
There is no independent ground truth in impurity analysis. There is only the human reference, which is why the test measured the technicians against each other as carefully as it measured the machine.
Result 1: the AI was the most consistent participant, by a clear margin
Every impurity fraction in a sample should sum to roughly 100%. How tightly a participant holds that line across ten different samples is a direct measure of internal stability.
| Participant | Range of sample totals | Spread |
|---|---|---|
| GrainODM | 98.87% – 100.01% | 1.14pp |
| Technician 1 | 98.57% – 100.00% | 1.43pp |
| Technician 4 | 98.20% – 100.36% | 2.16pp |
| Technician 2 | 97.66% – 100.48% | 2.82pp |
| Technician 3 | 95.60% – 99.78% | 4.18pp |
GrainODM was the tightest participant in the test. That matters operationally more than it sounds: it means the same consignment measured twice, on two different days, by two different shifts, produces the same number. There is no “difficult morning” factor in a model.
It is also worth being precise about what this metric is not. Consistency is not accuracy. A system can be reliably stable and still be reliably off. Which brings us to the next result.
Result 2: on accuracy, the average technician was closer to the reference
Measured as mean absolute error against the four-technician average:
| Participant | MAE vs. reference | Measurements |
|---|---|---|
| Technician 2 | 0.180pp | 209 |
| Technician 1 | 0.223pp | 210 |
| Technician 4 | 0.229pp | 210 |
| Technician 3 | 0.266pp | 210 |
| GrainODM | 0.491pp | 208 |
We are publishing this table because leaving it out would make the rest of the post untrustworthy. On this test, on this grain flow, GrainODM sat further from the human mean than any of the four technicians did.
Two pieces of context belong next to that number, and neither of them cancels it.
First, the average spread between the four technicians on a single measurement was 0.66pp - larger than GrainODM’s deviation from their mean. The reference itself carries substantial noise.
Second, the aggregate figure hides very large differences between categories. Which is the actually useful finding.
Result 3: 7 of 21 categories are already at laboratory level
In a third of the impurity categories, GrainODM’s measurements are practically indistinguishable from the technicians’:
| Impurity | Technician average | GrainODM | Difference |
|---|---|---|---|
| Ergot | 0.001% | 0.008% | 0.01pp |
| Stones | 0.007% | 0.029% | 0.02pp |
| Husks | 0.061% | 0.025% | 0.04pp |
| Maize | 0.415% | 0.419% | 0.05pp |
| Oats | 0.227% | 0.251% | 0.05pp |
| Other impurities | 0.124% | 0.114% | 0.08pp |
| Barley | 0.989% | 0.941% | 0.08pp |
These are not the easy categories. Ergot and stones are exactly the contaminants an inspection company cannot afford to miss, and the AI measures them inside the human noise band today.
Result 4: where it still gets things wrong
Six categories carry a bias large enough to matter commercially:
| Impurity | Technician average | GrainODM | Bias |
|---|---|---|---|
| Sound grains | 88.93% | 86.22% | -2.70pp |
| Darkened grains | 0.354% | 2.620% | +2.27pp |
| Small / shrivelled kernels | 1.588% | 0.526% | -1.06pp |
| Damaged grains | 0.059% | 0.738% | +0.68pp |
| Fusarium | 0.452% | 0.983% | +0.53pp |
| Triticale | 0.624% | 1.140% | +0.52pp |
The pattern is not random, and that is the encouraging part. Darkened, damaged and fusarium-affected kernels are one visual family - overlapping colour and texture cues that the classifier is currently separating too aggressively. Sound grains then absorb the residual as a mirror image of those errors.
There is a directional consequence worth stating plainly: the bias is one-sided against the supplier. Where GrainODM and the technicians disagreed on a quality class, the AI assigned the stricter class in the large majority of cases. A system that errs toward rejecting good cargo is not a neutral system, and it is not something an inspection company can deploy as a final verdict.
The fix is not a general model improvement. It is facility-specific training data: 50 to 100 real samples from the Sekargas Hamilton grain flow, weighted toward those categories, then targeted fine-tuning and a re-test on the same protocol with a statistically adequate sample count. That work runs 4 to 8 weeks.
Result 5: the reference has more noise than the industry admits
The most uncomfortable finding of the test was not about the AI.
- The four technicians agreed on the quality class in only 66% of cases. In a third of cases, at least one disagreed with the rest.
- On sound grain content, the spread between technicians on the same sample reached 5.69pp (94.16 / 88.60 / 91.20 / 88.47).
- On one sample, the fusarium class assigned by the four technicians was 1, 1, 4, 4: two experienced professionals called it best quality, two called it worst. GrainODM returned 1, agreeing with the first pair.
- A rushed 8-minute test was, in this sample, marginally more accurate than the full 30-minute one (MAE 0.188pp vs 0.222pp). At n=10 that may be a real focusing effect or an artefact, but it does not support the assumption that more time automatically means a better number.
None of this is a criticism of the laboratory. It is an honest picture of what visual impurity analysis is: a skilled human judgement with real variance, performed under time pressure, on a physically heterogeneous material. It is also the reason a stable, documented, repeatable measurement has value even when it is not yet the most accurate one in the room.
What would a 3-second pre-screen be worth at your throughput?
A full EN 15587 analysis runs about 30 minutes, a rushed one about 8. Estimate what an automated first pass could recover in lab hours at your own intake volume.
Try the ROI calculator →Where this leaves the deployment
Given the results above, the honest positioning is narrow and specific.
What GrainODM is used for today:
- Pre-screening at intake. A 3-second measurement on every load, triaging consignments into clear, borderline, and check-manually. Technician time goes to the samples that need judgement.
- Continuous monitoring. Every consignment gets a measurement and a photograph, rather than only the ones selected for full analysis.
- Audit trail. A timestamped, reopenable record with every object segmented and classified. This is documentation that manual grading simply did not produce.
What it is not used for:
- An autonomous replacement for laboratory analysis, particularly in the categories with a known bias. Those stay with the technician until calibration closes them and a re-test confirms it.
The review from the inspection side
Nikolaj’s assessment of the pilot process itself, on the record:
The pilot stage was constructive and results-oriented. The GrainODM team worked actively with our specialists, responded quickly to observations and sought to understand our real working processes. We value the fact that the solution was improved on the basis of practical user experience, and that communication throughout the project was open and effective.
- Nikolaj Stankevič, Head of Inspections, Sekargas Hamilton
On where the value lands for their own clients:
We see the greatest benefit for clients in data transparency, faster access to information and better traceability. Digitalised processes make it possible to share results more quickly, ensure consistent data management and provide more confidence in the decisions being made. In grain trading, where quality data often becomes an important part of commercial decisions, this is a significant advantage.
- Nikolaj Stankevič, Head of Inspections, Sekargas Hamilton
And his answer to the question we asked last, which is the one most quality managers are actually sitting with:
Process automation today is becoming not a competitive advantage but a natural direction of business evolution. The most important thing is to start from clearly identified processes and to choose a partner who understands the specifics of your operation. Our experience shows that properly implemented digitalisation increases efficiency, reduces the administrative burden, improves data quality and creates greater value for clients.
- Nikolaj Stankevič, Head of Inspections, Sekargas Hamilton
What we took away from it
A vendor-run demo would have shown the seven categories and stopped. A test run inside an accredited inspection laboratory showed all twenty-one, and the six that are wrong are more useful to us than the seven that are right.
Three things we would repeat at any facility considering this:
- Test against multiple technicians, not one. A single reference hides the variance that determines what “accurate” can even mean at your site.
- Report per category, not in aggregate. A single headline accuracy figure is close to meaningless when category sizes span four orders of magnitude.
- Deploy to the use case the data supports. Pre-screening and traceability are defensible on this evidence today. Autonomous grading in biased categories is not, and claiming otherwise is how pilots turn into disputes.
The calibration work at Sekargas Hamilton continues on real samples from their grain flow, with a defined re-test at the end of it.
Related reading: our validation against five lab technicians on 600+ wheat tests, the Allive hemp inspection case study, and a primer on EN 15587 besatz analysis for wheat.
Quotes from Nikolaj Stankevič were provided in writing in Lithuanian and are published here in translation. Test data: Sekargas Hamilton final testing round, April 2026, 10 wheat samples across 21 impurity categories, EN 15587.
Frequently Asked Questions
Sekargas Hamilton, UAB is one of the leading independent inspection and quality control companies in the Baltic region, based in Klaipėda, Lithuania, and part of the international J.S. Hamilton Group. It provides sampling, laboratory testing, cargo inspection and certification for agricultural and other commodity sectors, with GAFTA, FOSFA International and ISO/IEC 17025 accreditation.
Ten wheat samples were analysed across 21 impurity categories by four experienced Sekargas Hamilton technicians and by GrainODM, producing roughly 840 individual measurements. Testing followed EN 15587 in two modes: a full analysis of about 30 minutes and a rushed analysis of about 8 minutes that simulates peak-season conditions. Each technician performed an equal number of tests in both modes. The reference value for each measurement was the average of the four technicians.
No. On the accuracy metrics used in this test, the average technician was closer to the four-technician mean than GrainODM was. Where the AI led was consistency and speed: it produced the tightest impurity totals of any participant and returned results in about 3 seconds instead of 8 to 30 minutes. In 7 of the 21 categories its measurements already sit inside the human noise band.
The main gaps are in visually similar damaged kernels: darkened grains, damaged grains and fusarium-affected kernels, plus small or shrivelled kernels and triticale. These share colour and texture cues with sound grain, which is exactly where an image-based classifier needs facility-specific training data. The calibration plan runs 4 to 8 weeks on real samples from the Sekargas Hamilton grain flow.
As a pre-screening and documentation layer rather than a final verdict. Every load gets a measurement, an image and a timestamp in about 3 seconds, which triages consignments into clear, borderline and check-manually, and creates an audit trail that manual grading never produced. Categories with a known bias stay with the technician until calibration closes them.
Continue reading
From Tweezers to Imaging: Grain Impurity Analysis at a GAFTA-Approved Port Laboratory
Case StudiesHow AI Cut Allive's Hemp Seed Inspection From 30 Minutes to Seconds
Case StudiesAI vs. 5 Lab Technicians: What We Found After 4 Months and 600+ Wheat Tests
Case StudiesManual Wheat Sprout Detection Fails: AI vs. Human Eye
The New Standard in Grain Purity Analysis
Data, not guesswork. Learn how GrainODM sets a new benchmark for digital grain inspection.

