Why we think a photograph is good enough
An estimate from a photograph is a good guess, not a measurement. It is worse for some nutrients than for others.
Written and maintained by Mattias Geniar. Last updated .
- What the app does
- Puts a confidence badge on every meal, prints the assumption it made under each ingredient, and asks you a question instead of guessing what it cannot see.
- How much we believe it
- Enough to show a number, not enough to show it on its own. Portion size is where the error comes from, and a better camera does not fix that.
What the benchmarks say
The best public benchmark is ACETADA (Coburn et al., 2025). It has 806 meal images from a controlled feeding trial. Dietitians gave each item a weight in grams, checked against food weighed to 0.1 g. Against that, current vision models are off by about 166 to 211 kcal per meal on average. As a percentage, that is a mean absolute percentage error (MAPE) of 24 to 34%. For carbohydrate the error is 23 to 28 g (37 to 50%). Fridolfsson et al. (2025) tested 52 photographs against a calibrated scale and found an energy error of 35.8% for two of the three models tested.
In the study that produced Nutrition5k (Thames et al., CVPR 2021), nutritionists estimating weight from the same photographs were 41% out on average. Non-nutritionists were 53% out. Read that carefully before relying on it: it measures weight error, not energy error, and the sample is small, as the caption above says. A photograph is not accurate. Neither is a person looking at the same photograph.
What the app does with this is on the accuracy page: the confidence badge on each meal, the assumption printed under each ingredient, and the question it asks when it cannot see the cooking fat.
What actually changes the error
Not the prompt, and not the camera. Vedovelli et al. (2026) tested 40 vision-language models. The choice of model explained 99.6% of the difference in results. Taking photos from several angles gave no benefit (p = 0.182), and changing the prompt had no significant effect (all p > 0.05).
What does help is extra information. In Rodríguez-Jiménez et al. (2025), 195 dishes, carbohydrate error fell from 11.72 g (41.95% MAPE) with the image alone to 6.20 g (23.52%) once the model was told the ingredient amounts. Mu et al. (2025) found carbohydrate error dropping from 56.6% to 20.2% when the model was given the true food weight.
So the design follows the evidence. More angles do not improve the estimate in any measurable way, so the app does not ask for them. Portion and preparation are where the error is, so that is what it asks about. It asks once, and only in the two or three cases where the answer changes the number by more than the question costs you.
It is also why every ingredient carries the assumption it was estimated under. An estimate you can correct is worth more than one that is slightly better but that you cannot check.
The part we are least comfortable with
The problem is not average accuracy but variation. One preprint asked the same models about the same 13 photographs 26,904 times. Depending on the model, the answers varied by 2.4% to 11% on average (the median coefficient of variation). One plate of paella came back with anywhere between 55 g and 484 g of carbohydrate across repeats of the same image. The same preprint found that a model's own stated confidence had almost no relation to whether it was right. It is not peer reviewed. We cite it anyway because it is the only work that asks the question.
A figure we published and withdrew
We used to quote “18% median, 42% at the ninetieth percentile” from our own benchmark. That number comes from a real internal test over twenty meals. But it was measured against meals a person had already corrected by hand, not against weighed food. That is fine for comparing one model with another. It is not a measurement of accuracy, and we should not have presented it as one. When we noticed, we removed it from the about page, the terms, the home page FAQ in all ten languages and the accuracy page's meta description. This paragraph is the record of that. The ACETADA and Fridolfsson figures above are the real picture.