Decision 1 of 10

Why we think a photograph is good enough

An estimate from a photograph is a good guess, not a measurement. It is worse for some nutrients than for others.

Written and maintained by Mattias Geniar. Last updated .

What the app does
Puts a confidence badge on every meal, prints the assumption it made under each ingredient, and asks you a question instead of guessing what it cannot see.
How much we believe it
Enough to show a number, not enough to show it on its own. Portion size is where the error comes from, and a better camera does not fix that.

What the benchmarks say

The best public benchmark is ACETADA (Coburn et al., 2025). It has 806 meal images from a controlled feeding trial. Dietitians gave each item a weight in grams, checked against food weighed to 0.1 g. Against that, current vision models are off by about 166 to 211 kcal per meal on average. As a percentage, that is a mean absolute percentage error (MAPE) of 24 to 34%. For carbohydrate the error is 23 to 28 g (37 to 50%). Fridolfsson et al. (2025) tested 52 photographs against a calibrated scale and found an energy error of 35.8% for two of the three models tested.

How far off a calorie estimate from a photograph is Mean absolute percentage error. The bottom two rows measure weight rather than energy, so they are not the same claim.
Vision models, energy per meal 24–34%

ACETADA, 806 meal images against food weighed to 0.1 g

Vision models, energy per meal 35.8%

Fridolfsson et al. 2025, 52 photographs against a calibrated scale

Vision models, carbohydrate 37–50%

ACETADA, same images. Carbohydrate is harder than energy on every benchmark here

Nutritionists, weight, from the same photo 41%

Thames et al., CVPR 2021. Four professionals, ten plates

Non-nutritionists, weight 53%

The same study, sixteen amateurs

0 60% error
The human comparison from Nutrition5k is 10 simple plates shown to 16 amateurs and 4 professionals. The authors say themselves it “is not an exhaustive study”. It is on the same chart because it is the only measurement anyone has of the alternative, not because it is as good a measurement.

In the study that produced Nutrition5k (Thames et al., CVPR 2021), nutritionists estimating weight from the same photographs were 41% out on average. Non-nutritionists were 53% out. Read that carefully before relying on it: it measures weight error, not energy error, and the sample is small, as the caption above says. A photograph is not accurate. Neither is a person looking at the same photograph.

What the app does with this is on the accuracy page: the confidence badge on each meal, the assumption printed under each ingredient, and the question it asks when it cannot see the cooking fat.

What actually changes the error

Not the prompt, and not the camera. Vedovelli et al. (2026) tested 40 vision-language models. The choice of model explained 99.6% of the difference in results. Taking photos from several angles gave no benefit (p = 0.182), and changing the prompt had no significant effect (all p > 0.05).

What does help is extra information. In Rodríguez-Jiménez et al. (2025), 195 dishes, carbohydrate error fell from 11.72 g (41.95% MAPE) with the image alone to 6.20 g (23.52%) once the model was told the ingredient amounts. Mu et al. (2025) found carbohydrate error dropping from 56.6% to 20.2% when the model was given the true food weight.

Tell it what the food weighs and the error roughly halves Carbohydrate error, before and after the model is given portion information.
From the image alone 41.95%

Rodríguez-Jiménez et al. 2025, 195 dishes

With ingredient amounts supplied 23.52%

The same 195 dishes

From the image alone 56.6%

Mu et al. 2025

With the true food weight supplied 20.2%

The same meals

0 60% error
Two studies, both on carbohydrate, both showing the same thing: the hard part is judging the portion, not recognising the food. That is why the app asks you questions instead of just handing over a total, and why it writes down what it assumed.
A Snapkin screen showing a photograph of a meal with the question “was anything cooked in oil or butter?” and three answers to choose from.
The question, instead of the guess.

So the design follows the evidence. More angles do not improve the estimate in any measurable way, so the app does not ask for them. Portion and preparation are where the error is, so that is what it asks about. It asks once, and only in the two or three cases where the answer changes the number by more than the question costs you.

It is also why every ingredient carries the assumption it was estimated under. An estimate you can correct is worth more than one that is slightly better but that you cannot check.

The part we are least comfortable with

The problem is not average accuracy but variation. One preprint asked the same models about the same 13 photographs 26,904 times. Depending on the model, the answers varied by 2.4% to 11% on average (the median coefficient of variation). One plate of paella came back with anywhere between 55 g and 484 g of carbohydrate across repeats of the same image. The same preprint found that a model's own stated confidence had almost no relation to whether it was right. It is not peer reviewed. We cite it anyway because it is the only work that asks the question.

The same photograph, asked again and again Carbohydrate returned for one plate of paella, across repeats of the identical image.
One paella, same image, repeated 55–484 g

The full range returned. Median coefficient of variation across models ran from 2.4% to 11%

0 500 g of carbohydrate
26,904 queries over 13 photographs. A nine-fold spread on one dish is not something you can average away, and a model that does not know when it is wrong cannot warn you. This is a preprint and has not been peer reviewed.

A figure we published and withdrew

We used to quote “18% median, 42% at the ninetieth percentile” from our own benchmark. That number comes from a real internal test over twenty meals. But it was measured against meals a person had already corrected by hand, not against weighed food. That is fine for comparing one model with another. It is not a measurement of accuracy, and we should not have presented it as one. When we noticed, we removed it from the about page, the terms, the home page FAQ in all ten languages and the accuracy page's meta description. This paragraph is the record of that. The ACETADA and Fridolfsson figures above are the real picture.