The research behind Snapkin
Last updated 5 September 2026.
Snapkin tells you what to eat. That is a thing you should be allowed to argue with, so this page is where the arguments are written down: each big decision the app takes, what the evidence for it actually says, and what the evidence against it says. Every study is linked. Where a number is weak or a paper is small, it says so.
This page is in English only, deliberately. Ten machine translations of a clinical claim is ten ways to get one subtly wrong.
None of this is medical advice. Snapkin is a general wellness tool. It does not diagnose, treat, prevent or monitor any disease, and it is not a medical device. See the terms for what that means in practice.
1. Why we think a photograph is good enough
The honest position: an estimate from a photograph is a good guess, not a measurement, and it is worse for some nutrients than others.
The best-constructed public benchmark is ACETADA (Coburn et al., 2025): 806 meal images from a controlled-feeding crossover trial, where dietitians assigned gram-level consumed weights against food weighed to 0.1 g. Against that ground truth, current vision models land at a mean absolute error of roughly 166 to 211 kcal per meal (24 to 34% MAPE), with carbohydrate at 23 to 28 g (37 to 50%). Fridolfsson et al. (2025), 52 photographs against a calibrated scale, put energy MAPE at 35.8% for two of the three models tested.
For comparison, in the study that produced Nutrition5k, nutritionists estimating mass from the same photographs were 41% out on average, and non-nutritionists 53%. A photograph is not accurate. Neither is a person looking at the same photograph.
What actually moves the error
Not the prompt, and not the camera. Vedovelli et al. (2026) benchmarked 40 vision-language models and found model architecture accounted for 99.6% of performance variance, while multi-angle photography gave no benefit (p = 0.182) and prompt engineering showed no significant effect at all (all p > 0.05).
What does help is context. In Rodríguez-Jiménez et al. (2025), 195 dishes, carbohydrate error fell from 11.72 g (41.95% MAPE) with the image alone to 6.20 g (23.52%) once ingredient amounts were supplied. Mu et al. (2025) found carbohydrate MAPE dropping from 56.6% to 20.2% when the model was handed the true food weight. Portion estimation is the bottleneck, not food recognition, which already runs at 74 to 93%. That is why the app asks you questions back rather than just handing over a total, and why it writes down what it assumed.
The part we are least comfortable with
Variance, not average accuracy. One preprint asked the same models the same 13 photographs over 26,904 queries. Median coefficient of variation ranged from 2.4% to 11% depending on the model, and one paella came back anywhere between 55 g and 484 g of carbohydrate across repeats of the same image. It also found that a model's self-reported confidence had near-zero correlation with whether it was actually right. That is not peer reviewed, and we are citing it anyway because it is the only work that asks the question.
Where we disagree with our own marketing
We have quoted a figure of "18% median, 42% at the ninetieth percentile" from our own benchmark. That number comes from a real internal sweep, but it is measured against meals a person had already corrected by hand, over twenty meals. It is a fair way to compare one model against another. It is not a measurement of accuracy against weighed food, and we should not have presented it as one. That wording is still on the about and terms pages at the time of writing and is being fixed; treat the ACETADA and Fridolfsson figures above as the honest picture until it is.
2. Why protein comes first
Protein gets the dial you reach rather than the one you stay under. The strongest reason is body composition in a calorie deficit, and it is a smaller effect than most apps imply.
Wycherley et al. (2012), a meta-analysis of 24 randomised trials and 1,063 people on calorie-matched diets, found higher protein produced:
- body weight −0.79 kg (95% CI −1.50, −0.08)
- fat mass −0.87 kg (95% CI −1.26, −0.48)
- fat-free mass +0.43 kg preserved (95% CI 0.09, 0.78)
- resting energy expenditure +595.5 kJ/day preserved (95% CI 67.0, 1124.1)
Under one kilo of scale weight over about twelve weeks. The authors call it "modest" and they are right. The win is what the weight is made of, not how much of it there is. Their own satiety finding is weaker still: greater satiety with higher protein in only 3 of the 5 studies that measured it.
Longland et al. (2016) is the upper bound: 40 young men, four weeks, 2.4 vs 1.2 g/kg/day during a 40% deficit, with all food provided and supervised training six days a week. Lean mass +1.2 kg vs +0.1 kg, fat mass −4.8 kg vs −3.5 kg. The authors call it a proof-of-principle trial. It is not a forecast for somebody logging food on a phone.
On satiety specifically, Kohanmoo et al. (2020) pooled 49 acute and 19 long-term isocaloric trials: hunger fell about 7 mm and fullness rose about 10 mm on a 100 mm scale. Real, clearly measured, and small. Their long-term arm found no effect on any appetite outcome.
The evidence against
- At two years it makes no difference at all. POUNDS LOST (Sacks et al., NEJM 2009) randomised 811 adults to 15% vs 25% protein for two years: −3.0 kg vs −3.6 kg, p > 0.20. "Satiety, hunger, satisfaction with the diet, and attendance at group sessions were similar for all diets." What did predict weight loss was showing up: 0.2 kg per session attended.
- Adding protein does not reduce what you eat. Ben-Harchache et al. (2021), 22 studies and 857 older adults: a protein preload cut the following meal by 164 kJ, but once the supplement's own energy was counted, total intake was higher by 649 kJ (95% CI 438, 861). Protein works as a substitution at matched calories. As an addition, it does not.
- The lever may be the wrong way round. Gosby et al. (2011) found that dropping protein from 15% to 10% raised energy intake by 12%, while raising it from 15% to 25% did not change intake at all. That argues the goal is avoiding protein dilution, not maximising protein. Martens et al. (2013) found the opposite. The two disagree, and that disagreement is the actual state of the evidence.
One more caveat on our own numbers. The "about 30 g of protein per meal" rule in the recipe library comes from Schoenfeld & Aragon (2018), which recommends 0.4 g/kg/meal across at least four meals. That is a narrative review about muscle growth in resistance-trained people, not a meta-analysis and not about satiety or weight loss, and its lead author discloses a scientific advisory role with a protein supplement company. We use it because a defensible floor beats no floor, not because it is settled.
3. Why we track added sugar, and what we mean by it
The naming, first, because it matters
Snapkin says "added sugar" everywhere, because that is the phrase people use. The definition underneath is the World Health Organization's free sugars: sugars added by a manufacturer, a cook or the person eating, plus honey, syrups, and fruit juice and its concentrates. Not the sugar in whole fruit, in vegetables, or in plain milk and yoghurt.
Those two are not the same thing, and they differ in exactly one place. A US nutrition label's "Added Sugars" line (21 CFR 101.9) excludes 100% fruit juice. WHO counts it in full. So a glass of orange juice is 0 g of "added sugar" on an American label and roughly 21 g in Snapkin. We picked the wider definition on purpose, because juice is exactly the case people mean when they say they are cutting sugar.
Total sugars, the figure a European label actually declares, is the wrong number for this job entirely: it counts the lactose in yoghurt and the fructose in blueberries, and a limit that goes amber for a bowl of berries is a limit nobody keeps reading. The WHO guideline recommends free sugars below 10% of energy, with a conditional further reduction below 5%. Snapkin sets the daily ceiling at 10% of your calorie target.
How the number is worked out
Free sugars are not printed on any European label, so they have to be inferred. There is a published protocol for exactly this and we implement it rather than inventing one: Louie et al. (2015) laid out ten steps running from objective evidence down to category guesswork; Kibblewhite et al. (2017) restated them for WHO free sugars; Scapin et al. (2021) adapted them for a country whose labels declare neither total nor added sugars.
Scapin's validation against FDA-declared values on 930 products reached an ICC of 0.98, and an independent benchmark of Louie's method (Davies et al., 2022) put it at R² 0.97 with a mean absolute error of 1.26 g/100 g. Both of those are one specialist working through packets with the ingredient list in hand. The number that says what the protocol is really worth is the inter-rater one: two trained researchers applying the same written steps to the same 5,740 foods differ with a standard deviation of 4.6 g/100 g (Louie, Lei & Rangan, 2016).
The evidence against, and it is substantial
- Sugar does almost nothing at matched calories. Te Morenga et al. (BMJ 2013): reducing dietary sugar changes weight by −0.80 kg (95% CI −0.39 to −1.21) and increasing it by +0.75 kg (0.30 to 1.19), but there is no evidence of any effect when energy intake is held equal. Cutting sugar works by removing easy calories and by being a rule you can follow, not as a metabolic lever. We present the dial as a heuristic for that reason.
- Estimating sugar from a photograph is the weakest thing this app does. In the entire literature there is one peer-reviewed study reporting a sugar figure from food photographs: O'Hara et al. (2025), 114 photographs of weighed meals, where sugar came out 32% low against energy at +0.1% and protein at −2.7%. Nobody anywhere has separated added from total sugar in a photograph. Other groups decline to analyse sugar at all and say why: it hides in sauces and processed components and cannot be visually separated.
- "Carbs that turn into sugar" is a different thing and we do not track it. That is glycaemic load, not sugar, and Zeevi et al. (Cell 2015) put 800 people on continuous glucose monitors across 46,898 meals and found high variability between people responding to identical meals. There is no per-food number to put in a table, so we do not pretend to have one.
There is one encouraging detail in the O'Hara result. Sugar was the worst nutrient by mass and among the best by rank correlation (0.75 by mass, 0.84 as a share of energy). The model knows which meals are sugary; it lowballs how much. That is why the dial is a rough budget and never a precise reading, and why it carries no confidence badge.
4. Why track anything at all
The strongest evidence for logging is indirect, and worth stating plainly: no trial has ever randomised people to log or not log inside an otherwise identical programme. What exists compares logging methods, or compares whole programmes that happen to contain logging.
The best of it: Berry et al. (2021), a meta-analysis of 12 randomised trials of digital self-monitoring, found −2.87 kg (95% CI −3.78, −1.96) and energy intake down 182 kcal/day. Michie et al. (2009), a meta-regression over 122 evaluations and 44,747 people, found self-monitoring explained more between-study variation than any of the other 25 behaviour-change techniques examined. Harvey et al. (2019) found that people losing 10% or more logged 2.7 times a day against 1.7 for those losing under 10%, while the time spent logging did not discriminate at all.
And the classic caveat from the same literature: Burke et al. (2011) found the association consistently, then said out loud that "the level of evidence was weak because of methodologic limitations."
Prediabetes is where the evidence is strongest, and it is not about the app
The Diabetes Prevention Program (n = 3,234, mean follow-up 2.8 years) is the landmark. An intensive lifestyle programme targeting 7% weight loss and 150 minutes of activity a week cut the incidence of type 2 diabetes by 58% (95% CI 48–66) against placebo, beating metformin's 31% (17–43). Mean weight loss was 5.6 kg. The Finnish Diabetes Prevention Study (n = 522) found the same 58% reduction independently.
Dietary self-monitoring was part of both. In the DPP, Wing et al. (2004) reported that "dietary self-monitoring was positively related to meeting both weight loss and activity goals." In the Finnish study, dietary advice was built on three-day food records.
On how much weight loss is enough: the round "5 to 7%" figure is a programme target back-derived from what those two trials set and hit, not the output of a dose-finding study. The real relationship is continuous, with no cliff. Hamman et al. (2006) found a 16% reduction in diabetes risk for every kilogram lost (HR 0.42, 95% CI 0.35–0.51 per 5 kg), and weight loss was the dominant predictor. Magkos et al. (2016) showed 5% weight loss alone improving insulin sensitivity in adipose, liver and muscle.
Be clear about what this buys. At the 15-year follow-up, cumulative incidence was 55% in the lifestyle arm, 56% for metformin and 62% for placebo. The effect is largely delay, not prevention. More than half of every arm developed diabetes.
The evidence against, and the two numbers we like least
- Handing somebody an app does approximately nothing. Laing et al. (2014) randomised 212 primary-care patients to usual care with or without MyFitnessPal for six months. Weight difference: −0.30 kg (95% CI −1.50 to 0.95), p = 0.63. Their conclusion: apps "may be useful for persons who are ready to self-monitor calories," but introducing one "is unlikely to produce substantial weight change for most patients." Chew et al. (2022) put the pooled app effect at 2.18 kg at three months tapering to 1.63 kg at twelve, and called the effects "minimal in their current states." Adding a human coach roughly doubles it (Chew et al., 2023).
- Almost nobody keeps logging. Helander et al. (2014) followed 189,770 downloads of a free photo food-logging app. 2.58% used it actively, and the authors noted those who did "may already have been healthy eaters." That is the number this app is up against, and it is the reason Snapkin asks for a photograph rather than a database search.
-
Self-reported calories barely track real calories.
Freedman et al. (2014) pooled five
validation studies against doubly labelled water (n = 2,265): under-reporting averaged
15% on a single 24-hour recall and 28% on a food frequency questionnaire,
and the correlation between reported and true energy intake was
r = 0.21 to 0.31. Under-reporting also gets worse with body
weight, by a further 5 to 7 percentage points at BMI 30 versus 25. Photo-based methods
are not exempt: Ho et al.
(2020) found image-based logging 448 kcal/day short against doubly labelled
water.
The gap is mostly not the estimate. It is the meals nobody logs.
5. Tracking and eating disorders
We think anybody building a calorie tracker owes a clear answer here, so this section exists even though it argues against our own product.
The association is real in cross-sectional data. Levinson et al. (2017) surveyed 105 adults recently discharged from eating-disorder treatment: 74% had used MyFitnessPal, and of those, 73% said it contributed at least somewhat to their eating disorder, 30% said "very much." The authors are explicit that the data are retrospective and cannot establish cause.
Every study that has actually manipulated tracking has come out null. Hahn et al. (2021) randomised 200 undergraduate women to a month of MyFitnessPal or nothing: no effect on eating-disorder symptoms (β = −0.04, p = 0.17), though the sample was screened to be low-risk. Jospe et al. (2018) ran 250 adults for twelve months and found no differences (all p ≥ 0.164) in a group where 53.6% had binge eating at baseline. The pooled estimate across nine studies and 13,507 people is r = 0.13 (95% CI −0.02 to 0.28), p = 0.08, not significant (Roth et al., 2024).
The best summary of the state of play is Moody et al. (2025): cross-sectional studies show a reasonably consistent association, "however, this association was not replicated in experimental research," and it is "currently not possible to conclude if they increase disordered eating, or the direction of this relationship." Longitudinal work runs both ways: tracking predicts later disordered behaviour in one cohort, and compensatory behaviour predicts later tracking in another (O'Loughlin et al., 2023, adjusted OR 1.7).
What does reliably separate harm from no harm is why somebody is tracking. Two independent studies found that tracking for weight and shape reasons is associated with harm while tracking for health and fitness reasons is not.
No guideline we could find prohibits calorie tracking. The pattern is screen and tailor, not avoid. NICE NG69 actively prescribes self-monitoring of dietary intake inside eating-disorder treatment. The 2023 AAP guideline notes that eating disorders and other conditions "may preclude the use of certain tools that require intensive tracking." The clearest statement about apps like ours is NICE HTG700, which requires commercial weight-management services to have access to psychological support "to reduce the risk of harm, including from disordered eating."
Snapkin has no such support, and we are not going to pretend otherwise. If you have or have had an eating disorder, please talk to a clinician before using this or any tracker. Concretely, the app does not show a streak on the sugar limit, does not celebrate going under a ceiling, and shows what is left rather than counting upward, because a restriction metric with a streak on it is the shape that goes worst for the people most likely to be counting.
How this page is maintained
Every figure here was read from the paper or its abstract, not recalled. Where we could only reach an abstract, the page says what the abstract says and no more. Where a claim rests on a preprint or on a single small study, it is labelled. If you find something wrong, tell us at [email protected] and we will correct it here.