Evaluation Rules
A plain-language reference to every score, badge and watch-out you see across sativum.ai — so you can judge for yourself how a product was rated. This page is maintained by Manole 13 Labs (sativum.ai) and reflects the rules currently in production.
Updated · July 2026What just changed
We introduced a third orthogonal score, tightened salt / sugar rules, and made the scoring pipeline more transparent. The headline changes:
- New Authenticity Score (0–100) — for products matched to a canonical traditional dish (Ragù alla Bolognese, Hummus, Pesto Genovese…), measures how faithful the industrial version is to the original recipe. Surfaced as a tile in the KPI strip and folded into the Sativum star rating as a 7th KPI when present.
- New Culinary Score (0–100) — measures sensory density and cohort standing, independent of Clean and Health. An aged cheese or curry paste is no longer punished for being what it is. Feeds the new buy-reason chips (Great taste, Clean label, Everyday value, Indulgent, Skip). The breakdown now also exposes the peer count used for cohort standing, so "not enough peers yet" is auditable.
- Salt as #1 ingredient → Clean Score capped at 55 (45 in seasonings), even when the nutrition panel is present. The previous loophole that skipped the position check when salt grams were declared is closed.
- Quantitative sugar caps now apply to every category: ≥ 12.5 g/100 g caps Clean at 65, ≥ 22.5 g/100 g caps at 50.
- Added-sugar detection tightened — the sugar cap now fires only on explicit added / free sugars (sugar, glucose, fructose, HFCS, syrups, honey, malt, molasses, treacle, dextrose). Whole fruits, fruit pieces, cocoa, coconut, milk powder and polyols are no longer treated as added sugar — their naturally occurring sugars appear as informational notes only. This ends a class of false positives on fruit-forward and dairy-containing products.
- Health Score hard ceiling of +1 ("Good") when salt or free sugar is the first pack ingredient, salt ≥ 10 g/100 g or sugar ≥ 22.5 g/100 g — such products can no longer reach "Very Good" or "Excellent".
- Top-3 amplification (1.5×) on harmful ingredients in pack positions 1–3 for the Health Score, so position matters as much as presence.
- Clean Score audit trail — Stage-1 label-vision deductions (each with a reason, a −5 / −10 delta and the verbatim label token that grounds it) are surfaced in the score explainer. Legacy scans without a stored trail are labelled as such rather than silently reconstructed.
- Organic / bio detection now runs on every scan. Three trust tiers (Certified · Likely · Claimed) are surfaced as a leaf badge on tiles, hero and compare views, and add up to +10 Clean Score (after caps, hard-capped at 100). "Claimed" products carrying synthetic additives are flagged as greenwashing suspect.
Basics
sativum.ai helps you read food labels with more context: each product gets a Clean Score, ingredients get a Health Score, and the overall product receives one of three verdicts — Recommended, Situational or Not recommended. The rules below are deterministic and apply identically to every product, regardless of brand.
This is an informational guide, not medical, regulatory or nutritional advice. See our Privacy Policy and Terms of Service for the full disclaimer.
Clean Score (0–100)
The Clean Score measures label transparency and processing load — not calories, not macros. A high score means short, recognisable ingredient lists with few additives; a low score means the opposite.
67–100
34–66
0–33
What raises the score
- Short ingredient list with whole-food names
- No flavour enhancers, colourings or sweeteners
- Named fats (e.g. "olive oil") instead of "vegetable fat"
- Transparent allergen and origin declarations
What lowers the score
- High-concern additives with regulatory/health evidence (Southampton colours, nitrites, sulphites, BHA/BHT, glutamates, aspartame)
- Artificial sweeteners or artificial colours (any presence)
- Three or more moderate ultra-processing additives stacked
- Palm oil / palm fat; vague terms like "natural flavouring", "spices"
How the Clean Score is computed
The score starts from an AI label-reading estimate (0–100) and is then passed through a deterministic cap-and-stack pipeline. Every additive detected is looked up in an evidence-based registry that assigns a function class (sweetener, colour, preservative, flavour enhancer, emulsifier…) and a risk class (high / moderate / low / neutral) with a citation (EFSA re-evaluation, IARC monograph, Southampton study, JECFA opinion). We flag by function and evidence — not by raw presence: neutral natural colours (riboflavin, carotenes, curcumin, annatto, beetroot red) and low-risk technical aids (citric/ascorbic acid, lecithin, gums, pectin) never trigger a cap.
- AI score — initial 0–100 read from the ingredient list.
- Caps fire — each red-flag trigger imposes a maximum allowed score (see table below).
- Top-3 amplification — if a trigger ingredient sits in pack positions 1–3 (i.e. it is one of the main ingredients by mass), its cap tightens by an extra −10 points.
- Stack penalty — every additional cap beyond the first subtracts another −5 points (max −20).
- Final score = min(AI score, tightest cap) + stack penalty, clamped to 0–100.
Every community product exposes a collapsible “Why this score?” panel that shows the AI score, every triggered cap (with a top-3 badge when amplified), the stack penalty, and the final number — so the calculation is fully auditable.
Caps in force
| Trigger | Cap | If in top-3 |
|---|---|---|
| High-concern additive (Southampton colours E102/110/122/124/129, nitrites E249/250, sulphites E220–228, BHA/BHT E320/321, glutamates E620–635, aspartame E951, caramel-IV E150d) | 45 | 35 |
| ≥2 high-concern additives stacked | 35 | 25 |
| Artificial sweetener (sucralose E955, acesulfame E950, saccharin E954, cyclamate E952, neotame E961) | 55 | 45 |
| Intransparent term (“spices”, “natural flavouring”, “aroma”) | 60 | 50 |
| Palm oil / palm fat / palm kernel | 60 | 50 |
| ≥3 moderate ultra-processing additives (emulsifiers, phosphates, carrageenan, artificial flavourings, modified starch, maltodextrin, polyols, HVP, yeast extract) | 55 | 45 |
| 1–2 moderate additives | 70 | — |
| Low-risk additives only (citric/ascorbic acid, lecithin, gums, pectin) — informational, no cap | — | — |
| Neutral natural colours/vitamins (riboflavin, carotenes, curcumin, annatto, beetroot red) — not counted | — | — |
| Added sugar (any subcategory) | 70 | — |
| Added sugar in a seasoning/condiment | 60 | — |
| Sugar ≥ 12.5 / 22.5 g per 100 g | 65 / 50 | — |
| Sugar / syrup in pack position 1 / 2-3 (no panel) | 55 / 65 | — |
| Salt ≥ 5 g / 10 g / 20 g per 100 g | 60 / 45 / 30 | — |
| Salt in pack position 1 / 2-3 (always evaluated) | 55 / 65 (45 / 60 in seasonings) | — |
| ≥2 processed-filler patterns (maltodextrin, dried potato, milk powder…) | 60 | 50 |
Footnote — what does “cap” mean? A cap is a ceiling, not a subtraction. When a trigger fires, the product's score cannot exceed that number, regardless of what the AI initially gave it. Example: AI scores a product 78, but a flavour enhancer is detected → the score is pulled down to 55 (the cap). If multiple triggers fire, the lowest cap wins, and each additional trigger removes another −5 points via the stack penalty (max −20). The top-3 badge means the offending ingredient sits in pack positions 1–3 (a main component by mass), which tightens the cap by another −10.
Worked example — salt-first seasoning
Label: Salt, paprika, garlic, onion, parsley. Salt 62 g / 100 g.
AI score: 72.
Caps fired:
- Salt is #1 ingredient in a seasoning → cap 45
- Salt ≥ 20 g/100 g → cap 30
Tightest cap wins: 30. Stack penalty (1 extra cap): −5. Final Clean Score = 25.
Health Score (−10 … +10)
Each ingredient has a Health Score on a −10 to +10 scale, aggregated from peer-reviewed beneficial and harmful health effects weighted by intensity. Positive = net beneficial effects dominate; negative = net harmful effects dominate.
- Beneficial effects (antioxidant, anti-inflammatory, cardio-protective…) add positive points.
- Harmful effects (pro-inflammatory, allergenic, toxicity at dose…) subtract points.
- Each effect is weighted by its documented intensity in our database.
Product-level Health Score (community products)
For a finished product (a blend of many ingredients), single-ingredient scores are combined with a position-weighted formula — ingredients listed first count for much more than ingredients listed last, because EU labelling rules require descending order by mass.
- Harmonic position weight: ingredient at index i gets weight
wi = 1 / (i + 1). So position 1 = 1.00, position 2 = 0.50, position 3 = 0.33, position 10 = 0.09. - Bad-actor amplification: harmful ingredients (negative score) in positions 1–3 are multiplied by 1.5× — a flavour enhancer as the second ingredient hurts the product score much more than the same enhancer at the end.
- Stacked product penalties add on top of the weighted average: flavour enhancers, palm oil, ≥3 additives, high salt and high added sugar each subtract additional fixed points from the product Health Score.
- Final score = weighted average of ingredient scores + product penalties, clamped to −10 … +10. A hard ceiling of +1 applies whenever a severe lever fires: salt ≥ 10 g/100 g, sugar ≥ 22.5 g/100 g, or salt / free sugar as the first pack ingredient. Such products can never display "Very Good" or "Excellent".
Each product page exposes a collapsible “Why this score?” panel listing every ingredient with its position weight, its individual contribution and every penalty applied, so the result is fully traceable.
The Health Score reflects general evidence at typical culinary doses. It is not tailored to individual conditions, allergies or medications — always consult a qualified professional for personal advice.
Worked example — sugar-heavy syrup
Label: Glucose-fructose syrup, water, lemon juice, citric acid, natural flavouring. Sugar 58 g / 100 g.
Weighted ingredient average: −1.2.
Amplification: syrup at position 1 (harmful, top-3) → ×1.5 applied to its negative contribution.
Product penalties: added sugar (−2), sugar ≥ 22.5 g/100 g (−2).
Raw result: ≈ −5. Hard ceiling +1 doesn't bind (already below). Final Health Score = −5 ("Poor"). For a milder product the hard ceiling would clamp anything > +1 down to +1.
Culinary Score (0–100)
The Culinary Score is orthogonal to Clean and Health. It answers a different question: is this a good example of what it claims to be? Rewards sensory density and cohort standing so an aged cheese, a fish sauce or a curry paste is not punished for being intense — those products have to be intense to work.
80–100 · Exceptional
65–79 · Great
45–64 · Solid
0–44 · Mild / Bland
How it's computed
- Richness (0–60 pts) — sum of the top intensities across taste, aroma, mouth-stimulus and texture. Top-3 tastes are weighted ×1.5 because headline flavors define character. Normalized against a reference ceiling of 80 raw points.
- Cohort standing (0–40 pts) — percentile of this product's richness inside its subcategory. When cohort data is missing the component is skipped (the score then maxes out at 60).
- Final = richness + cohort, clamped 0–100. Label from Exceptional / Great / Solid / Mild / Bland.
Worked example — aged hard cheese
Richness raw ≈ 72 (deep umami, high aroma density, firm crystalline texture) → 72 / 80 × 60 = 54 pts.
Cohort percentile = 0.9 (top decile of hard cheeses) → 36 pts.
Final Culinary Score = 90 · "Exceptional". Clean and Health scores are computed separately and may be lower — that is expected and correct.
A high Culinary Score does not imply clean or healthy. Always read the three scores together — that is the whole point of keeping them independent.
Authenticity Score (0–100)
The Authenticity Score answers a specific question: how faithful is this industrial product to the canonical traditional recipe of the dish it evokes? It only appears when the product has been matched to a canonical dish (Ragù alla Bolognese, Hummus, Pesto Genovese, Guacamole, Tzatziki…). Products with no traditional counterpart don't get this KPI — the tile and star weight are dropped, not penalized.
85–100 · Traditional
70–84 · Faithful
50–69 · Adapted
30–49 · Industrial
0–29 · Far from original
How it's computed
Score = 60 · coreCoverage + 25 · optionalCoverage + 15 − deviationPenalty, clamped 0–100. Higher-tier canonical sources (PDO / PGI legal text, producer-of-origin) win over consensus writeups when the canonical recipe is researched.
- Core coverage (0–60 pts) — weighted fraction of the dish's core ingredients that are present on the label. Missing a core ingredient (e.g. tahini in hummus, basil in pesto) is the biggest hit.
- Traditional optional (0–25 pts) — fraction of the traditional but non-mandatory ingredients present. Missing these is fine, having them helps.
- Baseline (+15 pts) — everyone starts here; the penalties below eat into it.
- Deviation penalties — capped per category so no single deviation zeroes the score:
- Additives / E-numbers / flavourings — up to −30.
- Added sugars where the traditional recipe has none — up to −20.
- Wrong oil / fat basis (e.g. sunflower oil instead of olive) — up to −10.
- Dominant swap: non-canonical ingredient in pack positions 1–3 — up to −15.
Worked example — supermarket "hummus"
Canonical core: chickpeas, tahini, lemon juice, garlic, olive oil, salt.
Label: chickpeas, sunflower oil, water, lemon juice, garlic, salt, preservative (E202), acidity regulator.
Core coverage = 4/6 (missing tahini + olive oil) → 60 × 0.67 = 40 pts.
Deviations: sunflower oil in position 2 (wrong oil −10, dominant swap −15), E202 (additive −5), acidity regulator (additive −5) → −35 pts.
Raw = 40 + 15 − 35 = 20 · "Far from original".
Authenticity is orthogonal to Clean, Health and Culinary — a heavily adapted product can still be clean, or a traditional one can still be salt-heavy. It is folded into the Sativum star rating as an equal-weight 7th KPI whenever a canonical match exists.
Organic / Bio detection (0 / +5 / +10 Clean bonus)
Every scan is checked for organic evidence on the label — logo text, ingredient list, brand and declared claims — and classified into one of three trust tiers. A leaf badge is surfaced on the community grid tile, the product hero, the compare bar and the PDF export. Certified and Likely tiers add a bonus to the Clean Score (applied after all caps and hard-capped at 100).
Certified · +10
DE-ÖKO-006, FR-BIO-01). Highest trust.Likely · +5
Claimed · +0
Detection sources
- EU Bio-Logo + control code matched against the
organic_control_bodiestable — code and certifier name are stored on the observation and rolled up to the product. - Certifier keywords — a curated allow-list of trusted labels (see tier 2 above).
- Loose marketing keywords — multilingual (EN/DE/FR/IT/ES/RO/RU) with a small denylist to avoid false positives like "eco-friendly packaging" or "biological process".
Rollup & bonus
- Detection runs per observation — the highest tier across all scans of a product wins and is written to
aggregate_stats.organic. - The bonus (+10 / +5 / 0) is added after the Clean Score cap-and-stack pipeline, then clamped to 100 so it can never override a hard cap into an impossible number.
- Fake-claim / greenwashing suspects keep the +0 bonus and surface an amber pill so the user sees the mismatch immediately.
Buy-reason chips
The colored chips on each product page are deterministic. Given the three scores plus additive count, the same product always produces the same chips. Rules are evaluated in the priority order below and a product shows at most two chips.
| Chip | Fires when |
|---|---|
Great taste | Culinary Score ≥ 75 |
Clean label | Clean Score ≥ 75 and additive count ≤ 1 |
Everyday value | Clean Score ≥ 65 and Health Score ≥ +2 |
Indulgent | Culinary ≥ 70 and (Clean < 55 or Health < 0) — strong sensory value, worth it as a treat. |
Skip | Culinary < 40 and Clean < 40 — few reasons to buy. |
Verdict labels
Every community product receives one of three verdicts. The rule is deterministic: we compute it first, then the AI-written narrative is required to be consistent with it.
Clean and unproblematic
Fine occasionally, not a daily staple
Clear red flags on the label
Strengths & watch-outs (deterministic rules)
The bullet-point Strengths and Watch-outs shown on every community product page (and in the Compare view) are not free-form AI text. They come from a deterministic rule registry evaluated against the parsed label: if the signal is true, the callout fires — every time, on every product. The AI's own writeup is only kept when it doesn't duplicate a rule, which is why two products with the same label always surface the same positives and the same negatives.
Strengths (max 6)
| Rule | Fires when |
|---|---|
| No added MSG / flavour enhancers | No E620–E635, MSG, HVP, yeast extract or "flavour enhancer" found |
| No added sugar | No sugar, glucose, fructose, syrups, honey, agave found |
| No palm oil | No palm oil / palm fat / palm kernel found |
| No artificial colours | No E1xx azo dyes, tartrazine, sunset yellow etc. found |
| No artificial preservatives | No E2xx, sulphites, benzoates, nitrites, BHA/BHT found |
| No hydrogenated / trans fats | No "hydrogenated", "trans-fat" or "shortening" found |
| Transparent ingredient list | No vague terms ("spices", "aroma", "natural flavouring") |
| Short ingredient list ({n} items) | ≤ 8 ingredients |
| No / only one additive E-number | additive_count ≤ 1 |
| No declared allergens | allergen_count = 0 |
| Good fibre content | fibre ≥ 6 g / 100 g |
| Low salt | salt ≤ 0.3 g / 100 g (sodium converted if salt missing) |
Watch-outs (max 6)
| Rule | Fires when |
|---|---|
| Extremely high salt ({x} g / 100 g) | salt ≥ 5 g / 100 g |
| High salt ({x} g / 100 g) | 1.5 ≤ salt < 5 g / 100 g |
| Salt is the first / main ingredient | First ingredient is salt / sea salt / sodium chloride |
| Added sugar in a seasoning / condiment | Added sugar present and subcategory ∈ {seasonings, condiments, rubs, stock-cubes, marinades} |
| Contains added sugar | Explicit added / free-sugar token detected (sugar, glucose, fructose, HFCS, syrups, honey, malt, molasses, treacle, dextrose). Whole fruit, cocoa, coconut, milk powder and polyols are excluded. |
| Contains palm oil | Palm oil / palm fat / palm kernel detected |
| Contains MSG / flavour enhancer | MSG, glutamate, HVP, yeast extract, "flavour enhancer", E620–E635 |
| Stacked ultra-processing additives ({n} moderate-risk) | ≥ 3 moderate-risk additives from the registry (emulsifiers, phosphates, carrageenan, artificial flavourings, modified starch, maltodextrin, polyols, HVP, yeast extract) |
| Contains artificial colours | E1xx azo dye or named colour detected |
| Contains artificial preservatives | E2xx, sulphites, benzoates, nitrites, BHA/BHT detected |
| Contains hydrogenated / trans fats | "Hydrogenated", "trans-fat" or "shortening" detected |
| Intransparent terms | "Spices", "aroma" or "natural flavouring" present |
| Multiple processed fillers | ≥ 2 fillers (maltodextrin, dried potato, milk powder…) |
| Broad "may contain" warning | ≥ 3 major allergen tokens inside a "may contain" / "traces of" clause |
The strengths and watch-outs for the same axis are mutually exclusive — a product either gets No palm oil or Contains palm oil, never both. Salt tiers are also mutually exclusive: the highest applicable tier wins. Cap reasons from the Clean Score (see above) are merged into the watch-out list so the UI always explains why the score was capped.
Watch-out triggers (Clean Score impact)
These triggers automatically cap or lower a product's verdict. They are not moral judgements — they exist because each one materially changes what the product is.
| Trigger | Why it matters |
|---|---|
| Flavour enhancers (E620–E650) | Mask low-quality base ingredients; force a "Not recommended" verdict. |
| Multiple additives | A stacked additive profile is a strong processing marker. |
| Palm oil / palm fat | Sustainability and fatty-acid profile concerns; caps verdict to Situational at best. |
| Intransparent ingredient list | Vague terms ("natural flavouring", "spices") block a clean verdict. |
| Allergen ambiguity | "May contain" without clear sourcing raises consumer risk. |
Adaptive weighting by product type
The Sativum star rating blends up to seven KPIs (Health, Clean, Nutrition, Culinary, Allergens, Daily-use, Authenticity). Rating a bottle of wine, a bottle of olive oil and a jar of tomato sauce with the same equal weights is misleading — some KPIs are structurally not meaningful for certain product types (e.g. per-100 g Nutri-Grade for pure oils always resolves to E).
We therefore first infer an evaluation archetype from the subcategory (with light heuristics like ABV %, "beer" / "cider" mentions, alcoholic-beverage keywords) and apply a per-archetype weight profile. Suppressed KPIs are dropped from the weighted mean and shown as informational ("not weighted") rows in the explainer so nothing is hidden.
| Archetype | Health | Clean | Nutri-Grade | Culinary | Allergens | Daily-use | Authenticity |
|---|---|---|---|---|---|---|---|
| Wine / spirits | 15 | 20 | — | 25 | 10 | — | 30 |
| Beer / cider | 20 | 20 | — | 20 | 10 | 10 | 20 |
| Oils, fats & vinegars | 15 | 25 | — | 20 | 5 | 15 | 20 |
| Cheese & aged dairy | 15 | 20 | 10 | 20 | 15 | 5 | 15 |
| Herbs, spices & seasonings | 10 | 30 | — | 25 | 15 | — | 20 |
| Sauces & condiments | 20 | 25 | 10 | 20 | 10 | 5 | 10 |
| Supplements | 35 | 30 | — | — | 15 | 20 | — |
| Infant & baby food | 30 | 35 | 15 | 5 | 15 | — | — |
| Confectionery & desserts | 25 | 25 | 20 | 15 | 10 | 5 | — |
| Everyday food (default) | Equal weighting across all present KPIs (pro-rata renormalized). | ||||||
Numbers are nominal weights — renormalized to 100 % across the KPIs that are actually present for a given product. "—" means the KPI is suppressed for that archetype and does not contribute to the weighted mean; it still appears in the explainer as informational.
Worked example — table wine (12.5 % ABV). Health, Clean, Culinary, Allergens and Authenticity contribute; Nutri-Grade and Daily-use are suppressed. If the wine scores +4 Health, 70 Clean, 78 Culinary, "Free" Allergens (100), and "Faithful" 82 Authenticity, the weighted index is (60·15 + 70·20 + 78·25 + 100·10 + 82·30) / 100 ≈ 78 → 3.9★. The same wine under equal weighting would drag in a "Grade E" nutrition tile (20/100) and a null Daily-use, misrepresenting the product.
Sensoric evaluation
The qualitative panel on each product page — Sensory character, Culinary use, Best for, Skip if — is generated in two steps:
- The deterministic verdict above is computed first.
- An AI narrative is generated and required to stay consistent with the verdict — it can never contradict it.
This keeps the prose readable while guaranteeing the badge and the text always tell the same story.
Data sources & method
- Community contributions — label scans and product entries submitted by users.
- Ingredient database — sativum.ai's curated chemistry, sensory and health-effect dataset.
- Multi-model AI pipeline — Perplexity, OpenAI, x.AI and Gemini are used in sequence for descriptive copy, with deterministic logic on top for any scoring decision.
- Scores and verdicts recompute as the underlying ingredient data is refined.
Cuisine data & auto-linking
- Source — a strict-JSON LLM pass produces each cuisine's description, cultural notes, dietary traditions, six flavor metrics (spice/heat/umami/fermentation/acidity/sweetness), a signature-ingredient list and a typical-ingredient list.
- Ingredient matching — every AI-suggested ingredient is normalized (diacritics stripped, form descriptors like "ground/dried/fresh" removed) and matched against the canonical
ingredientstable and itsingredient_name_translations. - Plausibility guard — cuisines that resolve fewer than 5 ingredients are logged to
data_quality_flagsfor human review instead of being silently saved. - Auto-relink — a trigger on the
ingredientstable enqueues every new or renamed ingredient into a relink queue; a scheduled job replays the queue and adds missing links to the cuisines whose stored payload references that ingredient. New ingredients propagate to cuisine coverage without manual work. - Refresh — admins can re-enrich a single cuisine or bulk-backfill the whole taxonomy from Admin → Cuisines. Runs are timestamped in
last_enrichment_at.
Limitations & feedback
sativum.ai is not a certification body. The verdicts on this site are not regulatory approvals, not medical advice, and not a substitute for reading the label yourself. Products evolve; recipes change; our database is improving daily.
If you spot a mistake or want to challenge a verdict, please write to hello@sativum.ai with a photo of the current label — we re-evaluate on request.
Last reviewed: August 2026. Maintained by Manole 13 Labs SRL, Curtea de Arges, Romania.
