Skip to main content
R&DishBack to workspace

Data Sources

This page publishes the identity, license, attribution wording, and snapshot of every asset R&Dish relies on for evidence. The license links on the evidence cards point to the matching section on this page.

Evidence has weight. Approved literature and government guidance carry the most; structured data issued by governments comes next; and the internally derived tools for identification, observation, and retrieval carry the least. This page sets out, in order, the assets that build that hierarchy and shows which badge each asset’s facts wear on screen. A badge’s fill density (solid, solid-line tint, dashed tint, outline) draws the weight of trust directly.

The five epistemic-type badges

Each badge’s fill grows lighter as you move down, meaning the trust grows lighter.

Source factA sentence from approved literature or government guidance. Shown only when it matches, byte for byte, a bundle registered in the corpus. Solid fill: the only one, and the heaviest.
Data figureA fact returned by an external API. Government-issued data is also accepted as a safety premise; crowd data is credited only as far as the fact of its existence and labeling (marked with a “Crowd data” supplementary tag on the card).
Your recordAn experiment result the user logged. The sentence the server normalizes becomes evidence for the next proposal: the only badge that wears the chef’s own color (terracotta).
Data observationAn observation surfaced by a model or data. Because it is not a causal fact, a proposal built from this badge alone is marked down as an “exploratory direction.” Dashed border: it signifies the tentativeness of observation.
DeductionThe conclusion of a registered inference rule. Shown only when all its premises hold. Outline only: no background, and the lightest of all.
T1a: Internal primary, human-approved

The heaviest tier: human-approved original texts

This tier stands alone. When a query resolves here, no other tier is consulted.

Approved evidence corpus

Source fact

A snapshot bundle of human-approved literature and government guidance

The only place the claim sentences of an experiment proposal come from. No other asset wears the “Source fact” badge. It is also the source that establishes safety-related premises (temperature, preservation, allergens, and so on).

LicenseMixed, per source (recorded individually in the approval manifest)
AttributionNo single attribution statement: each source’s origin and license are recorded individually in the evidence card’s source row (e.g., FDA Food Code 2022 §3-501).
OriginNo external public URL: this bundle is managed by the internal approval pipeline.
Snapshot / versionAPPROVED_CORPUS_MANIFEST v3, 13 literature & government-guidance items, latest snapshot 2026-07-18
T1b: Internal primary, authoritative structured

Structured data issued by governments

Supplements figure and composition queries left unresolved at T1a.

USDA FoodData Central

Data figure

The nutrition and composition database issued by the U.S. Department of Agriculture (USDA): Branded / Foundation / SR Legacy / FNDDS

Used for queries about the nutritional composition and figures of an ingredient. Because it is data the government issues directly, it carries the highest trust among structured data and is also accepted as a safety-related premise.

LicenseCC0 1.0
AttributionCC0 carries no legal attribution obligation. To make the origin clear, we show the wording below on the evidence cards and on this page (citation recommended, not required).
Data: USDA FoodData Central, CC0 1.0
Snapshot / versionSnapshot 2026-05-14, FDC ID 173430 et al.
T2: Internal derived

The lightest tier: tools for identification, observation, and retrieval

The three assets in this tier supply no facts. They find candidates (e5), identify an ingredient by canonical ID (FoodOn), or observe relationships within the data (epicure). When any of their results reaches the user as a sentence, it must be a normalized string the server assembled, and even then it is shown only with the tentative weight of a “Data observation.” If a query is still unresolved after T2, it is settled as unresolved: there is no next tier.

FoodOn

Data observation

Food ontology: identification asset (a canonical ID system for ingredient and food entities)

The identification layer that normalizes an ingredient name the user gives (e.g., “fresh cream,” “whipping cream”) to a canonical ID. It produces no claim sentence on its own and is shown as a “Data observation” only when an entity-resolution result surfaces on screen.

LicenseCC-BY 4.0 (attribution required)
AttributionAuthor attribution (recommended form):
FoodOn (Food Ontology), CC-BY 4.0

Dooley, D. M., Griffiths, E. J., Gosal, G. S., Buttigieg, P. L., Hoehndorf, R., Lange, M. C., Schriml, L. M., Brinkman, F. S. L., & Hsiao, W. W. L. (2018). FoodOn: a harmonized food ontology to increase global food traceability, quality control and data integration. npj Science of Food, 2, 18. https://doi.org/10.1038/s41538-018-0032-6

Snapshot / versionFoodOn release 2026-06-01, pinned to OBO PURL

epicure

Data observation

Ingredient-relationship model, observation asset (co-occurrence COOC + chemistry CHEM + core CORE relations)

Produces reproducible pairing evidence: the observation that two ingredients stand in a COOC, CHEM, or CORE relationship across three pinned embedding spaces. Model evidence is a first-class basis for pairing directions. It is a different category from literature facts, not a lesser grade of them: it is never promoted to a “Source fact,” and it can never support a food-safety premise, because the model carries no safety information.

LicenseCC-BY 4.0 (attribution required)
AttributionRequired attribution:
epicure ingredient-relation model, Kaikaku, CC-BY 4.0
MethodTop-32 cosine kNN pairing edges derived from three pinned embedding spaces (cooc / chem / core), reproducible against the published epicure_neighbors table: 171,840 edges = 1,790 ingredients × 32 neighbors × 3 spaces. Anyone can re-derive the shipped edges from the pinned repositories below and compare.
Snapshot / versionKaikaku/epicure-cooc@03edd311, epicure-chem@2461ef3f, epicure-core@d31ebb5a (pinned commits; full hashes in the approval manifest)

multilingual-e5-small

No badge: not evidence

Multilingual embedding model: tool asset (a retriever that finds candidates within the corpus and ontologies)

A search tool, not evidence. Similarity scores are used only for ranking and are never promoted to a measure of trust, and there is no path by which this asset’s output reaches the user as a sentence, which is why it wears no badge.

LicenseMIT
AttributionModel card notice (for reference):
intfloat/multilingual-e5-small, MIT

The MIT License (MIT) Copyright (c) Microsoft Corporation

Wang, L., Yang, N., Huang, X., Yang, L., Majumder, R., & Wei, F. (2024). Multilingual E5 Text Embeddings: A Technical Report. arXiv:2402.05672.

Snapshot / versionpinned by sha256 (hash value in the approval manifest)
Deduction

The fifth badge comes from a rule, not an asset

The “Deduction” badge comes not from the assets above but from the system’s own registered inference rules. Only the fixed literal sentences in the rule registry are used as evidence, and because they are not external data they are not subject to license attribution. The first release ships with a single rule.

R-CONTROLLED-COMPARISON: A controlled comparison is justified. Results are not guaranteed.