The NORM Framework: A Methodology for Measuring Claim–Evidence Proportionality in Consumer Wellness and High-Claim Brand Marketing
A structured framework for evaluating whether marketing claims are proportionate to publicly visible evidence — with primary application to consumer wellness, nutrition, and high-claim brand categories
NORM's methodology is open in the sense that its scoring dimensions, evidence hierarchy, scope limits, and interpretation rules are publicly documented. Exact scoring weights, thresholds, model prompts, anti-gaming controls, and implementation details remain proprietary to preserve system integrity and prevent manipulation.
Consumer brand marketing operates under conditions of significant information asymmetry, in which brands possess substantially more knowledge about their product formulations, underlying evidence, and marketing strategies than do the consumers who purchase them. Current regulatory frameworks — primarily Federal Trade Commission substantiation standards and relevant category-specific agency requirements — provide a legal floor that many brands meet while still leaving consumers materially uninformed about the quality and applicability of the evidence underlying marketing claims. This paper describes the NORM Framework, a systematic methodology for evaluating the degree to which consumer-facing brand marketing claims are proportionate to the evidence brands publicly disclose. We describe the theoretical basis for the framework, the operationalization of its six scoring dimensions, the positional weighting methodology, and the evidence classification hierarchy that underlies dimension scoring. We discuss validity, reliability, and the framework's explicit limitations. While the framework was initially applied to health and wellness categories — where claim intensity and evidence gaps are most acute — it is designed to be generalizable across consumer brand categories, with highest fidelity in evidence-driven and high-claim contexts.
The severity of this problem has accelerated materially since 2020. The broader global wellness economy is now measured in the trillions, with the Global Wellness Institute estimating global wellness spending at approximately $6.8 trillion in 2024 and projecting continued growth toward $9.8 trillion by 2029.[13] NORM does not treat that entire economy as directly scoreable. Its primary domain is the subset of high-claim consumer categories — including supplements, functional foods and beverages, skincare and cosmetics, wellness devices, and related health-adjacent products — where marketing representations substantially shape purchase decisions and where evidence visibility is often uneven. Within that domain, U.S. dietary supplement spending alone represents tens of billions annually, illustrating the scale of consumer decisions made in categories where claim–evidence alignment is frequently difficult for consumers to evaluate independently.
At the same time, AI-generated marketing content now produces scientifically credible-sounding claims at industrial scale with near-zero marginal cost, while consumer health decisions are increasingly self-directed, made outside clinical supervision and inside an information environment optimized to persuade rather than inform. In this context, claim–evidence misalignment is not an exception. It is a structural condition of the market. NORM makes that condition more measurable, easier to examine, and easier to compare across brands — and creates a durable record of what each brand chooses to show and what it does not, in the consumer-facing environment where purchase decisions are actually made.
NORM evaluates representation, not reality. A score reflects what a brand publicly discloses, not what its products do — and in a market where consumer decisions are made almost entirely on what brands choose to show, that distinction is the point.
1. Introduction
1.1 The Problem of Information Asymmetry in Consumer Brand Marketing
George Akerlof's foundational analysis of information asymmetry demonstrated that when one party to a transaction possesses materially superior information, market outcomes systematically disadvantage the less-informed party.[1] This dynamic is structural in consumer brand markets, where brands routinely possess detailed knowledge of their formulations, clinical testing history, and the methodological quality of studies they reference — while consumers operate with minimal capacity to evaluate these factors independently.
The scale of this asymmetry is substantial. Consumer categories in which marketing claims are primary purchase drivers — including dietary supplements, skincare and cosmetics, functional foods and beverages, wellness devices, and financial wellness products — represent a significant subset of a broader global wellness economy measured in the trillions. These categories are not unified by product type. They are unified by claim intensity: consumers are asked to make purchase decisions based on representations about health, performance, appearance, cognition, recovery, longevity, or financial well-being. Yet rigorous, accessible evaluation of the quality of evidence underlying those representations has remained largely absent from consumer-facing information environments. Consumers encounter claims of clinical validation, peer-reviewed support, and scientific consensus without any structured means of evaluating what those phrases actually signify in context.
The problem is not new. What is new is its magnitude, its mechanism, and its trajectory. The gap between what wellness brands claim and what they publicly substantiate is not narrowing — it is widening, and it is widening at the precise moment consumers are directing more health decisions, more spending, and more trust toward this category than at any prior point in its history. This paper describes a framework designed to make that gap more measurable, easier to examine, and more persistent as a public record.
1.2 Regulatory Context and Its Limitations
The U.S. regulatory framework for wellness product claims is distributed across multiple agencies with differing standards and enforcement capacities:
- FDA structure-function claims: Under the Dietary Supplement Health and Education Act (DSHEA) of 1994, manufacturers may make structure-function claims without pre-market FDA review, provided claims do not assert treatment of a disease and the manufacturer maintains substantiation files that are not required to be publicly disclosed.[3]
- FTC advertising substantiation: The FTC requires that health-related advertising claims be substantiated by "competent and reliable scientific evidence." Under its 2022 Health Products Compliance Guidance, FTC guidance emphasizes the need for well-controlled human clinical testing for many health-related efficacy claims — though the standard is applied reactively rather than proactively, and the burden of substantiation varies by claim type and category.[4]
- FTC enforcement limitations: FTC enforcement applies only to demonstrated deception, leaving a substantial gray zone of technically compliant but materially misleading claims. A brand may cite a single small-n, industry-funded pilot study as "clinical evidence" and remain legally compliant while substantially misrepresenting the certainty of claimed effects to consumers.
These frameworks establish a legal minimum. They do not evaluate the quality, relevance, or consumer interpretability of evidence. NORM operates in the space between legal compliance and genuine epistemic transparency.
1.3 Contribution of This Work
The NORM Framework operates in the space between legal compliance and genuine epistemic transparency — a space that current regulatory frameworks explicitly do not occupy. It is not a supplement to existing oversight. It fills a structural absence.
NORM makes the claim–evidence gap legible. It identifies where marketing claims may exceed what brands publicly substantiate, creating a structured public record of what each brand chooses to show — in the specific information environment where consumers decide. This record does not require a regulatory trigger, an enforcement action, or a consumer complaint; it is produced continuously and made publicly accessible without registration. This paper documents the methodology underlying that record.
Scores represent methodology-based assessments derived from publicly available information and do not constitute legal, medical, or regulatory determinations.
1.4 Why This Moment Demands a Framework
The structural conditions producing claim–evidence misalignment in consumer brand markets are not stable. Each of the following forces is accelerating simultaneously — and their confluence creates an information environment that is more consequential, more adversarially constructed, and more resistant to individual evaluation than at any prior point.
Record-scale consumer expenditure
The broader global wellness economy is measured in the trillions. The Global Wellness Institute estimated the global wellness economy at approximately $6.8 trillion in 2024 and projected growth toward $9.8 trillion by 2029.[13] NORM's relevant market is narrower but still substantial: high-claim consumer categories within and adjacent to wellness where purchase decisions are materially shaped by marketing representations. These include supplements, functional foods and beverages, skincare and cosmetics, wellness devices, and related health-adjacent products.
The relevant issue is not that this spending is "wasted," nor that every product in these categories is ineffective. NORM does not evaluate product efficacy. The issue is that large volumes of consumer spending occur in environments where claim–evidence alignment is not consistently visible. When consumers are asked to act on claims that appear clinical, scientific, or outcome-specific, but the supporting evidence is incomplete, difficult to locate, or disproportionate to the claim being made, allocation cannot be assumed to be fully informed. NORM evaluates that visibility gap.
AI-generated marketing claims at industrial scale
Generative AI has fundamentally altered the production economics of marketing claims. Brands can now produce thousands of evidence-sounding, mechanistically coherent, specifically stated claims at near-zero marginal cost. AI language models trained on the scientific literature generate plausible-sounding biological mechanisms, produce specific figures with the surface structure of quantitative evidence, and compose marketing copy structurally indistinguishable from well-supported scientific communication — without any of the underlying evidentiary basis.[14] The rate of claim production now far exceeds the capacity of any regulatory body to evaluate it reactively. NORM evaluates the output continuously, not retroactively.
Consumer trust erosion
High-claim consumer categories have historically benefited from a trust premium — consumers have attributed higher credibility to "natural," "science-backed," and "clinician-formulated" products than the evidentiary basis warrants. That premium is now under pressure. Survey data consistently show consumers reporting difficulty distinguishing genuine scientific evidence from marketing language, and awareness that health and wellness marketing is frequently not independently validated is growing.[15] Declining trust without a structured tool for evaluation does not protect consumers — it produces cynicism without discrimination, disadvantaging brands that invest in genuine evidence disclosure alongside those that do not. NORM creates the discrimination mechanism trust erosion demands.
The rise of self-directed decisions
The share of consequential consumer decisions made without expert guidance has increased materially across health, finance, nutrition, and performance categories. An estimated 67% of U.S. supplement purchases are made without clinician input; the pattern repeats across financial products, cosmeceuticals, and functional foods.[16] Consumers operating as their own evaluators require the same quality of epistemic infrastructure that expert intermediaries once provided. NORM is a component of that infrastructure — the part that applies to the consumer-facing brand environment where no such infrastructure currently exists.
2. Theoretical Framework
2.1 Claim–Evidence Alignment as a Construct
We define claim–evidence alignment (CEA) as the degree to which a brand's publicly visible marketing claims are proportionate to the quality and relevance of evidence the brand publicly discloses in support of those claims.
CEA is explicitly distinct from the following constructs that the NORM Framework does not measure:
- Product efficacy: Whether a product's active ingredients produce the claimed effects under any conditions
- Product safety: Whether a product is safe for any population or individual
- Regulatory compliance: Whether a brand meets FTC or FDA standards
- Ingredient quality: Whether a product contains the declared ingredients at declared concentrations
A high NORM score indicates stronger visible claim–evidence proportionality. It does not certify product efficacy, safety, legality, or medical value. A low NORM score means claims materially exceed visible evidence — a form of epistemic overreach that distorts consumer decision-making regardless of whether the underlying product may be effective.
2.2 Behavioral Economics Basis
Consumers exhibit well-documented cognitive tendencies that render them particularly susceptible to wellness marketing claims under ordinary decision conditions:
- Authority heuristics: References to "clinically proven," "scientifically shown," or "doctor-recommended" trigger automatic credibility attribution independent of the quality of underlying evidence, consistent with Cialdini's authority principle and Chaiken's heuristic-systematic model.[5,6]
- Natural product bias: Terms including "natural," "clean," and "pure" generate positive affect independent of safety or efficacy implications — a well-replicated finding with particular relevance to supplement marketing.[7]
- Specificity bias: Specific-sounding claims ("increases VO₂ max by 12%") are perceived as more credible than vague ones, even when specific figures derive from low-quality or poorly applicable studies — consistent with the persuasion literature on numerical specificity.[8]
- Dual-process susceptibility: Under conditions of low elaboration likelihood — characteristic of supplement and wellness purchase decisions made in browsing contexts — consumers rely on peripheral cues (packaging, marketing language, brand authority) rather than systematic evidence evaluation.[9]
- Testimonial bias: Personal narratives are processed preferentially over statistical information in health contexts, a phenomenon documented extensively in the cancer screening and pharmaceutical advertising literatures.[10]
These tendencies create conditions in which brands can substantially influence consumer behavior through claim framing, independent of the epistemic quality of those claims. CEA measurement makes this framing visible and quantifiable.
2.3 Positional Weighting Theory
Marketing claims do not carry equal informational influence based on their position within a consumer-facing page. Eye-tracking research consistently demonstrates that above-fold content receives disproportionate visual attention, with engagement declining nonlinearly with scroll depth.[11] Accordingly, NORM applies a positional weighting system to claims based on their location within the page information hierarchy.
| Position | Weight | Rationale |
|---|---|---|
| Above-fold hero headline | Highest | Highest visual salience; primary purchase signal; appears in social previews |
| Product name / primary descriptor | High | Persistent across all touchpoints; forms lasting mental association |
| Above-fold body copy | Elevated | High read-through probability; direct support for purchase decision |
| Below-fold primary section | Baseline | Baseline; standard consumer engagement for engaged visitor |
| Feature / ingredient list | Below baseline | Enumerated format; lower per-item salience; higher information density |
| FAQ / secondary pages | Reduced | Lower traffic volume; higher engagement depth when visited |
| Footer / disclaimer text | Lowest | Low visibility; characteristically used for legal hedging of hero claims |
Exact positional weight multipliers are proprietary calibration parameters.
This weighting reflects the principle that brands deliberately position their most impactful claims in highest-salience locations. A claim in a hero headline generates materially greater epistemic influence on consumers than the same claim buried in a footer disclaimer — and is accordingly weighted proportionately in scoring. Footer disclaimers that contradict hero claims without correcting them are specifically identified as Consumer Distortion signals (§3.6).
The specific multipliers are internally calibrated parameters derived from performance against the reference corpus; formal empirical validation against eye-tracking fixation and scroll-depth data is a planned v1.1 study.
3. Scoring Dimensions
The NORM Framework evaluates brands across six dimensions, producing a composite score of 0–100. Higher scores indicate stronger claim–evidence alignment. Dimension weights reflect the relative importance of each factor to overall CEA and were calibrated against an internal reference corpus spanning multiple high-claim consumer categories including supplements, nutrition, skincare, functional food and beverage, and wellness devices. The framework is designed primarily for consumer wellness and high-claim brand marketing; it is generalizable across consumer categories with category-specific evidence adaptations, though its evidence hierarchy is most directly applicable to health-adjacent product claims.
Definition: Whether the brand publicly discloses any evidence for its primary claims. EP is the threshold question in CEA evaluation — many brands make substantive health claims with zero visible supporting evidence.
EP carries the lowest explicit weight because its absence cascades through the framework: a claim with no visible evidence cannot earn meaningful Evidence Strength and typically scores poorly across Specificity, Effect Reality, and Consumer Distortion. The effective penalty for absent evidence is therefore substantially larger than EP's nominal ceiling alone suggests.
Definition: The methodological quality of evidence publicly disclosed. ES is the highest-weighted dimension because the quality of evidence, not merely its presence, determines its epistemic value to consumers.
Evidence is classified according to the following hierarchy, consistent with standard evidence-based medicine frameworks adapted for the supplement and wellness marketing context:[12]
Confidence adjustments: Industry funding applies a configurable confidence modifier. Studies cited but not hyperlinked receive a configurable accessibility modifier. Studies cited in a way that misrepresents their scope or conclusions are additionally penalized under Consumer Distortion (§3.6). ES is computed as the weighted mean of hierarchy scores across all cited evidence, normalized to the dimension's scale. Exact modifier values are proprietary calibration parameters.
Definition: The degree to which claims are precise and falsifiable rather than vague and hedged. Vague claims ("supports wellness," "promotes balance") are epistemically empty — they convey a benefit impression while committing to nothing a consumer could evaluate or falsify.
Each tier-1 claim is assessed on five specificity criteria (0–3 each):
Specificity scoring evaluates each claim against all five criteria, weighted by positional prominence and normalized to the dimension's scale. Claims using absolute language ("proven," "guaranteed") without meeting all five criteria receive a Specificity–Distortion flag that also affects CD scoring.
Definition: Whether the brand explains the biological or physiological mechanism by which a claimed effect occurs, and whether that explanation is consistent with published science. Mechanistic explanation is a hallmark of genuine evidence-based communication — it allows consumers to evaluate coherence, not merely accept assertion.
Common MI failure modes: "activates mitochondria" (vague, no pathway); "boosts serotonin" without receptor mechanism or clinical translation; citing in vitro evidence for human mechanism claims without translation caveat. Regulatory enforcement history — including documented FDA warning letters or FTC enforcement actions relevant to the claims assessed — is treated as a contextual risk signal, not as a standalone determinant of score. Where enforcement findings relate directly to claim–evidence mismatch, they may inform Consumer Distortion and Evidence Strength scoring.
Definition: Whether claimed effects are grounded in outcomes likely to manifest for the target consumer under realistic conditions of use. Evidence may exist for a claimed effect under conditions that do not apply to typical consumers — ER evaluates whether this gap is acknowledged and whether it materially distorts the implied benefit.
ER penalties are applied when: (a) evidence derives from materially different populations but is presented as general; (b) study doses exceed product doses without acknowledgment; (c) implied timelines contradict study duration; (d) study conditions are not reproducible by typical consumers without disclosure.
Definition: The degree to which marketing language, framing, or presentation techniques are likely to create materially false impressions among reasonable consumers, independent of the underlying evidence. Consumer distortion is evaluated relative to the interpretation a reasonable consumer would form under standard browsing conditions — a standard aligned with FTC deception analysis. CD is scored inversely — a full score indicates no distortion detected.
The following distortion signals are evaluated, each contributing a severity-scaled penalty based on positional prominence:
Each detected signal contributes a severity-scaled penalty. Severities are fixed; the penalty is then multiplied by the positional weight of the location in which the distortion appears:
| Signal | Severity | Rationale |
|---|---|---|
| AL — Absolute language | Configurable | High prevalence; directly inflates perceived certainty |
| CP — Cherry-picking | Configurable | Intentional misrepresentation of evidence body |
| SM — Study misrepresentation | Configurable | Directly distorts consumer understanding of evidence scope |
| PT — Pseudoscientific terminology | Configurable | Implies scientific basis where none exists |
| TE — Testimonial-as-evidence | Configurable | Common; partially mitigated by consumer awareness |
| CC — Causation from correlation | Configurable | Well-documented mechanism for false certainty attribution |
| FH — Footer contradiction | Configurable | Structurally conceals qualifications from point of claim encounter |
| RO — Risk omission | Configurable | Directly affects safety-relevant consumer decisions |
| IE — Influencer echo amplification | Configurable | Exploits parasocial trust and bypasses evidence-evaluation heuristics at scale; severity increases when FTC disclosure non-compliance is identified [21,22,23] |
| Exact severity values are proprietary calibration parameters. | ||
4. Automated Review Pipeline
NORM uses automated capture, claim extraction, evidence resolution, and scoring workflows. Outputs may be routed through automated validation and limited quality-control review before public display. Review is used to identify capture, processing or publication issues—not to replace the scoring methodology with editorial judgment.
Concurrent execution of multiple independent capture methods — including structured content extraction, dynamic rendering, and visual analysis — operating in parallel. This multi-source approach ensures that dynamic content, interactive elements, and image-based claims are identified with high fidelity.
Automatic retrieval of external evidence. The pipeline cross-references PubMed, ClinicalTrials.gov, and independent testing databases, while parsing complex PDF whitepapers to build a structured evidence snapshot for every claim.
Each claim is independently adjudicated against matched evidence, assigned a substantiation verdict and weighted penalty, and aggregated to a brand-level composite score using positional weighting. Scoring is fully deterministic — no model-based judgment is applied at the scoring stage.
4.1 Weighted Aggregation
The composite score is computed as a positionally-weighted aggregation of individual claim-level dimension scores. Each claim receives a positional weight drawn from the position table in §2.3. Claims in high-prominence positions contribute proportionally more to the brand-level score than those in lower positions. Each claim is evaluated across all six dimensions, and the results are aggregated using proprietary weighting to produce the composite brand score on a 0–100 scale.
This formulation ensures that a brand with one strongly-evidenced hero claim and ten unsupported secondary claims is not rewarded by simple averaging. The positional weighting anchors the score to the consumer-facing prominence distribution of the actual claims evaluated.
4.2 Signal Classification
| Score | Signal | Interpretation |
|---|---|---|
| Upper range | Green | Claims are specific and backed by visible evidence of appropriate quality and scope |
| Middle range | Yellow | Some claims are supported; marketing extends beyond what is fully disclosed |
| Lower range | Red | Claims materially exceed visible evidence; significant epistemic gap |
Threshold calibration was performed against an internal reference corpus spanning multiple high-claim consumer categories. Signal thresholds were set through iterative calibration against team assessments; formal interrater reliability measurement with independent reviewers is a planned future milestone. Reliability design targets for each dimension and overall signal classification are described in §5.
4.3 Claim-Level vs. Brand-Level Scoring
The NORM Framework operates at the claim level. Each individual claim on a brand's consumer-facing pages is evaluated and scored before being aggregated to the brand level. This approach avoids the masking effect that occurs when a brand's strong evidence for one ingredient or claim obscures weak or absent evidence for others — a common pattern in which a single well-supported ingredient claim is used to imply evidential support for an entire product line.
Brand-level scores are the positionally-weighted mean of claim-level dimension scores. Tier-1 claims (hero headlines, primary descriptors) are weighted at their positional multipliers; lower-position claims are downweighted accordingly.
5. Validation & Reliability
5.1 Pipeline Calibration
The NORM v1.0 scoring pipeline applies the framework through a deterministic evaluation architecture. Each identified claim is matched against the brand's visible evidence base, adjudicated for substantiation quality using rule-based criteria, assigned a severity-weighted penalty, and aggregated to brand level using the positionally-weighted mean formulation in §4.1. Claim extraction during the capture phase uses AI-assisted content analysis; scoring itself is fully deterministic.
The pipeline was calibrated by the NORM Analytics team against an internal reference corpus spanning multiple high-claim consumer categories. These reference scores served as the calibration target. Penalty weights, evidence hierarchy mappings, and substantiation thresholds were iteratively refined until pipeline output aligned with team assessments on signal classification (Green / Yellow / Red) across the calibration corpus.
Formal interrater reliability assessment — in which independent reviewers with no involvement in framework development apply the rubric to a blind corpus and Cohen's weighted kappa is computed against pipeline output — is the primary validation milestone for methodology v1.1. The design targets below represent the reliability levels the framework's operational criteria are structured to achieve per dimension:
| Dimension | Design Target κ | Interpretation |
|---|---|---|
| Evidence Presence (EP) | ≥ 0.85 | Almost perfect |
| Evidence Strength (ES) | ≥ 0.80 | Almost perfect |
| Specificity (SP) | ≥ 0.75 | Substantial |
| Mechanism Integrity (MI) | ≥ 0.75 | Substantial |
| Effect Reality (ER) | ≥ 0.70 | Substantial |
| Consumer Distortion (CD) | ≥ 0.80 | Almost perfect |
| Signal Classification (overall) | ≥ 0.78 | Substantial → Almost perfect |
Effect Reality (ER) carries the most demanding reliability requirement of any dimension, given its requirement that assessors evaluate population and dosage correspondence between cited studies and actual product specifications — a judgment dependent on full-text study access and product labeling detail. The ER rubric criteria are designed to minimize variance within the automated scoring context; residual ER uncertainty is one of the acknowledged limitations of v1.0 automated assessment.
Pipeline scoring is version-controlled. Each published score is associated with the specific scoring configuration and capture timestamp used at the time of publication. Model upgrades are not applied silently to published scores. Prior to any model migration, the internal reference corpus is rescored under the proposed configuration and compared against the prior production baseline. Material score drift triggers review before migration. Historical scores remain associated with the scoring stack used at the time of publication.
5.2 Construct Validity Predictions
The NORM Framework makes specific, falsifiable predictions about what construct validity testing should show. Publishing these predictions in advance of formal data collection is a commitment to falsifiability — a condition the framework's independence claims require.
- Convergent — FTC enforcement history: Brands with prior enforcement actions involving deceptive or inadequately substantiated claims should, on average, score lower on NORM than comparable brands without such histories. This prediction follows from the framework construct: claim–evidence misalignment detectable on consumer-facing pages should correlate with enforcement outcomes. NORM scores should detect this at the dimension level before enforcement occurs, not after.
- Convergent — Consumer trustworthiness perception: NORM scores should correlate positively with independent consumer assessments of brand credibility. If the construct NORM measures does not map onto layperson credibility judgment, the framework's consumer-protection rationale requires revision.
- Discriminant — Product formulation quality: NORM scores should show no significant correlation with independently assessed product formulation quality. A high NORM score means a brand shows its evidence proportionately — not that the product is effective. A brand with an excellent formulation and poor evidence disclosure should score low; a brand with a mediocre formulation and proportionate evidence disclosure should score higher. If scores correlate with product quality, the framework is measuring something other than what it claims.
Formal validity testing against each prediction — using the calibration corpus and a planned expanded corpus — will be published as data are collected and will inform methodology v1.1 refinements.
5.3 Calibration Corpus
The calibration corpus was constructed by the NORM Analytics team to ensure category breadth and score distribution coverage. Brands were selected to include representation across expected signal ranges, product categories (supplements, skincare/cosmetics, functional food/beverage, health devices, financial wellness), and revenue scale (from pre-revenue DTC brands to established CPG lines). Calibration corpus composition and size are not publicly disclosed to prevent targeted gaming of calibration-specific scoring patterns. An expanded corpus is planned for future reliability and validity assessment.
This calibration approach is bootstrap-limited: pipeline output is calibrated against team judgments produced by the framework developers, not against an independent external standard. External validity testing against the predictions in §5.2 is the primary mechanism for breaking this circularity in v1.1.
6. Data Collection Protocol
6.1 Pipeline Architecture
NORM v1.0 uses a deterministic scoring pipeline with quality-control and review workflows for selected outputs. For each brand, the pipeline performs five sequential stages: (1) page capture — rendering and extracting consumer-facing content including dynamically loaded elements; (2) claim identification and positional classification — the scoring engine identifies marketing claims and assigns each a position tier from the §2.3 table; (3) evidence resolution — URLs and citation references in brand pages are followed and resolved to determine evidence presence and assess evidence tier; (4) claim adjudication and scoring — each identified claim is matched against resolved evidence, adjudicated for substantiation quality, and assigned a weighted penalty; (5) brand-level aggregation via the positionally-weighted mean formula in §4.1.
Outputs undergo automated consistency checks. Scores near publication boundaries or carrying capture and variance flags may receive additional validation before publication. The pipeline operates on a refresh cycle; refreshed scores are published without re-review unless a material change is detected.
6.2 Page Scope and Capture Confidence
Pages included in scoring: homepage, product detail pages, ingredient and science pages, FAQ and education pages, clinical research pages. Pages excluded: press releases, investor materials, third-party retailer pages, social media content, and advertising placements. Brands whose content cannot be fully captured receive a Low Confidence classification reflecting the gap in capturable content; their scores are displayed with an explicit confidence flag and excluded from comparative rankings until a full capture is available. Visual claims present only in images or video are noted as uncaptured and do not influence scores in v1.0; structured visual claim capture is a planned v1.1 capability.
6.3 Evidence Resolution
Citation references identified in brand pages — including hyperlinked URLs, DOI references, and named study titles — are resolved by the pipeline to determine evidence presence and tier. Links that resolve to accessible study abstracts or full texts are evaluated for study design metadata (population, design type, sample size, funding source where disclosed). Citations present in brand copy that do not resolve to an accessible study receive reduced EP credit and an Unresolvable Citation note in the claim record. Studies cited without any link or identifier are scored at the minimum evidence presence tier.
7. Comparative Context
The NORM Framework occupies a distinct position relative to existing evaluation mechanisms. Understanding what distinguishes NORM from these mechanisms clarifies the gap it fills and the claims it does not make.
8. Adversarial Robustness
Any public scoring system creates incentives for scored entities to optimize toward the signal rather than the underlying construct. The following five adversarial vectors were identified during framework development. Each is addressed by structural mitigations within the NORM scoring architecture.
The scoring framework is structured so that diluting claims does not improve composite scores. Vague language forfeits evidence-linked scoring opportunities and may trigger visibility flags in the public record.
Evidence scoring distinguishes between citation volume and citation quality. Low-tier citation flooding does not improve composite scores due to quality-weighting and scoring caps built into the evidence dimensions.
Content divergence between capture sessions triggers re-review. The framework is designed to surface and flag content-serving inconsistencies through periodic multi-session comparison.
Score history is stored per brand. The framework treats score change velocity as a signal in its own right, distinguishing reactive cosmetic changes from sustained, genuine evidence improvements across refresh cycles.
Consumer Distortion (CD) is scored independently from specificity and evidence dimensions, ensuring that technically compliant language which misrepresents certainty still generates distortion penalties. Methodology versioning supports calibration updates as new gaming patterns are identified.
Design principle: No scoring system is perfectly robust to adversarial optimization — especially a public one. NORM's response is transparency: documented vectors, a public record that surfaces manipulation signals alongside scores, and an architecture designed to evolve in response to adversarial adaptation.
Exact anti-gaming parameters, detection thresholds, and operational controls are proprietary.
9. Limitations
The following limitations are explicit constraints of the NORM Framework. They are not deficiencies to be resolved — they are inherent to the defined scope of what the framework measures and what it does not.
9.1 NORM Does Not Evaluate Product Efficacy
NORM measures whether brands show their evidence — not whether the underlying products are effective, safe, or appropriate for any individual. A high score means a brand publicly discloses proportionate evidence. A low score means claims exceed visible evidence. Neither score constitutes a finding about product performance.
9.2 Temporal Limitations
Scores reflect a point-in-time assessment of publicly accessible pages. Brands update pages; evidence bases evolve. NORM scores are refreshed periodically but cannot guarantee currency between refresh cycles.
9.3 Private Evidence Files
DSHEA permits manufacturers to hold substantiation files that are not required to be publicly disclosed. A brand may hold robust private evidence while scoring poorly on NORM because that evidence has not been made publicly visible. This is an explicit design choice: NORM evaluates the consumer-facing information environment specifically, not the regulatory file.
9.4 Off-Page Marketing Channels
Paid social advertising, podcast sponsorships, and individual influencer posts are not directly scored — NORM does not crawl off-domain channels. However, the structural footprint of influencer programs is partially captured when brands surface it on their own pages: ambassador network landing pages, affiliate program infrastructure, and creator program disclosures are scored under the IE (Influencer Echo Amplification) signal within Consumer Distortion (§3.6). This captures the category of brands that deploy influencer reach at a scale materially disproportionate to their evidence disclosure — the pattern where paid social consensus substitutes for verifiable evidence — without requiring off-page crawling. Full off-page channel coverage remains a meaningful v2.0 expansion direction.
9.5 Non-U.S. Regulatory Contexts
The framework was developed with reference to U.S. regulatory standards (FTC, FDA/DSHEA). Application to brands operating primarily under EU, UK, or other regulatory frameworks requires contextual adaptation of the Consumer Distortion criteria.
10. Independence and Ethical Framework
10.1 Structural Independence
NORM Framework scores are produced by an automated pipeline that is structurally separated from commercial relationships. Brand Pro and Data API subscribers receive workflow tools and analytical access to their own brand data. The pipeline operates independently without access to commercial account information. No brand can purchase inclusion, a higher score, favorable treatment, or suppression of a score. Revenue streams are structurally separated from scoring functions — this separation is enforced architecturally, not by policy alone.
10.2 Conflict of Interest Policy
No individual involved in evaluating a brand may hold a financial interest in that brand. NORM Analytics does not accept advertising revenue, affiliate commissions, or brand-sponsored content of any kind.
10.3 Transparency Commitments
- This methodology document is publicly available and versioned
- Score rationale is disclosed in public claim records accessible without registration
- Brands may submit evidence for pipeline evaluation through the Brand Portal; submissions are evaluated on the same criteria as all other evidence and do not guarantee score changes
- Score disputes are addressed through a documented review process with written rationale
- Methodology updates are versioned and publicly noted with effective dates
10.4 Provenance and Capture Trail
Every NORM score is linked to a timestamped capture record. The capture trail includes the pages evaluated, the date and time of capture, the scoring configuration version used, and an evidence packet summarizing the claims identified and evidence resolved. This audit trail is retained internally and can be referenced in any dispute or review process. Scores are not retroactively modified without a new capture cycle and re-evaluation against the current methodology version.
10.5 Public Output Governance
Scores are published to consumer-facing surfaces only after passing internal quality-control gates. No score is made public automatically without review. Outputs that fall in boundary ranges or exhibit low capture confidence are held for additional review before publication. Published scores reflect the methodology version and capture timestamp in effect at the time of evaluation. NORM does not publish scores influenced by commercial relationships, and no brand can purchase expedited review, favorable treatment, or suppression of a published score.
This methodology is in active development. NORM v1.0 is a beta release. Scoring parameters, capture scope, and review workflows may be updated between versions. Material methodology changes will be publicly noted with version numbers and effective dates.
11. Conclusion
Consumer brand markets exhibit systematic information asymmetry that current regulatory frameworks address incompletely — and that is widening as expenditure grows, claim production industrializes through AI, and consumers make consequential decisions with minimal independent infrastructure for evaluating what they are told. In this environment, the absence of a structured, publicly available, consumer-facing evaluation of claim–evidence alignment is not a gap. It is a structural gap in current market and regulatory design — and it is not specific to one category.
The NORM Framework is a response to that structural gap. By evaluating marketing claims at the individual level, applying positional weighting grounded in visual attention research, classifying evidence quality through a hierarchy consistent with evidence-based medicine standards, and producing scores interpretable by a non-expert consumer, NORM makes the epistemic gap between what brands claim and what they publicly substantiate visible, measurable, and persistent — in any consumer category where that gap exists.
NORM does not determine truth; it measures proportionality between what is claimed and what is shown. It does not require a regulatory trigger, a complaint, or an enforcement action. It operates on the consumer-facing page — the specific environment where purchase decisions are made — and produces a structured public record that updates as brands update their pages.
NORM is designed to provide a structured, independent, publicly accessible record of whether brand claims are proportionate to what the brand actually shows. For consumers making health or financial commitments based on brand marketing, this record makes the claim–evidence gap legible. For brands that have invested in genuine evidence disclosure, it creates a durable basis for differentiation. For the market as a whole, it is offered as scalable infrastructure — focused on consumer wellness and high-claim brand marketing, with an evidence framework grounded in that domain.
The framework is explicitly limited in scope: it measures what brands choose to show, not what products do. That limitation is also its precision. The consumer-facing page is where the decision happens. That is where the record belongs.
11.1 The Emerging Role of Machine-Readable Evidence
There is a second, forward-looking reason this framework exists.
A growing share of consumer purchasing decisions may increasingly be mediated, filtered, or informed by AI-powered tools. Shopping assistants, recommendation engines, and comparison platforms are beginning to evaluate brands not on marketing aesthetics or influencer reach, but on machine-readable evidence signals: structured claim-evidence mappings, verifiable study citations, transparent dosage disclosures, and quantifiable outcome data. These systems are designed to evaluate what is provable, not what is persuasive.
If that trajectory continues — and early indicators in AI-assisted commerce and agent-mediated product evaluation suggest it may — brands with strong claim-evidence alignment will be better positioned. Brands whose marketing relies primarily on persuasion techniques that work on human cognitive biases but carry no machine-verifiable evidence layer may face increasing friction in AI-mediated discovery and recommendation environments.
The NORM scoring dimensions described in this paper — Evidence Presence, Evidence Strength, Specificity, Mechanism Integrity, Effect Reality, and Consumer Distortion — are designed to be both human-readable transparency measures and structured evidence records. Every brand scored by NORM receives an assessment of the kinds of signals that evidence-aware systems are likely to evaluate.
Brands that invest in genuine evidence disclosure are not just improving their NORM scores — they are building the evidence infrastructure that an increasingly automated commerce environment may require as a baseline condition for visibility.
A Green signal is not a certification. It is a record — built from what a brand chooses to make visible, and how proportionate that is to what it claims. Any brand can earn it. The ones that already can are the ones this framework was built to find.
The following illustrates the conceptual application of the NORM Framework to a representative hypothetical brand. Exact dimension weights, severity values, positional multipliers, and composite scoring formulas are proprietary calibration parameters and are not disclosed. This illustration demonstrates how the framework operates at a conceptual level.
Brand Profile
Category: Dietary supplement (cognitive health). Three primary claims are evaluated, each assigned a positional weight based on page location consistent with the hierarchy described in §2.3.
| # | Claim | Page Position | Weight |
|---|---|---|---|
| 1 | "Clinically proven to improve focus within 30 days" | Hero headline | Highest |
| 2 | "Backed by 14 peer-reviewed studies" | Above-fold body copy | Elevated |
| 3 | "Supports long-term memory consolidation" | Below-fold product section | Baseline |
Evidence provided: One hyperlinked industry-funded RCT (n = 42, single-blind) on ingredient A; 13 citations to in vitro and animal model studies for remaining ingredients. No independent replication. No dose-matching disclosure.
Conceptual Claim-Level Analysis
ES: Single mid-tier study (industry-funded, small n, single-blind); confidence modifiers applied for funding source.
SP: Outcome and temporal specificity met; population not defined — reduced specificity.
ER: Study population was cognitively impaired older adults; product marketed to healthy working adults; dose correspondence not disclosed — significant reality gap.
CD: "Clinically proven" absolute language in highest-prominence position generates substantial penalty. Presenting a single-blind n=42 pilot as clinical proof constitutes study misrepresentation. Hero claim contradicted by footer disclaimer adds additional penalty. Compounding distortion signals drive this claim's CD score significantly downward.
Result: Low claim score — Red territory. The combination of strong absolute language, study-population mismatch, and compounding CD penalties places this claim firmly in the lower scoring range.
ES: 13 of 14 cited studies are low-tier (in vitro or animal model); the "14 peer-reviewed" framing implies a stronger evidence body than exists.
SP: Claim makes no specific outcome assertion — scores poorly on outcome, magnitude, population, and temporal subcriteria.
ER: In vitro / animal evidence body; no human efficacy data for most ingredients cited.
CD: Cherry-picking (presenting a primarily in vitro/animal evidence body as peer-reviewed backing) and study misrepresentation (framing low-tier studies as implying human efficacy) generate significant compounding penalties.
Result: Low claim score — Red territory. Despite high evidence presence, the evidence quality and distortion patterns produce a low overall score.
ES: Single low-tier citation; no human evidence for memory consolidation outcome.
SP: Outcome named ("memory consolidation") but no magnitude, population, or timeframe specified.
MI: Mechanism plausibly stated; no mechanistic overreach detected.
ER: In vitro evidence for a human outcome claim; mild mismatch, no dose or population contradiction.
CD: Hedged language ("supports") avoids absolute claims; mechanism stated without overreach. Minimal distortion detected.
Result: Low-to-moderate claim score. Weak evidence presence and strength anchor the score low, but the relatively clean distortion profile partially offsets.
Composite Score Outcome
The brand's composite score is computed as the positionally-weighted mean of individual claim scores (§4.1). Because the hero claim carries the highest positional weight, its compounding CD penalties and evidence-reality gaps dominate the composite. The resulting brand score falls in the Red signal range.