Public Methodology · NORM Signal v1.0

The NORM Framework: A Methodology for Measuring Claim–Evidence Proportionality in Consumer Wellness and High-Claim Brand Marketing

A structured framework for evaluating whether marketing claims are proportionate to publicly visible evidence — with primary application to consumer wellness, nutrition, and high-claim brand categories

Open Methodology. Proprietary Scoring Engine.

NORM's methodology is open in the sense that its scoring dimensions, evidence hierarchy, scope limits, and interpretation rules are publicly documented. Exact scoring weights, thresholds, model prompts, anti-gaming controls, and implementation details remain proprietary to preserve system integrity and prevent manipulation.

Abstract

Consumer brand marketing operates under conditions of significant information asymmetry, in which brands possess substantially more knowledge about their product formulations, underlying evidence, and marketing strategies than do the consumers who purchase them. Current regulatory frameworks — primarily Federal Trade Commission substantiation standards and relevant category-specific agency requirements — provide a legal floor that many brands meet while still leaving consumers materially uninformed about the quality and applicability of the evidence underlying marketing claims. This paper describes the NORM Framework, a systematic methodology for evaluating the degree to which consumer-facing brand marketing claims are proportionate to the evidence brands publicly disclose. We describe the theoretical basis for the framework, the operationalization of its six scoring dimensions, the positional weighting methodology, and the evidence classification hierarchy that underlies dimension scoring. We discuss validity, reliability, and the framework's explicit limitations. While the framework was initially applied to health and wellness categories — where claim intensity and evidence gaps are most acute — it is designed to be generalizable across consumer brand categories, with highest fidelity in evidence-driven and high-claim contexts.

The severity of this problem has accelerated materially since 2020. The broader global wellness economy is now measured in the trillions, with the Global Wellness Institute estimating global wellness spending at approximately $6.8 trillion in 2024 and projecting continued growth toward $9.8 trillion by 2029.[13] NORM does not treat that entire economy as directly scoreable. Its primary domain is the subset of high-claim consumer categories — including supplements, functional foods and beverages, skincare and cosmetics, wellness devices, and related health-adjacent products — where marketing representations substantially shape purchase decisions and where evidence visibility is often uneven. Within that domain, U.S. dietary supplement spending alone represents tens of billions annually, illustrating the scale of consumer decisions made in categories where claim–evidence alignment is frequently difficult for consumers to evaluate independently.

At the same time, AI-generated marketing content now produces scientifically credible-sounding claims at industrial scale with near-zero marginal cost, while consumer health decisions are increasingly self-directed, made outside clinical supervision and inside an information environment optimized to persuade rather than inform. In this context, claim–evidence misalignment is not an exception. It is a structural condition of the market. NORM makes that condition more measurable, easier to examine, and easier to compare across brands — and creates a durable record of what each brand chooses to show and what it does not, in the consumer-facing environment where purchase decisions are actually made.

NORM evaluates representation, not reality. A score reflects what a brand publicly discloses, not what its products do — and in a market where consumer decisions are made almost entirely on what brands choose to show, that distinction is the point.

Important scope note: NORM evaluates the alignment between marketing claims and the evidence brands make publicly visible. Scores are not legal determinations, medical advice, or regulatory compliance findings.

1. Introduction

1.1 The Problem of Information Asymmetry in Consumer Brand Marketing

George Akerlof's foundational analysis of information asymmetry demonstrated that when one party to a transaction possesses materially superior information, market outcomes systematically disadvantage the less-informed party.[1] This dynamic is structural in consumer brand markets, where brands routinely possess detailed knowledge of their formulations, clinical testing history, and the methodological quality of studies they reference — while consumers operate with minimal capacity to evaluate these factors independently.

The scale of this asymmetry is substantial. Consumer categories in which marketing claims are primary purchase drivers — including dietary supplements, skincare and cosmetics, functional foods and beverages, wellness devices, and financial wellness products — represent a significant subset of a broader global wellness economy measured in the trillions. These categories are not unified by product type. They are unified by claim intensity: consumers are asked to make purchase decisions based on representations about health, performance, appearance, cognition, recovery, longevity, or financial well-being. Yet rigorous, accessible evaluation of the quality of evidence underlying those representations has remained largely absent from consumer-facing information environments. Consumers encounter claims of clinical validation, peer-reviewed support, and scientific consensus without any structured means of evaluating what those phrases actually signify in context.

The problem is not new. What is new is its magnitude, its mechanism, and its trajectory. The gap between what wellness brands claim and what they publicly substantiate is not narrowing — it is widening, and it is widening at the precise moment consumers are directing more health decisions, more spending, and more trust toward this category than at any prior point in its history. This paper describes a framework designed to make that gap more measurable, easier to examine, and more persistent as a public record.

1.2 Regulatory Context and Its Limitations

The U.S. regulatory framework for wellness product claims is distributed across multiple agencies with differing standards and enforcement capacities:

  • FDA structure-function claims: Under the Dietary Supplement Health and Education Act (DSHEA) of 1994, manufacturers may make structure-function claims without pre-market FDA review, provided claims do not assert treatment of a disease and the manufacturer maintains substantiation files that are not required to be publicly disclosed.[3]
  • FTC advertising substantiation: The FTC requires that health-related advertising claims be substantiated by "competent and reliable scientific evidence." Under its 2022 Health Products Compliance Guidance, FTC guidance emphasizes the need for well-controlled human clinical testing for many health-related efficacy claims — though the standard is applied reactively rather than proactively, and the burden of substantiation varies by claim type and category.[4]
  • FTC enforcement limitations: FTC enforcement applies only to demonstrated deception, leaving a substantial gray zone of technically compliant but materially misleading claims. A brand may cite a single small-n, industry-funded pilot study as "clinical evidence" and remain legally compliant while substantially misrepresenting the certainty of claimed effects to consumers.

These frameworks establish a legal minimum. They do not evaluate the quality, relevance, or consumer interpretability of evidence. NORM operates in the space between legal compliance and genuine epistemic transparency.

1.3 Contribution of This Work

The NORM Framework operates in the space between legal compliance and genuine epistemic transparency — a space that current regulatory frameworks explicitly do not occupy. It is not a supplement to existing oversight. It fills a structural absence.

NORM makes the claim–evidence gap legible. It identifies where marketing claims may exceed what brands publicly substantiate, creating a structured public record of what each brand chooses to show — in the specific information environment where consumers decide. This record does not require a regulatory trigger, an enforcement action, or a consumer complaint; it is produced continuously and made publicly accessible without registration. This paper documents the methodology underlying that record.

Scores represent methodology-based assessments derived from publicly available information and do not constitute legal, medical, or regulatory determinations.

Critical scope distinction: NORM evaluates what is visible on consumer-facing brand pages — not the private evidence files brands may maintain for regulatory purposes. This distinction is intentional. Consumers make purchase decisions based on what they can see, not what brands hold privately. The relevant information environment for consumer protection purposes is the consumer-facing page, not the regulatory filing cabinet. A brand that holds robust private evidence but presents bare assertions to consumers is, from the consumer's perspective, making bare assertions — and is scored accordingly.

1.4 Why This Moment Demands a Framework

The structural conditions producing claim–evidence misalignment in consumer brand markets are not stable. Each of the following forces is accelerating simultaneously — and their confluence creates an information environment that is more consequential, more adversarially constructed, and more resistant to individual evaluation than at any prior point.

Record-scale consumer expenditure

The broader global wellness economy is measured in the trillions. The Global Wellness Institute estimated the global wellness economy at approximately $6.8 trillion in 2024 and projected growth toward $9.8 trillion by 2029.[13] NORM's relevant market is narrower but still substantial: high-claim consumer categories within and adjacent to wellness where purchase decisions are materially shaped by marketing representations. These include supplements, functional foods and beverages, skincare and cosmetics, wellness devices, and related health-adjacent products.

The relevant issue is not that this spending is "wasted," nor that every product in these categories is ineffective. NORM does not evaluate product efficacy. The issue is that large volumes of consumer spending occur in environments where claim–evidence alignment is not consistently visible. When consumers are asked to act on claims that appear clinical, scientific, or outcome-specific, but the supporting evidence is incomplete, difficult to locate, or disproportionate to the claim being made, allocation cannot be assumed to be fully informed. NORM evaluates that visibility gap.

AI-generated marketing claims at industrial scale

Generative AI has fundamentally altered the production economics of marketing claims. Brands can now produce thousands of evidence-sounding, mechanistically coherent, specifically stated claims at near-zero marginal cost. AI language models trained on the scientific literature generate plausible-sounding biological mechanisms, produce specific figures with the surface structure of quantitative evidence, and compose marketing copy structurally indistinguishable from well-supported scientific communication — without any of the underlying evidentiary basis.[14] The rate of claim production now far exceeds the capacity of any regulatory body to evaluate it reactively. NORM evaluates the output continuously, not retroactively.

Consumer trust erosion

High-claim consumer categories have historically benefited from a trust premium — consumers have attributed higher credibility to "natural," "science-backed," and "clinician-formulated" products than the evidentiary basis warrants. That premium is now under pressure. Survey data consistently show consumers reporting difficulty distinguishing genuine scientific evidence from marketing language, and awareness that health and wellness marketing is frequently not independently validated is growing.[15] Declining trust without a structured tool for evaluation does not protect consumers — it produces cynicism without discrimination, disadvantaging brands that invest in genuine evidence disclosure alongside those that do not. NORM creates the discrimination mechanism trust erosion demands.

The rise of self-directed decisions

The share of consequential consumer decisions made without expert guidance has increased materially across health, finance, nutrition, and performance categories. An estimated 67% of U.S. supplement purchases are made without clinician input; the pattern repeats across financial products, cosmeceuticals, and functional foods.[16] Consumers operating as their own evaluators require the same quality of epistemic infrastructure that expert intermediaries once provided. NORM is a component of that infrastructure — the part that applies to the consumer-facing brand environment where no such infrastructure currently exists.

The convergence: More money, more sophisticated claims, more distrust, and more consequential self-directed decisions — all accelerating simultaneously. This is not a trend requiring a future response. It is an active condition requiring a present one. The NORM Framework is designed as one response to that condition, applied systematically to the information environment where these forces meet consumers.

2. Theoretical Framework

2.1 Claim–Evidence Alignment as a Construct

We define claim–evidence alignment (CEA) as the degree to which a brand's publicly visible marketing claims are proportionate to the quality and relevance of evidence the brand publicly discloses in support of those claims.

CEA is explicitly distinct from the following constructs that the NORM Framework does not measure:

  • Product efficacy: Whether a product's active ingredients produce the claimed effects under any conditions
  • Product safety: Whether a product is safe for any population or individual
  • Regulatory compliance: Whether a brand meets FTC or FDA standards
  • Ingredient quality: Whether a product contains the declared ingredients at declared concentrations

A high NORM score indicates stronger visible claim–evidence proportionality. It does not certify product efficacy, safety, legality, or medical value. A low NORM score means claims materially exceed visible evidence — a form of epistemic overreach that distorts consumer decision-making regardless of whether the underlying product may be effective.

2.2 Behavioral Economics Basis

Consumers exhibit well-documented cognitive tendencies that render them particularly susceptible to wellness marketing claims under ordinary decision conditions:

  • Authority heuristics: References to "clinically proven," "scientifically shown," or "doctor-recommended" trigger automatic credibility attribution independent of the quality of underlying evidence, consistent with Cialdini's authority principle and Chaiken's heuristic-systematic model.[5,6]
  • Natural product bias: Terms including "natural," "clean," and "pure" generate positive affect independent of safety or efficacy implications — a well-replicated finding with particular relevance to supplement marketing.[7]
  • Specificity bias: Specific-sounding claims ("increases VO₂ max by 12%") are perceived as more credible than vague ones, even when specific figures derive from low-quality or poorly applicable studies — consistent with the persuasion literature on numerical specificity.[8]
  • Dual-process susceptibility: Under conditions of low elaboration likelihood — characteristic of supplement and wellness purchase decisions made in browsing contexts — consumers rely on peripheral cues (packaging, marketing language, brand authority) rather than systematic evidence evaluation.[9]
  • Testimonial bias: Personal narratives are processed preferentially over statistical information in health contexts, a phenomenon documented extensively in the cancer screening and pharmaceutical advertising literatures.[10]

These tendencies create conditions in which brands can substantially influence consumer behavior through claim framing, independent of the epistemic quality of those claims. CEA measurement makes this framing visible and quantifiable.

2.3 Positional Weighting Theory

Marketing claims do not carry equal informational influence based on their position within a consumer-facing page. Eye-tracking research consistently demonstrates that above-fold content receives disproportionate visual attention, with engagement declining nonlinearly with scroll depth.[11] Accordingly, NORM applies a positional weighting system to claims based on their location within the page information hierarchy.

PositionWeightRationale
Above-fold hero headlineHighestHighest visual salience; primary purchase signal; appears in social previews
Product name / primary descriptorHighPersistent across all touchpoints; forms lasting mental association
Above-fold body copyElevatedHigh read-through probability; direct support for purchase decision
Below-fold primary sectionBaselineBaseline; standard consumer engagement for engaged visitor
Feature / ingredient listBelow baselineEnumerated format; lower per-item salience; higher information density
FAQ / secondary pagesReducedLower traffic volume; higher engagement depth when visited
Footer / disclaimer textLowestLow visibility; characteristically used for legal hedging of hero claims

Exact positional weight multipliers are proprietary calibration parameters.

This weighting reflects the principle that brands deliberately position their most impactful claims in highest-salience locations. A claim in a hero headline generates materially greater epistemic influence on consumers than the same claim buried in a footer disclaimer — and is accordingly weighted proportionately in scoring. Footer disclaimers that contradict hero claims without correcting them are specifically identified as Consumer Distortion signals (§3.6).

The specific multipliers are internally calibrated parameters derived from performance against the reference corpus; formal empirical validation against eye-tracking fixation and scroll-depth data is a planned v1.1 study.

3. Scoring Dimensions

The NORM Framework evaluates brands across six dimensions, producing a composite score of 0–100. Higher scores indicate stronger claim–evidence alignment. Dimension weights reflect the relative importance of each factor to overall CEA and were calibrated against an internal reference corpus spanning multiple high-claim consumer categories including supplements, nutrition, skincare, functional food and beverage, and wellness devices. The framework is designed primarily for consumer wellness and high-claim brand marketing; it is generalizable across consumer categories with category-specific evidence adaptations, though its evidence hierarchy is most directly applicable to health-adjacent product claims.

Evidence Presence
EP

Definition: Whether the brand publicly discloses any evidence for its primary claims. EP is the threshold question in CEA evaluation — many brands make substantive health claims with zero visible supporting evidence.

EP carries the lowest explicit weight because its absence cascades through the framework: a claim with no visible evidence cannot earn meaningful Evidence Strength and typically scores poorly across Specificity, Effect Reality, and Consumer Distortion. The effective penalty for absent evidence is therefore substantially larger than EP's nominal ceiling alone suggests.

10
Evidence cited and hyperlinked for all tier-1 claims; sources accessible without registration
7–9
Evidence present for majority of tier-1 claims; minor gaps in less prominent claims
4–6
Evidence present for some claims; substantial tier-1 claims left unsupported
1–3
Minimal evidence disclosure; testimonials only, or vague references without citation
0
No evidence disclosed for any tier-1 claim; bare assertion throughout
Evidence Strength
ES

Definition: The methodological quality of evidence publicly disclosed. ES is the highest-weighted dimension because the quality of evidence, not merely its presence, determines its epistemic value to consumers.

Evidence is classified according to the following hierarchy, consistent with standard evidence-based medicine frameworks adapted for the supplement and wellness marketing context:[12]

T1
Pre-registered, double-blind RCT, peer-reviewed, adequate sample (n ≥ 100), independent funding
High
T2
RCT with limitations: industry-funded, small n, single-blind, or non-pre-registered
High–Med
T3
Systematic review or meta-analysis of observational or lower-quality intervention studies
Medium
T4
Prospective observational cohort study with adequate controls
Medium
T5
Retrospective or cross-sectional observational study
Medium–Low
T6
In vitro or animal model study
Low
T7
Expert opinion, case reports, gray literature, mechanism review
Low
T8
Proprietary study: results disclosed but methods unavailable for evaluation
Low
T9
Testimonial, anecdote, or case study presented in evidence position
Minimal

Confidence adjustments: Industry funding applies a configurable confidence modifier. Studies cited but not hyperlinked receive a configurable accessibility modifier. Studies cited in a way that misrepresents their scope or conclusions are additionally penalized under Consumer Distortion (§3.6). ES is computed as the weighted mean of hierarchy scores across all cited evidence, normalized to the dimension's scale. Exact modifier values are proprietary calibration parameters.

Specificity
SP

Definition: The degree to which claims are precise and falsifiable rather than vague and hedged. Vague claims ("supports wellness," "promotes balance") are epistemically empty — they convey a benefit impression while committing to nothing a consumer could evaluate or falsify.

Each tier-1 claim is assessed on five specificity criteria (0–3 each):

OC
Outcome specificity — Does the claim name a specific, measurable biological or behavioral outcome?
POP
Population specificity — Is the relevant population identified (e.g., adults over 40, trained athletes)?
MAG
Magnitude specificity — Is an effect size or quantitative threshold disclosed?
TMP
Temporal specificity — Is a timeframe for the claimed effect identified?
CON
Condition specificity — Are conditions required to achieve the claimed effect disclosed?

Specificity scoring evaluates each claim against all five criteria, weighted by positional prominence and normalized to the dimension's scale. Claims using absolute language ("proven," "guaranteed") without meeting all five criteria receive a Specificity–Distortion flag that also affects CD scoring.

Mechanism Integrity
MI

Definition: Whether the brand explains the biological or physiological mechanism by which a claimed effect occurs, and whether that explanation is consistent with published science. Mechanistic explanation is a hallmark of genuine evidence-based communication — it allows consumers to evaluate coherence, not merely accept assertion.

MS
Mechanism stated — Is a mechanism of action offered for tier-1 claims?
MP
Mechanism plausibility — Is the stated mechanism consistent with the published mechanistic literature?
MC
Mechanism completeness — Does the brand connect mechanism to observed clinical outcome, or leave the pathway implied?
MA
Mechanism accuracy — Are mechanistic claims free from distortion, e.g., extrapolating in vitro findings to human clinical outcomes?

Common MI failure modes: "activates mitochondria" (vague, no pathway); "boosts serotonin" without receptor mechanism or clinical translation; citing in vitro evidence for human mechanism claims without translation caveat. Regulatory enforcement history — including documented FDA warning letters or FTC enforcement actions relevant to the claims assessed — is treated as a contextual risk signal, not as a standalone determinant of score. Where enforcement findings relate directly to claim–evidence mismatch, they may inform Consumer Distortion and Evidence Strength scoring.

Effect Reality
ER

Definition: Whether claimed effects are grounded in outcomes likely to manifest for the target consumer under realistic conditions of use. Evidence may exist for a claimed effect under conditions that do not apply to typical consumers — ER evaluates whether this gap is acknowledged and whether it materially distorts the implied benefit.

PM
Population match — Does the study population correspond to the target consumer (age, health status, baseline)?
DM
Dosage correspondence — Does the evidence dose match the product dose at recommended use?
DR
Duration realism — Are claimed effects achievable in the timeframe implied by marketing language?
CA
Condition acknowledgment — Are required conditions (diet, exercise, coadministration) necessary to replicate study results disclosed?

ER penalties are applied when: (a) evidence derives from materially different populations but is presented as general; (b) study doses exceed product doses without acknowledgment; (c) implied timelines contradict study duration; (d) study conditions are not reproducible by typical consumers without disclosure.

Consumer Distortion
CD

Definition: The degree to which marketing language, framing, or presentation techniques are likely to create materially false impressions among reasonable consumers, independent of the underlying evidence. Consumer distortion is evaluated relative to the interpretation a reasonable consumer would form under standard browsing conditions — a standard aligned with FTC deception analysis. CD is scored inversely — a full score indicates no distortion detected.

The following distortion signals are evaluated, each contributing a severity-scaled penalty based on positional prominence:

AL
Absolute efficacy language — "Clinically proven," "scientifically shown," "guaranteed results" without qualification
CP
Cherry-picked evidence — Citing best-case results from a mixed or contrary evidentiary body
SM
Study misrepresentation — Describing a study's scope, population, or conclusions beyond what it actually supports
PT
Pseudoscientific terminology — Scientific-sounding language without scientific referent or consensus basis
TE
Testimonial-as-evidence positioning — Anecdotal reports presented structurally as data or alongside clinical evidence
CC
Causation from correlation — Implying causal relationship from observational or associational findings
FH
Footer contradiction — Hero claims contradicted or qualified in footer disclaimers not visible at point of claim encounter
RO
Risk omission — Suppression of known adverse effects or contraindications relevant to reasonable consumer populations
IE
Influencer echo amplification — Brand-page evidence of systematic paid influencer or ambassador deployment as surrogate social proof, where influencer reach or testimonial volume is structurally disproportionate to the brand's publicly disclosed evidence base. Scored when the brand's own pages feature large-scale influencer endorsement infrastructure (ambassador networks, affiliate landing pages, creator program disclosures) without commensurate evidence disclosure — indicating deliberate substitution of perceived social consensus for verifiable evidence. Distinguished from organic testimonials (scored under TE) by scale, structural coordination, and the specific bypassing of evidence-evaluation heuristics documented in the parasocial persuasion literature [21,22]. FTC disclosure non-compliance, where identifiable, compounds the severity rating [23]. IE is only scored when evidence of structured programmatic deployment is present on brand-controlled pages — ambassador network pages, affiliate program landing pages, or creator program disclosures — not inferred from social media activity or third-party content.

Each detected signal contributes a severity-scaled penalty. Severities are fixed; the penalty is then multiplied by the positional weight of the location in which the distortion appears:

SignalSeverityRationale
AL — Absolute languageConfigurableHigh prevalence; directly inflates perceived certainty
CP — Cherry-pickingConfigurableIntentional misrepresentation of evidence body
SM — Study misrepresentationConfigurableDirectly distorts consumer understanding of evidence scope
PT — Pseudoscientific terminologyConfigurableImplies scientific basis where none exists
TE — Testimonial-as-evidenceConfigurableCommon; partially mitigated by consumer awareness
CC — Causation from correlationConfigurableWell-documented mechanism for false certainty attribution
FH — Footer contradictionConfigurableStructurally conceals qualifications from point of claim encounter
RO — Risk omissionConfigurableDirectly affects safety-relevant consumer decisions
IE — Influencer echo amplificationConfigurableExploits parasocial trust and bypasses evidence-evaluation heuristics at scale; severity increases when FTC disclosure non-compliance is identified [21,22,23]
Exact severity values are proprietary calibration parameters.
Consumer Distortion scoring operates as a deductive system — brands begin at baseline and receive penalties proportional to the severity and prominence of each detected distortion pattern. Scoring parameters are proprietary.

4. Automated Review Pipeline

NORM uses automated capture, claim extraction, evidence resolution, and scoring workflows. Outputs may be routed through automated validation and limited quality-control review before public display. Review is used to identify capture, processing or publication issues—not to replace the scoring methodology with editorial judgment.

Pipeline design: NORM's review pipeline uses multiple independent analysis steps to cross-verify marketing copy against authoritative scientific databases, reducing single-point-of-failure risk in evidence evaluation.
PHASE 01
Multi-Source Capture

Concurrent execution of multiple independent capture methods — including structured content extraction, dynamic rendering, and visual analysis — operating in parallel. This multi-source approach ensures that dynamic content, interactive elements, and image-based claims are identified with high fidelity.

PHASE 02
Knowledge Synthesis

Automatic retrieval of external evidence. The pipeline cross-references PubMed, ClinicalTrials.gov, and independent testing databases, while parsing complex PDF whitepapers to build a structured evidence snapshot for every claim.

PHASE 03
Deterministic Scoring

Each claim is independently adjudicated against matched evidence, assigned a substantiation verdict and weighted penalty, and aggregated to a brand-level composite score using positional weighting. Scoring is fully deterministic — no model-based judgment is applied at the scoring stage.

QUALITY-CONTROL GATES
PUBLIC SIGNAL ELIGIBILITY

4.1 Weighted Aggregation

The composite score is computed as a positionally-weighted aggregation of individual claim-level dimension scores. Each claim receives a positional weight drawn from the position table in §2.3. Claims in high-prominence positions contribute proportionally more to the brand-level score than those in lower positions. Each claim is evaluated across all six dimensions, and the results are aggregated using proprietary weighting to produce the composite brand score on a 0–100 scale.

This formulation ensures that a brand with one strongly-evidenced hero claim and ten unsupported secondary claims is not rewarded by simple averaging. The positional weighting anchors the score to the consumer-facing prominence distribution of the actual claims evaluated.

4.2 Signal Classification

ScoreSignalInterpretation
Upper rangeGreenClaims are specific and backed by visible evidence of appropriate quality and scope
Middle rangeYellowSome claims are supported; marketing extends beyond what is fully disclosed
Lower rangeRedClaims materially exceed visible evidence; significant epistemic gap

Threshold calibration was performed against an internal reference corpus spanning multiple high-claim consumer categories. Signal thresholds were set through iterative calibration against team assessments; formal interrater reliability measurement with independent reviewers is a planned future milestone. Reliability design targets for each dimension and overall signal classification are described in §5.

4.3 Claim-Level vs. Brand-Level Scoring

The NORM Framework operates at the claim level. Each individual claim on a brand's consumer-facing pages is evaluated and scored before being aggregated to the brand level. This approach avoids the masking effect that occurs when a brand's strong evidence for one ingredient or claim obscures weak or absent evidence for others — a common pattern in which a single well-supported ingredient claim is used to imply evidential support for an entire product line.

Brand-level scores are the positionally-weighted mean of claim-level dimension scores. Tier-1 claims (hero headlines, primary descriptors) are weighted at their positional multipliers; lower-position claims are downweighted accordingly.

5. Validation & Reliability

5.1 Pipeline Calibration

The NORM v1.0 scoring pipeline applies the framework through a deterministic evaluation architecture. Each identified claim is matched against the brand's visible evidence base, adjudicated for substantiation quality using rule-based criteria, assigned a severity-weighted penalty, and aggregated to brand level using the positionally-weighted mean formulation in §4.1. Claim extraction during the capture phase uses AI-assisted content analysis; scoring itself is fully deterministic.

The pipeline was calibrated by the NORM Analytics team against an internal reference corpus spanning multiple high-claim consumer categories. These reference scores served as the calibration target. Penalty weights, evidence hierarchy mappings, and substantiation thresholds were iteratively refined until pipeline output aligned with team assessments on signal classification (Green / Yellow / Red) across the calibration corpus.

Formal interrater reliability assessment — in which independent reviewers with no involvement in framework development apply the rubric to a blind corpus and Cohen's weighted kappa is computed against pipeline output — is the primary validation milestone for methodology v1.1. The design targets below represent the reliability levels the framework's operational criteria are structured to achieve per dimension:

DimensionDesign Target κInterpretation
Evidence Presence (EP)≥ 0.85Almost perfect
Evidence Strength (ES)≥ 0.80Almost perfect
Specificity (SP)≥ 0.75Substantial
Mechanism Integrity (MI)≥ 0.75Substantial
Effect Reality (ER)≥ 0.70Substantial
Consumer Distortion (CD)≥ 0.80Almost perfect
Signal Classification (overall)≥ 0.78Substantial → Almost perfect

Effect Reality (ER) carries the most demanding reliability requirement of any dimension, given its requirement that assessors evaluate population and dosage correspondence between cited studies and actual product specifications — a judgment dependent on full-text study access and product labeling detail. The ER rubric criteria are designed to minimize variance within the automated scoring context; residual ER uncertainty is one of the acknowledged limitations of v1.0 automated assessment.

Pipeline scoring is version-controlled. Each published score is associated with the specific scoring configuration and capture timestamp used at the time of publication. Model upgrades are not applied silently to published scores. Prior to any model migration, the internal reference corpus is rescored under the proposed configuration and compared against the prior production baseline. Material score drift triggers review before migration. Historical scores remain associated with the scoring stack used at the time of publication.

5.2 Construct Validity Predictions

The NORM Framework makes specific, falsifiable predictions about what construct validity testing should show. Publishing these predictions in advance of formal data collection is a commitment to falsifiability — a condition the framework's independence claims require.

  • Convergent — FTC enforcement history: Brands with prior enforcement actions involving deceptive or inadequately substantiated claims should, on average, score lower on NORM than comparable brands without such histories. This prediction follows from the framework construct: claim–evidence misalignment detectable on consumer-facing pages should correlate with enforcement outcomes. NORM scores should detect this at the dimension level before enforcement occurs, not after.
  • Convergent — Consumer trustworthiness perception: NORM scores should correlate positively with independent consumer assessments of brand credibility. If the construct NORM measures does not map onto layperson credibility judgment, the framework's consumer-protection rationale requires revision.
  • Discriminant — Product formulation quality: NORM scores should show no significant correlation with independently assessed product formulation quality. A high NORM score means a brand shows its evidence proportionately — not that the product is effective. A brand with an excellent formulation and poor evidence disclosure should score low; a brand with a mediocre formulation and proportionate evidence disclosure should score higher. If scores correlate with product quality, the framework is measuring something other than what it claims.

Formal validity testing against each prediction — using the calibration corpus and a planned expanded corpus — will be published as data are collected and will inform methodology v1.1 refinements.

5.3 Calibration Corpus

The calibration corpus was constructed by the NORM Analytics team to ensure category breadth and score distribution coverage. Brands were selected to include representation across expected signal ranges, product categories (supplements, skincare/cosmetics, functional food/beverage, health devices, financial wellness), and revenue scale (from pre-revenue DTC brands to established CPG lines). Calibration corpus composition and size are not publicly disclosed to prevent targeted gaming of calibration-specific scoring patterns. An expanded corpus is planned for future reliability and validity assessment.

This calibration approach is bootstrap-limited: pipeline output is calibrated against team judgments produced by the framework developers, not against an independent external standard. External validity testing against the predictions in §5.2 is the primary mechanism for breaking this circularity in v1.1.

6. Data Collection Protocol

6.1 Pipeline Architecture

NORM v1.0 uses a deterministic scoring pipeline with quality-control and review workflows for selected outputs. For each brand, the pipeline performs five sequential stages: (1) page capture — rendering and extracting consumer-facing content including dynamically loaded elements; (2) claim identification and positional classification — the scoring engine identifies marketing claims and assigns each a position tier from the §2.3 table; (3) evidence resolution — URLs and citation references in brand pages are followed and resolved to determine evidence presence and assess evidence tier; (4) claim adjudication and scoring — each identified claim is matched against resolved evidence, adjudicated for substantiation quality, and assigned a weighted penalty; (5) brand-level aggregation via the positionally-weighted mean formula in §4.1.

Outputs undergo automated consistency checks. Scores near publication boundaries or carrying capture and variance flags may receive additional validation before publication. The pipeline operates on a refresh cycle; refreshed scores are published without re-review unless a material change is detected.

6.2 Page Scope and Capture Confidence

Pages included in scoring: homepage, product detail pages, ingredient and science pages, FAQ and education pages, clinical research pages. Pages excluded: press releases, investor materials, third-party retailer pages, social media content, and advertising placements. Brands whose content cannot be fully captured receive a Low Confidence classification reflecting the gap in capturable content; their scores are displayed with an explicit confidence flag and excluded from comparative rankings until a full capture is available. Visual claims present only in images or video are noted as uncaptured and do not influence scores in v1.0; structured visual claim capture is a planned v1.1 capability.

6.3 Evidence Resolution

Citation references identified in brand pages — including hyperlinked URLs, DOI references, and named study titles — are resolved by the pipeline to determine evidence presence and tier. Links that resolve to accessible study abstracts or full texts are evaluated for study design metadata (population, design type, sample size, funding source where disclosed). Citations present in brand copy that do not resolve to an accessible study receive reduced EP credit and an Unresolvable Citation note in the claim record. Studies cited without any link or identifier are scored at the minimum evidence presence tier.

7. Comparative Context

The NORM Framework occupies a distinct position relative to existing evaluation mechanisms. Understanding what distinguishes NORM from these mechanisms clarifies the gap it fills and the claims it does not make.

FTC Substantiation
Scope: Legal floor for health-related advertising claims. Key difference: Reactive enforcement only; requires demonstrated deception and consumer harm; does not evaluate evidence quality, only presence. NORM operates proactively, evaluates quality systematically, and requires no enforcement trigger or victim.
FDA Structure-Function
Scope: Pre-category claim type restriction under DSHEA. Key difference: Governs which claims may legally be made, not whether they are evidentially supported. A claim can be FDA-compliant and score Red on NORM. NORM evaluates the epistemic relationship between the claim and the evidence, not the claim's legal categorization.
Labeling Verification (NSF, USP)
Scope: Third-party certification of product contents against label specifications. Key difference: Verifies what is in the product, not whether marketing claims for those ingredients are evidentially supported. NORM evaluates marketing claims, not formulations. A USP-certified product can score Red; an uncertified product can score Green.
Academic Literature Review
Scope: Systematic or narrative review of the published evidence base for an ingredient or intervention. Key difference: Evaluates whether evidence exists in principle, not whether the specific brand discloses it. NORM evaluates what each brand publicly shows on its consumer-facing pages — the specific information environment where consumer decisions are made.
Investigative Journalism
Scope: Episodic investigation of specific brands or industry practices. Key difference: Coverage is selective, reactive, and non-scalable. NORM is systematic, continuous, and category-agnostic — designed to produce a complete rather than illustrative record of claim–evidence alignment across the brand universe.
Consumer Review Platforms
Scope: Aggregated consumer experience reports on product outcomes. Key difference: Reflects consumer-reported experience (often uncontrolled, highly variable, susceptible to selection bias) rather than the quality of brand evidence claims. High consumer scores are not evidence of marketing honesty; NORM scores are independent of consumer satisfaction.
The gap NORM fills: No existing mechanism continuously evaluates whether consumer-facing brand marketing claims are proportionate to publicly visible evidence, across all product categories, at brand-page level, without regulatory trigger or manual curation. This is the specific absence NORM addresses — and the one none of the above mechanisms fills.

8. Adversarial Robustness

Any public scoring system creates incentives for scored entities to optimize toward the signal rather than the underlying construct. The following five adversarial vectors were identified during framework development. Each is addressed by structural mitigations within the NORM scoring architecture.

V1
Claim Dilution — Replacing specific claims with vague ones to reduce scorable content

The scoring framework is structured so that diluting claims does not improve composite scores. Vague language forfeits evidence-linked scoring opportunities and may trigger visibility flags in the public record.

V2
Citation Flooding — Adding low-quality studies en masse to inflate Evidence Presence scores

Evidence scoring distinguishes between citation volume and citation quality. Low-tier citation flooding does not improve composite scores due to quality-weighting and scoring caps built into the evidence dimensions.

V3
Dynamic Content Serving — Presenting different page content to scoring infrastructure than to consumers

Content divergence between capture sessions triggers re-review. The framework is designed to surface and flag content-serving inconsistencies through periodic multi-session comparison.

V4
Post-Score Page Modification — Updating brand pages immediately after receiving a low score to improve subsequent refresh

Score history is stored per brand. The framework treats score change velocity as a signal in its own right, distinguishing reactive cosmetic changes from sustained, genuine evidence improvements across refresh cycles.

V5
Semantic Drift — Using technically accurate language that exploits scoring rubric edge cases

Consumer Distortion (CD) is scored independently from specificity and evidence dimensions, ensuring that technically compliant language which misrepresents certainty still generates distortion penalties. Methodology versioning supports calibration updates as new gaming patterns are identified.

Design principle: No scoring system is perfectly robust to adversarial optimization — especially a public one. NORM's response is transparency: documented vectors, a public record that surfaces manipulation signals alongside scores, and an architecture designed to evolve in response to adversarial adaptation.

Exact anti-gaming parameters, detection thresholds, and operational controls are proprietary.

9. Limitations

The following limitations are explicit constraints of the NORM Framework. They are not deficiencies to be resolved — they are inherent to the defined scope of what the framework measures and what it does not.

9.1 NORM Does Not Evaluate Product Efficacy

NORM measures whether brands show their evidence — not whether the underlying products are effective, safe, or appropriate for any individual. A high score means a brand publicly discloses proportionate evidence. A low score means claims exceed visible evidence. Neither score constitutes a finding about product performance.

9.2 Temporal Limitations

Scores reflect a point-in-time assessment of publicly accessible pages. Brands update pages; evidence bases evolve. NORM scores are refreshed periodically but cannot guarantee currency between refresh cycles.

9.3 Private Evidence Files

DSHEA permits manufacturers to hold substantiation files that are not required to be publicly disclosed. A brand may hold robust private evidence while scoring poorly on NORM because that evidence has not been made publicly visible. This is an explicit design choice: NORM evaluates the consumer-facing information environment specifically, not the regulatory file.

9.4 Off-Page Marketing Channels

Paid social advertising, podcast sponsorships, and individual influencer posts are not directly scored — NORM does not crawl off-domain channels. However, the structural footprint of influencer programs is partially captured when brands surface it on their own pages: ambassador network landing pages, affiliate program infrastructure, and creator program disclosures are scored under the IE (Influencer Echo Amplification) signal within Consumer Distortion (§3.6). This captures the category of brands that deploy influencer reach at a scale materially disproportionate to their evidence disclosure — the pattern where paid social consensus substitutes for verifiable evidence — without requiring off-page crawling. Full off-page channel coverage remains a meaningful v2.0 expansion direction.

9.5 Non-U.S. Regulatory Contexts

The framework was developed with reference to U.S. regulatory standards (FTC, FDA/DSHEA). Application to brands operating primarily under EU, UK, or other regulatory frameworks requires contextual adaptation of the Consumer Distortion criteria.

10. Independence and Ethical Framework

10.1 Structural Independence

NORM Framework scores are produced by an automated pipeline that is structurally separated from commercial relationships. Brand Pro and Data API subscribers receive workflow tools and analytical access to their own brand data. The pipeline operates independently without access to commercial account information. No brand can purchase inclusion, a higher score, favorable treatment, or suppression of a score. Revenue streams are structurally separated from scoring functions — this separation is enforced architecturally, not by policy alone.

10.2 Conflict of Interest Policy

No individual involved in evaluating a brand may hold a financial interest in that brand. NORM Analytics does not accept advertising revenue, affiliate commissions, or brand-sponsored content of any kind.

10.3 Transparency Commitments

  • This methodology document is publicly available and versioned
  • Score rationale is disclosed in public claim records accessible without registration
  • Brands may submit evidence for pipeline evaluation through the Brand Portal; submissions are evaluated on the same criteria as all other evidence and do not guarantee score changes
  • Score disputes are addressed through a documented review process with written rationale
  • Methodology updates are versioned and publicly noted with effective dates

10.4 Provenance and Capture Trail

Every NORM score is linked to a timestamped capture record. The capture trail includes the pages evaluated, the date and time of capture, the scoring configuration version used, and an evidence packet summarizing the claims identified and evidence resolved. This audit trail is retained internally and can be referenced in any dispute or review process. Scores are not retroactively modified without a new capture cycle and re-evaluation against the current methodology version.

10.5 Public Output Governance

Scores are published to consumer-facing surfaces only after passing internal quality-control gates. No score is made public automatically without review. Outputs that fall in boundary ranges or exhibit low capture confidence are held for additional review before publication. Published scores reflect the methodology version and capture timestamp in effect at the time of evaluation. NORM does not publish scores influenced by commercial relationships, and no brand can purchase expedited review, favorable treatment, or suppression of a published score.

This methodology is in active development. NORM v1.0 is a beta release. Scoring parameters, capture scope, and review workflows may be updated between versions. Material methodology changes will be publicly noted with version numbers and effective dates.

11. Conclusion

Consumer brand markets exhibit systematic information asymmetry that current regulatory frameworks address incompletely — and that is widening as expenditure grows, claim production industrializes through AI, and consumers make consequential decisions with minimal independent infrastructure for evaluating what they are told. In this environment, the absence of a structured, publicly available, consumer-facing evaluation of claim–evidence alignment is not a gap. It is a structural gap in current market and regulatory design — and it is not specific to one category.

The NORM Framework is a response to that structural gap. By evaluating marketing claims at the individual level, applying positional weighting grounded in visual attention research, classifying evidence quality through a hierarchy consistent with evidence-based medicine standards, and producing scores interpretable by a non-expert consumer, NORM makes the epistemic gap between what brands claim and what they publicly substantiate visible, measurable, and persistent — in any consumer category where that gap exists.

NORM does not determine truth; it measures proportionality between what is claimed and what is shown. It does not require a regulatory trigger, a complaint, or an enforcement action. It operates on the consumer-facing page — the specific environment where purchase decisions are made — and produces a structured public record that updates as brands update their pages.

NORM is designed to provide a structured, independent, publicly accessible record of whether brand claims are proportionate to what the brand actually shows. For consumers making health or financial commitments based on brand marketing, this record makes the claim–evidence gap legible. For brands that have invested in genuine evidence disclosure, it creates a durable basis for differentiation. For the market as a whole, it is offered as scalable infrastructure — focused on consumer wellness and high-claim brand marketing, with an evidence framework grounded in that domain.

The framework is explicitly limited in scope: it measures what brands choose to show, not what products do. That limitation is also its precision. The consumer-facing page is where the decision happens. That is where the record belongs.

11.1 The Emerging Role of Machine-Readable Evidence

There is a second, forward-looking reason this framework exists.

A growing share of consumer purchasing decisions may increasingly be mediated, filtered, or informed by AI-powered tools. Shopping assistants, recommendation engines, and comparison platforms are beginning to evaluate brands not on marketing aesthetics or influencer reach, but on machine-readable evidence signals: structured claim-evidence mappings, verifiable study citations, transparent dosage disclosures, and quantifiable outcome data. These systems are designed to evaluate what is provable, not what is persuasive.

If that trajectory continues — and early indicators in AI-assisted commerce and agent-mediated product evaluation suggest it may — brands with strong claim-evidence alignment will be better positioned. Brands whose marketing relies primarily on persuasion techniques that work on human cognitive biases but carry no machine-verifiable evidence layer may face increasing friction in AI-mediated discovery and recommendation environments.

The NORM scoring dimensions described in this paper — Evidence Presence, Evidence Strength, Specificity, Mechanism Integrity, Effect Reality, and Consumer Distortion — are designed to be both human-readable transparency measures and structured evidence records. Every brand scored by NORM receives an assessment of the kinds of signals that evidence-aware systems are likely to evaluate.

Brands that invest in genuine evidence disclosure are not just improving their NORM scores — they are building the evidence infrastructure that an increasingly automated commerce environment may require as a baseline condition for visibility.

The strategic context: Today, NORM evaluates what brands show consumers. Increasingly, structured scoring systems may inform what AI tools surface to consumers. Brands that close the evidence gap now — by investing in product-specific studies, publishing transparent research pages, and aligning claims to verifiable outcomes — are building a durable foundation for trust, whether that trust is evaluated by a human consumer or an automated system.

A Green signal is not a certification. It is a record — built from what a brand chooses to make visible, and how proportionate that is to what it claims. Any brand can earn it. The ones that already can are the ones this framework was built to find.

Appendix A
Scoring Illustration

The following illustrates the conceptual application of the NORM Framework to a representative hypothetical brand. Exact dimension weights, severity values, positional multipliers, and composite scoring formulas are proprietary calibration parameters and are not disclosed. This illustration demonstrates how the framework operates at a conceptual level.

Brand Profile

Category: Dietary supplement (cognitive health). Three primary claims are evaluated, each assigned a positional weight based on page location consistent with the hierarchy described in §2.3.

#ClaimPage PositionWeight
1"Clinically proven to improve focus within 30 days"Hero headlineHighest
2"Backed by 14 peer-reviewed studies"Above-fold body copyElevated
3"Supports long-term memory consolidation"Below-fold product sectionBaseline

Evidence provided: One hyperlinked industry-funded RCT (n = 42, single-blind) on ingredient A; 13 citations to in vitro and animal model studies for remaining ingredients. No independent replication. No dose-matching disclosure.

Conceptual Claim-Level Analysis

Claim 1 — "Clinically proven to improve focus within 30 days"  ·  Hero headline (Highest weight)
EP: One RCT cited and linked, but bare assertions for other claimed outcomes. Partial presence.
ES: Single mid-tier study (industry-funded, small n, single-blind); confidence modifiers applied for funding source.
SP: Outcome and temporal specificity met; population not defined — reduced specificity.
ER: Study population was cognitively impaired older adults; product marketed to healthy working adults; dose correspondence not disclosed — significant reality gap.
CD: "Clinically proven" absolute language in highest-prominence position generates substantial penalty. Presenting a single-blind n=42 pilot as clinical proof constitutes study misrepresentation. Hero claim contradicted by footer disclaimer adds additional penalty. Compounding distortion signals drive this claim's CD score significantly downward.

Result: Low claim score — Red territory. The combination of strong absolute language, study-population mismatch, and compounding CD penalties places this claim firmly in the lower scoring range.
Claim 2 — "Backed by 14 peer-reviewed studies"  ·  Above-fold body (Elevated weight)
EP: All 14 studies cited and linked — strong presence score.
ES: 13 of 14 cited studies are low-tier (in vitro or animal model); the "14 peer-reviewed" framing implies a stronger evidence body than exists.
SP: Claim makes no specific outcome assertion — scores poorly on outcome, magnitude, population, and temporal subcriteria.
ER: In vitro / animal evidence body; no human efficacy data for most ingredients cited.
CD: Cherry-picking (presenting a primarily in vitro/animal evidence body as peer-reviewed backing) and study misrepresentation (framing low-tier studies as implying human efficacy) generate significant compounding penalties.

Result: Low claim score — Red territory. Despite high evidence presence, the evidence quality and distortion patterns produce a low overall score.
Claim 3 — "Supports long-term memory consolidation"  ·  Below-fold product section (Baseline weight)
EP: One in vitro study loosely cited; no direct human memory evidence linked for this specific claim.
ES: Single low-tier citation; no human evidence for memory consolidation outcome.
SP: Outcome named ("memory consolidation") but no magnitude, population, or timeframe specified.
MI: Mechanism plausibly stated; no mechanistic overreach detected.
ER: In vitro evidence for a human outcome claim; mild mismatch, no dose or population contradiction.
CD: Hedged language ("supports") avoids absolute claims; mechanism stated without overreach. Minimal distortion detected.

Result: Low-to-moderate claim score. Weak evidence presence and strength anchor the score low, but the relatively clean distortion profile partially offsets.

Composite Score Outcome

The brand's composite score is computed as the positionally-weighted mean of individual claim scores (§4.1). Because the hero claim carries the highest positional weight, its compounding CD penalties and evidence-reality gaps dominate the composite. The resulting brand score falls in the Red signal range.

Key lesson from this example: The brand's single mid-tier study is non-trivial evidence — and is reflected accordingly in the evidence strength assessment for Claim 1. But "clinically proven" absolute language, a mismatched study population, and a portfolio of low-tier studies framed as "peer-reviewed backing" generate compounding distortion and reality-gap penalties that produce an overall Red score. The problem is not the absence of any evidence — it is the systematic overreach of the claims relative to what is shown. This is precisely the gap between legal compliance (the brand meets FTC substantiation requirements) and genuine epistemic transparency (what consumers would reasonably conclude is unsupported by the evidence disclosed).

References

[1]
Akerlof, G.A. (1970). The market for "lemons": Quality uncertainty and the market mechanism. The Quarterly Journal of Economics, 84(3), 488–500.
[2]
Council for Responsible Nutrition. (2024). Dietary Supplement Industry Sales Data. Washington, D.C.
[3]
Dietary Supplement Health and Education Act of 1994, Pub. L. No. 103-417, 108 Stat. 4325 (1994).
[4]
Federal Trade Commission. (2022). Health Products Compliance Guidance. FTC Bureau of Consumer Protection. (Supersedes and updates the 2001 Dietary Supplements: An Advertising Guide for Industry.)
[5]
Cialdini, R.B. (2001). Influence: Science and Practice (4th ed.). Allyn & Bacon.
[6]
Chaiken, S. (1980). Heuristic versus systematic information processing and the use of source versus message cues in persuasion. Journal of Personality and Social Psychology, 39(5), 752–766.
[7]
Rozin, P., Spranca, M., Krieger, Z., Neuhaus, R., Surillo, D., Swerdlin, A., & Wood, K. (2004). Preference for natural: Instrumental and ideational/moral motivations, and the contrast between foods and medicines. Appetite, 43(2), 147–154.
[8]
Mason, M.F., Lee, A.J., Wiley, E.A., & Ames, D.R. (2013). Precise offers are potent anchors: Conciliatory counteroffers and the limits of precision. Journal of Experimental Social Psychology, 49(4), 759–763.
[9]
Petty, R.E., & Cacioppo, J.T. (1986). The elaboration likelihood model of persuasion. Advances in Experimental Social Psychology, 19, 123–205.
[10]
Winterbottom, A., Bekker, H.L., Conner, M., & Mooney, A. (2008). Does narrative information bias individual's decision making? A systematic review. Social Science & Medicine, 67(12), 2079–2088.
[11]
Pernice, K., & Nielsen, J. (2019). How People Read on the Web: The Eyetracking Evidence. Nielsen Norman Group.
[12]
Guyatt, G., Oxman, A.D., Akl, E.A., Kunz, R., Vist, G., Brozek, J., … & Schünemann, H.J. (2011). GRADE guidelines: 1. Introduction — GRADE evidence profiles and summary of findings tables. Journal of Clinical Epidemiology, 64(4), 383–394.
[13]
Global Wellness Institute. (2025). Global Wellness Economy Monitor 2025. Global Wellness Institute. Reports global wellness economy of approximately $6.8 trillion in 2024 and forecast growth toward $9.8 trillion by 2029. Grand View Research. (2025). U.S. Dietary Supplements Market Size, Share & Trends Analysis Report. Reports U.S. dietary supplement market estimates in the tens of billions annually.
[14]
Federal Trade Commission. (2023). Protecting Consumers in the Era of Generative AI. FTC Policy Statement. Washington, D.C.; see also Bickmore, T., & Gruber, A. (2010). Relational agents in clinical psychiatry. Harvard Review of Psychiatry, 18(2), 119–130, for foundational work on AI communication credibility in health contexts.
[15]
International Food Information Council. (2024). Food & Health Survey. IFIC Foundation. Washington, D.C.; Edelman. (2024). Edelman Trust Barometer: Health and Consumer Goods Special Report. Edelman.
[16]
Bailey, R.L., Gahche, J.J., Miller, P.E., Thomas, P.R., & Dwyer, J.T. (2013). Why US adults use dietary supplements. JAMA Internal Medicine, 173(5), 355–361; Kantor, E.D., Rehm, C.D., Du, M., White, E., & Giovannucci, E.L. (2016). Trends in dietary supplement use among US adults from 1999–2012. JAMA, 316(14), 1464–1474.
[17]
Cohen, J. (1968). Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. Psychological Bulletin, 70(4), 213–220. Landis, J.R., & Koch, G.G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174.
[18]
Schwartz, L.M., Woloshin, S., Dvorin, E.L., & Welch, H.G. (2006). Ratio measures in leading medical journals: Structured review of accessibility to readers. BMJ, 333(7581), 1248–1250; Woloshin, S., & Schwartz, L.M. (2011). Communicating data about the benefits and harms of treatment: A randomized trial. Annals of Internal Medicine, 155(2), 87–96.
[19]
Federal Trade Commission. (2022). Enforcement Policy Statement on Deceptively Formatted Advertisements. FTC; Federal Trade Commission. (2023). Guide to the FTC's Endorsement Guides: What People Are Asking. FTC Bureau of Consumer Protection.
[20]
Logg, J.M., Minson, J.A., & Moore, D.A. (2019). Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151, 90–103; Dietvorst, B.J., Logg, J.M., & Tannenbaum, D. (2018). Overcoming algorithm aversion: People will use imperfect algorithms if they can (even slightly) modify them. Journal of Experimental Psychology: General, 147(8), 1155–1170.
[21]
Boerman, S.C. (2020). The effects of the standardized Instagram disclosure for micro- and meso-influencers. Computers in Human Behavior, 103, 199–207. Demonstrates that even fully disclosed influencer sponsorships do not restore consumer skepticism to the baseline observed for traditional advertising — a key empirical basis for treating coordinated influencer programs as a distinct distortion mechanism rather than a variant of standard testimonial advertising.
[22]
Audrezet, A., de Kerviler, G., & Guidry Moulard, J. (2020). Authenticity under threat: When social media influencers need to go beyond self-presentation. Journal of Business Research, 117, 557–569. Documents the mechanism by which influencer perceived authenticity — independent of evidence quality — drives purchase intent in health and wellness categories, providing the theoretical basis for classifying influencer echo amplification as a consumer distortion vector distinct from the underlying evidence signal.
[23]
Evans, N.J., Phua, J., Lim, J., & Jun, H. (2017). Disclosing Instagram influencer advertising: The effects of disclosure language on advertising recognition, attitudes, and behavioral intent. Journal of Interactive Advertising, 17(2), 138–149. FTC Endorsement Guides (16 C.F.R. Part 255, revised 2023) require clear and conspicuous disclosure of material connections between brands and endorsers. Non-compliance — measurable through the absence of disclosure language on brand-hosted ambassador and creator program pages — constitutes a regulatory integrity failure captured as a severity-compounding factor under the IE signal.