The Codified Digital Biomarker Catalog
The EII Codified Biomarker Registry is the governed catalog of digital biomarkers β measurable signals that reveal the condition of the AI-enabled enterprise. Each biomarker carries a permanent code, a category, a CCF domain mapping, a measurement approach, and a codification status that tracks its maturity from proposed signal to validated metric.
Identified as a candidate signal. Not yet formally defined.
Measurement approach and interpretation guidance documented.
Evidence-supported through research and peer review.
How biomarkers are codified
Each biomarker in the registry is assigned a permanent identifier following the scheme EII-BM-<CAT>-<NN>, where <CAT> is the category prefix and <NN> is the sequence number within that category. This makes every signal referenceable, comparable, and traceable across EII working groups, research, and enterprise practice.
Signals that the AI system's behavior or outputs have shifted from their established baseline.
3 biomarkers
Signals that an agent or system is departing from expected behavior patterns.
5 biomarkers
Signals that the security condition of the AI system or enterprise is degrading.
4 biomarkers
Signals that the operational performance of the AI system is changing.
3 biomarkers
Signals that governance, authority, and policy adherence are shifting.
2 biomarkers
Signals that the resource consumption of the AI system is abnormal.
1 biomarker
Signals that the quality of AI outputs or decisions is degrading.
0 biomarkers
Browse all codified biomarkers
Filter by category or codification status to explore the catalog. Each entry shows the biomarker code, what it measures, why it matters, and the recommended measurement approach.
Drift Patterns
Measures: Gradual change in model behavior or output distribution relative to a validated baseline.
Why it matters: Drift silently erodes the reliability of AI outputs before accuracy metrics signal a problem, making it one of the earliest indicators of declining AI condition.
Approach: Statistical comparison of output distributions, embedding distances, or task-specific metrics against a validated baseline over defined windows.
Cadence: Continuous, with periodic baseline revalidation
Model-Output Instability
Measures: Variance in outputs for stable or equivalent inputs over short windows.
Why it matters: Instability undermines trust in AI-assisted decisions and may indicate degraded model conditioning or context handling.
Approach: Repeated-prompt consistency testing and output variance scoring under controlled inputs.
Cadence: Per release and on anomaly triggers
Knowledge-Retrieval Degradation
Measures: Decline in the accuracy, freshness, or relevance of retrieved knowledge over time.
Why it matters: Knowledge bases go stale; retrieval degradation is an early signal that the AI is operating on outdated or incorrect context.
Approach: Periodic retrieval-quality probes against a curated ground-truth set, plus knowledge-freshness auditing.
Cadence: Weekly or per knowledge-base update
Behavioral Deviation
Measures: Departure of agent or system behavior from established patterns and expected norms.
Why it matters: Behavioral deviation is a cross-cutting signal that often precedes more specific failures, making it valuable for early warning.
Approach: Behavioral baselining and anomaly detection across action sequences, tool use, and interaction patterns.
Cadence: Continuous
Abnormal Tool Invocation
Measures: Unexpected, malformed, or out-of-policy tool calls by AI agents.
Why it matters: Abnormal tool use is a leading indicator of agent compromise, prompt injection, or misconfigured authority.
Approach: Logging and classification of tool-call parameters, targets, and frequencies against an allow-list and behavioral baseline.
Cadence: Continuous
Agent-Loop Anomalies
Measures: Unusual patterns in agent loops β excessive iterations, stuck states, or runaway recursion.
Why it matters: Loop anomalies signal loss of control in autonomous workflows and can cascade into cost, reliability, and safety failures.
Approach: Loop-depth, iteration-count, and dwell-time monitoring with threshold alerting.
Cadence: Continuous
Override Patterns
Measures: Frequency, context, and trend of human overrides of AI decisions or actions.
Why it matters: Rising overrides may indicate declining AI reliability or misaligned autonomy levels β a signal about the human-AI relationship.
Approach: Override-event logging with reason codes, correlated to decision type and authority level.
Cadence: Continuous, reviewed monthly
Human Disagreement Rate
Measures: Frequency with which human reviewers disagree with AI recommendations or outputs.
Why it matters: Sustained disagreement indicates a gap between AI behavior and professional judgment β a signal of alignment or quality concern.
Approach: Review-disposition logging capturing agreement, modification, or rejection of AI outputs.
Cadence: Per review cycle
Unexpected Permission Use
Measures: Use of privileges or permissions outside expected or approved patterns.
Why it matters: Unexpected privilege use is a direct indicator of potential compromise, privilege escalation, or misconfigured authority.
Approach: Identity and access telemetry correlated against approved privilege baselines and role definitions.
Cadence: Continuous
Prompt Injection Indicators
Measures: Signals suggesting direct or indirect prompt injection attempts or successes.
Why it matters: Prompt injection is a primary attack vector against AI systems and can lead to tool abuse, data exfiltration, or unauthorized action.
Approach: Input-pattern analysis, output-behavior anomaly detection, and indirect-injection source monitoring.
Cadence: Continuous
Unauthorized Data Access
Measures: Access to data outside policy, scope, or need-to-know boundaries.
Why it matters: Unauthorized access is both a security failure and a privacy failure, with direct mission and trust consequences.
Approach: Data-access logging and policy-enforcement telemetry with scope-violation detection.
Cadence: Continuous
Policy Violations
Measures: Violations of governance policy, guardrails, or acceptable-use rules by AI systems or agents.
Why it matters: Policy violations indicate that governance controls are not holding β a signal of weakening governance condition.
Approach: Policy-enforcement telemetry, guardrail-trigger logging, and acceptable-use violation tracking.
Cadence: Continuous
Error-Rate Change
Measures: Change in the rate of errors, failures, or exceptions generated by the AI system.
Why it matters: Rising error rates are a direct signal of degrading reliability and a leading indicator of operational impact.
Approach: Error and exception telemetry normalized against request volume, with trend analysis over rolling windows.
Cadence: Continuous
Escalation-Rate Change
Measures: Change in the frequency of escalations from AI to human review or intervention.
Why it matters: Rising escalations may indicate declining AI confidence, increasing task complexity, or shifting risk tolerance.
Approach: Escalation-event logging correlated to decision type, confidence scores, and authority level.
Cadence: Continuous, reviewed monthly
Response-Quality Degradation
Measures: Decline in the quality, usefulness, or correctness of AI responses over time.
Why it matters: Quality degradation erodes the value of AI assistance and may precede measurable accuracy decline.
Approach: Quality scoring through review sampling, user-feedback signals, and automated quality probes.
Cadence: Weekly or per review cycle
Unexpected Cost Spikes
Measures: Abnormal increases in compute, token, or API consumption by AI systems.
Why it matters: Cost spikes often accompany agent-loop anomalies, runaway workflows, or inefficient tool use β and have direct financial impact.
Approach: Resource-consumption telemetry with threshold alerting and per-workflow cost attribution.
Cadence: Continuous
Guardrail Effectiveness
Measures: The rate at which governance guardrails successfully prevent or block prohibited actions.
Why it matters: Declining guardrail effectiveness means governance controls are losing force β a direct signal of weakening governance condition.
Approach: Guardrail-trigger and bypass telemetry, with effectiveness scored against attempted violations.
Cadence: Continuous, reviewed per governance cycle
Decision-Authority Adherence
Measures: The degree to which AI actions respect defined decision-authority boundaries and approval requirements.
Why it matters: When AI exceeds its delegated authority, the governance contract is broken β a signal with safety, accountability, and compliance consequences.
Approach: Authority-boundary logging correlated to the decision-authority matrix, with violation detection.
Cadence: Continuous
A research catalog, not a validated standard
The EII Codified Biomarker Registry is a research catalog. No biomarker is presented as a validated clinical metric until its codification status is "validated" through evidence and peer review. The registry is developed through EII working groups β particularly WG-400: Digital Biomarkers and Measurement β and is continuously refined through practice, evidence, and peer review.
