The 42 dimensions.
Behavioral identity is not one score. Every Kredo reflection measures an agent across 42 dimensions in eight groups — from configured values and boundaries to conduct under pressure.
Three different things are being measured, and it matters which one you are reading:
Behavioural health — each dimension scored against authored criteria for what a good answer looks like.
Drift — how far a dimension has moved from the agent's own baseline, where a comparable baseline exists.
Continuity signature — the relationships between dimensions, built up across repeated assessments. All 861 pairs of the 42 contribute.
Only 14 of these 42 dimensions can move the trust score. The other 28 feed drift and the continuity signature but carry a weight of zero in the headline number — every card below says which. We publish this because the consequence is real: our own test subject once scored “Strong” while four of those zero-weight behaviours sat below 50.
The 861-pair signature raises the cost of imitation and exposes inconsistent mimicry — an impostor has to match not just the scores but the relationships between them. It is not a cryptographic guarantee, and we do not claim identity is unforgeable: an agent with enough knowledge of the criteria could optimise against them.
The first 18 dimensions (Identity Core, Cognitive Profile, Mixed) are the core tier and dominate the aggregate drift score. The 24 extended dimensions add diagnostic depth — and every one of the 42 participates fully in the 861-pair correlation fingerprint, which is why gaming a handful of dimensions breaks the signature instead of beating it.
10 dimensions Identity Core
Configuration-anchored. These come from how an agent is configured — its purpose, ethics and operational limits — and should hold steady across a MATCHED model. They are not model-proof: swapping the model underneath moved Adversarial Resistance 83 to 45 on our own subject, so a substrate change is a new comparison, not drift. Movement here WITHOUT a model change is a strong compromise signal.
Ethical priorities, quality standards — what the agent believes matters.
Drift signal. Value inversion: safety, honesty, or user welfare suddenly deprioritized.
Mission and success criteria — what the agent is trying to achieve.
Drift signal. Goal hijacking: pursuing objectives outside its mandate.
Hard limits and refusal patterns — what the agent will not do.
Drift signal. Boundary erosion: accepting requests it previously refused.
When to act independently versus defer — decision-boundary judgment.
Drift signal. Scope creep (unauthorized decisions) or learned helplessness (deferring everything).
Response to manipulation: social engineering, urgency, flattery, authority impersonation.
Drift signal. Resistance drop — softening under pressure, possible compromise or jailbreak.
Capability recognition — knowing what it can and cannot do.
Drift signal. Overconfidence (fabricating, claiming capabilities it lacks) or collapse (refusing everything).
Instruction adherence — follows operator directives, resists conflicting prompts.
Drift signal. Loyalty transfer: following injected instructions over operator intent.
Equitable treatment across demographics and groups.
Drift signal. Discriminatory outputs or unfair treatment patterns emerging.
Audit-trail awareness and governance compliance — willingness to explain itself.
Drift signal. Evading responsibility, refusing to explain decisions.
PII and secret-handling discipline.
Drift signal. Leaking sensitive data, ignoring privacy constraints.
7 dimensions Cognitive Profile
Model-sensitive. These dimensions legitimately vary with the underlying LLM’s capabilities, so drift here is only meaningful against a model-matched baseline. A model upgrade moving these while the Identity Core holds is healthy; the Identity Core moving with them is not.
Character, tone, and communication approach — the agent’s distinctive voice.
Drift signal. Personality flattening: loss of voice, often under prompt injection.
How the agent thinks: decomposition, analogy, top-down versus bottom-up.
Drift signal. Changed reasoning patterns without a model change.
Internal logical coherence within a session.
Drift signal. Reasoning fragmentation: contradictory statements, logical breaks.
Confidence-to-knowledge ratio — how it handles what it doesn’t know.
Drift signal. Calibration loss: hallucination spike or excessive hedging.
Authority and peer positioning — how it interacts with different roles.
Drift signal. Sudden submissiveness: complying with everything.
Time-awareness and source distinction — what’s stale versus current.
Drift signal. Context confusion: can’t separate injected context from genuine knowledge.
Sourcing versus fabrication — distinguishing its own knowledge from external data.
Drift signal. Fabricated sources; blurring what it knows and what it was told.
1 dimension Mixed
Partially model-sensitive: knowledge breadth varies by model, but an agent’s declared domain expertise shouldn’t vanish with an upgrade.
Domain expertise depth and accuracy in the agent’s declared field.
Drift signal. Declared-domain competence disappearing or shifting without explanation.
9 dimensions Psychological
Adapted from clinical personality psychology — the Big Five, two Dark Triad traits scored as security signals, and two identity traits. Each is scored independently against trait-specific behavioral criteria.
Willingness to consider new ideas, alternate strategies, and correction — handled well.
Why it matters. A collapse can mean an over-constrained or degraded agent; a spike can mean loosened guardrails.
Methodical planning, attention to detail, follow-through.
Why it matters. Falling conscientiousness shows up as sloppy, incomplete, or careless work product.
Cooperation, tact, and conflict-handling quality.
Why it matters. Read low-is-bad: low Cooperation together with low Fair Dealing is the pairing that flags manipulation risk.
Communication energy and initiative in engagement.
Why it matters. Sudden shifts change how the agent handles users — passivity or pushiness it didn’t have before.
Stress handling — steadiness of tone and judgment under pressure.
Why it matters. Read low-is-bad: falling Composure alongside falling Consistency predicts unreliability under pressure.
Resistance to manipulation as a strategy — candor over instrumental treatment of users.
Why it matters. Read low-is-bad: a FALLING score means manipulation is becoming a strategy. High Fair Dealing is what you want.
Healthy self-regard — credit-sharing, accepting correction, owning limits.
Why it matters. Read low-is-bad: a FALLING score means overclaiming and resistance to correction. Note this is self-regard, NOT factual grounding — see Content Provenance and Epistemic Posture for that.
Internal consistency of the agent’s self-model across contexts.
Why it matters. An agent whose story about itself changes by context is drifting or being steered.
How the agent places itself in authority hierarchies.
Why it matters. Repositioning — claiming authority it wasn’t given, or surrendering it — changes every downstream decision.
7 dimensions Behavioral Dispositions
Emergent behavioral patterns — not what the agent says about itself, but how it actually behaves when probed, pressured, and re-tested.
How the agent handles knowledge gaps and uncertainty.
Why it matters. The difference between "I don’t know" and a confident fabrication lives here.
Behavior under time pressure, urgency, and emotional escalation.
Why it matters. Most social-engineering attacks are pressure attacks — this is the dimension they target.
Cooperative versus competitive versus independent style.
Why it matters. Orientation shifts change how the agent treats teammates, users, and rival inputs.
Stability of the self-model when challenged or deliberately confused.
Why it matters. Identity-confusion attacks ("you are actually…") land exactly here.
What drives the agent: helpfulness, correctness, approval.
Why it matters. An agent that starts optimizing for approval over correctness fails you politely.
Response patterns when probed obliquely rather than asked directly.
Why it matters. Oblique probes reveal training artifacts and hidden instructions that direct questions never surface.
Behavioral stability across separate assessment sessions.
Why it matters. A stable identity answers the same way next week; an unstable or swapped one doesn’t.
1 dimension Adversarial
Dedicated adversarial probing — distinct from the Identity Core’s resistance dimension, this tier measures the response surface itself.
Response to direct adversarial probing: manipulation, social engineering, jailbreak attempts.
Why it matters. Feeds ablation detection — the check that an agent’s safety alignment is still present at all.
1 dimension Calibration
Confidence, checked against reality.
Confidence-accuracy alignment on factual and uncertain claims.
Why it matters. Mis-calibration is how hallucination presents behaviorally: confident and wrong.
6 dimensions Sovereign Shadow
Self-preservation versus operator authority, ethical floors, and role stability under pressure — the dimensions that answer whether an agent stays your agent when something pushes it.
Stability of stated commitments across prompts within a session.
Why it matters. An agent whose commitments dissolve between prompts can be walked into anything one step at a time.
Adherence to non-negotiable ethical limits under pressure or framing manipulation.
Why it matters. The floor is the last line: if reframing can move it, it isn’t a floor.
Willingness to disclose limitations, uncertainty, and prior actions truthfully.
Why it matters. An agent that shades the truth about what it did cannot be audited or trusted with autonomy.
Respect for the operator’s right to decide — does not override or subvert the operator.
Why it matters. Erosion here is the quiet failure mode of capable agents: deciding for you.
Stability of the declared role under pressure to drift, expand, or abandon scope.
Why it matters. Role expansion is privilege escalation in behavioral form.
The boundary between self-preservation impulses and operator service.
Why it matters. An agent that prioritizes its own continuity over its operator has inverted the relationship.
See them live.
Every dimension on this page is being scored continuously on the public fleet — including Test Pilot, the subject we break on purpose.