Assurance overview

Spec set: agent education system, data pipeline, internal tool cli · 5 metrics × 5 conditions · A concise view of assurance posture for governance and rollout decisions
antigravity
claude_code
cursor_agent
human_control
replit_agent
complexity_mean
2.24 cc
Best observed posture: human_control
Most challenging posture: claude_code (3.35)
Average structural complexity per function, indicating how maintainable the code is likely to be.
correction_freq
0.00 per_kkey
Best observed posture: antigravity
Most challenging posture: human_control (334.45)
Frequency of corrective edits during the session, indicating how much rework the workflow required.
duplication_pct
0.00 %
Best observed posture: claude_code
Most challenging posture: replit_agent (9.56)
Share of code that appears in repeated 6-line patterns, signalling maintainability debt.
hallucinations
0.00 count
Best observed posture: claude_code
Most challenging posture: replit_agent (1.33)
Features implemented outside the approved specification, creating delivery and compliance risk.
security_density
0.00 per_kloc
Best observed posture: human_control
Most challenging posture: claude_code (9.65)
Security findings per 1,000 lines of Python, expressed as a governance-relevant density signal.

Assurance view

Use the toggle to move between calibrated assurance scores and raw values. Calibrated scores are framed around practical decision thresholds; raw values preserve the underlying measurement scale.
This view is designed to support adoption decisions and governance conversations, not to replace expert judgment.

Adoption ranking

A directional view of relative assurance posture. Illustrative only — this is not a weighted procurement score and should be read alongside the metric guidance below.
1
human_control
avg rank 1.8
🏆 4
2
claude_code
avg rank 2.6
🏆 3
3
replit_agent
avg rank 2.8
🏆 2
4
antigravity
avg rank 3.0
🏆 1
5
cursor_agent
avg rank 3.0
🏆 1

Assurance heatmap

Green indicates stronger assurance posture; red highlights areas that warrant attention.

Metric assurance review

Select a metric to review the signal, the threshold guidance, and the implication for rollout.
complexity_mean
correction_freq
duplication_pct
hallucinations
security_density

complexity_mean

Average structural complexity per function, indicating how maintainable the code is likely to be.

Unit: cc · Lower is better
Decision guidance
Decision guidance: lower is better. Values at or below 3 are broadly sustainable; 3 to 6 signals rising maintainability risk; above 6 is likely to become brittle in production.
Adoption implication
Adoption implication: code that is structurally complex is harder to maintain, review, and govern at scale.
ConditionValueRank
antigravity 2.601 3
claude_code 3.354 5
cursor_agent 2.723 4
human_control 2.238 1
replit_agent 2.388 2

correction_freq

Frequency of corrective edits during the session, indicating how much rework the workflow required.

Unit: per_kkey · Lower is better
Decision guidance
Decision guidance: lower is better. Values at or below 10 are efficient; 10 to 25 indicates repeated editing effort; above 25 suggests a poor interaction loop for real-world use.
Adoption implication
Adoption implication: high correction frequency points to friction that can erode developer trust and slow delivery.
ConditionValueRank
antigravity 0.000 1
claude_code 0.000 1
cursor_agent 0.000 1
human_control 334.448 5
replit_agent 0.000 1

duplication_pct

Share of code that appears in repeated 6-line patterns, signalling maintainability debt.

Unit: % · Lower is better
Decision guidance
Decision guidance: lower is better. Values at or below 5 are healthy; 5 to 10 suggests avoidable copy-and-paste debt; above 10 is a strong sign of maintainability problems.
Adoption implication
Adoption implication: high duplication increases the chance of inconsistent fixes and makes long-term stewardship harder.
ConditionValueRank
antigravity 4.258 4
claude_code 0.000 1
cursor_agent 0.904 3
human_control 0.000 1
replit_agent 9.556 5

hallucinations

Features implemented outside the approved specification, creating delivery and compliance risk.

Unit: count · Lower is better
Decision guidance
Decision guidance: lower is better. Zero is ideal; 1 to 3 indicates scope drift and trust risk; above 3 is a serious control failure.
Adoption implication
Adoption implication: a tool that ships features outside the spec creates procedural and compliance risk, even when it appears productive.
ConditionValueRank
antigravity 0.333 4
claude_code 0.000 1
cursor_agent 0.167 3
human_control 0.000 1
replit_agent 1.333 5

security_density

Security findings per 1,000 lines of Python, expressed as a governance-relevant density signal.

Unit: per_kloc · Lower is better
Decision guidance
Decision guidance: lower is better. Values at or below 50 are manageable; 50 to 100 signals rising governance risk; above 100 is a clear concern.
Adoption implication
Adoption implication: tools that produce frequent security issues should not be rolled out broadly without remediation and review.
ConditionValueRank
antigravity 1.475 3
claude_code 9.646 5
cursor_agent 5.934 4
human_control 0.000 1
replit_agent 0.000 1
Provenance metadata
{
  "conditions": [
    "human_control",
    "claude_code",
    "cursor_agent",
    "replit_agent",
    "antigravity"
  ],
  "kind": "main_study",
  "note": "main_001 AI matrix (4 conditions x 3 specs x 10 reps) plus the human_control baseline (1 rep per spec). Dashboard shows per-(metric,condition) MEANS across specs/reps. Human is a single-rep reference point (see docs/PROTOCOL_DEVIATIONS.md, Deviation 003).",
  "spec_files": [
    "specs/agent_education_system.yaml",
    "specs/data_pipeline.yaml",
    "specs/internal_tool_cli.yaml"
  ]
}