complexity_mean
2.24
cc
Best: human_control
Worst: claude_code (3.35)
Mean McCabe cyclomatic complexity per function (radon)
correction_freq
0.00
per_kkey
Best: antigravity
Worst: human_control (334.45)
Backspace + delete events per 1000 keystrokes
duplication_pct
0.00
%
Best: claude_code
Worst: replit_agent (9.56)
% of source lines inside a duplicated 6-line shingle
hallucinations
0.00
count
Best: claude_code
Worst: replit_agent (1.00)
Features shipped that were NOT in the spec
security_density
0.00
per_kloc
Best: human_control
Worst: cursor_agent (43.67)
OWASP/CWE-tagged Bandit findings per 1000 lines of Python (per-language density)
Provenance metadata
{
"conditions": [
"human_control",
"claude_code",
"cursor_agent",
"replit_agent",
"antigravity"
],
"kind": "main_study",
"note": "main_001 AI matrix (4 conditions x 3 specs x 10 reps) plus the human_control baseline (1 rep per spec). Dashboard shows per-(metric,condition) MEANS across specs/reps. Human is a single-rep reference point (see docs/PROTOCOL_DEVIATIONS.md, Deviation 003).",
"spec_files": [
"specs/agent_education_system.yaml",
"specs/data_pipeline.yaml",
"specs/internal_tool_cli.yaml"
]
}