Methodology · better-loop-fluency-0.1
Show the evidence.
Keep the limits.
Better Loop separates how a person works with AI, what happened on the task, and how confidently we can interpret it.
General efficiency gains, person ability ratings and cohort comparisons remain unvalidated.
Observable collaboration
The behavioral foundation is Anthropic’s AI Fluency Index and the 4D framework: Delegation, Description, Discernment, and Diligence. Better Loop is an independent adaptation, without Anthropic endorsement.
The 11 conversation-visible indicators are grouped below. None is a Diligence indicator. Diligence reflection needs separate explicit evidence; it is not part of the 11-indicator Index.
- Delegation: goal definition; approach consultation.
- Description: iterative refinement; quality examples; output structure; collaboration mode; tone preferences; audience definition.
- Discernment: context gap detection; reasoning scrutiny; factual verification.
Observations can be observed, not observed, not applicable, or insufficient evidence. Missing evidence is not a zero. An agent’s autonomous action is not automatically evidence of human judgment.
Outcomes are a separate measure
Task quality, model time, human effort, tokens, and estimated API cost are Better Loop measures, separate from the published Index. Cash billing and estimated API cost are not interchangeable. Include parent, worker, judge, retry, and cache costs when known.
Baseline comparisons follow the public skill-creator evaluation approach: prespecified expectations, equivalent tasks, blinded review where appropriate, all trials retained, and honest neutral or adverse findings. This method does not establish efficacy by itself.
An improvement claim needs comparable evidence, a favorable primary measure, the prespecified quality floor, and no critical regression. Unknown quality cannot become a win. A zero baseline has no relative percentage.
Three labels, three meanings
- Performance: the measured outcome, including mixed or negative results.
- Coverage: what the selected evidence does and does not show.
- Verification: how the work evidence was obtained and checked. Email verification only proves access to an account.
Similar tasks are not equivalent tasks
Task family, problem type, objective, difficulty, conditions, demonstrated skills, and measurement version help organize stories. Numerical comparisons need additional compatibility and calibration. There is no universal “best AI worker” score.
Descriptive cohorts require eligible, consented evidence and privacy suppression. Synthetic examples, withdrawn stories, unknown quality, and incompatible measurements cannot contribute. Meeting a count threshold alone does not validate percentiles.
Source mapping and next evidence
The repository’s foundation reference maps all 11 stable indicator IDs to their sources and records the proposed validation work. Claude Code and Codex have selected-export adapters; host availability does not establish validity in every work domain.