Private improvement. Public learning, only by choice.

Methodology · better-loop-fluency-0.1

Show the evidence.
Keep the limits.

Better Loop separates how a person works with AI, what happened on the task, and how confidently we can interpret it.

Research foundation, not a validated ranking system.

General efficiency gains, person ability ratings and cohort comparisons remain unvalidated.

Observable collaboration

The behavioral foundation is Anthropic’s AI Fluency Index and the 4D framework: Delegation, Description, Discernment, and Diligence. Better Loop is an independent adaptation, without Anthropic endorsement.

The 11 conversation-visible indicators are grouped below. None is a Diligence indicator. Diligence reflection needs separate explicit evidence; it is not part of the 11-indicator Index.

Observations can be observed, not observed, not applicable, or insufficient evidence. Missing evidence is not a zero. An agent’s autonomous action is not automatically evidence of human judgment.

Outcomes are a separate measure

Task quality, model time, human effort, tokens, and estimated API cost are Better Loop measures, separate from the published Index. Cash billing and estimated API cost are not interchangeable. Include parent, worker, judge, retry, and cache costs when known.

Baseline comparisons follow the public skill-creator evaluation approach: prespecified expectations, equivalent tasks, blinded review where appropriate, all trials retained, and honest neutral or adverse findings. This method does not establish efficacy by itself.

An improvement claim needs comparable evidence, a favorable primary measure, the prespecified quality floor, and no critical regression. Unknown quality cannot become a win. A zero baseline has no relative percentage.

Three labels, three meanings

Similar tasks are not equivalent tasks

Task family, problem type, objective, difficulty, conditions, demonstrated skills, and measurement version help organize stories. Numerical comparisons need additional compatibility and calibration. There is no universal “best AI worker” score.

Descriptive cohorts require eligible, consented evidence and privacy suppression. Synthetic examples, withdrawn stories, unknown quality, and incompatible measurements cannot contribute. Meeting a count threshold alone does not validate percentiles.

Source mapping and next evidence

The repository’s foundation reference maps all 11 stable indicator IDs to their sources and records the proposed validation work. Claude Code and Codex have selected-export adapters; host availability does not establish validity in every work domain.