Private improvement. Public learning, only by choice.

Your work. Your next useful decision.

Get better at
working with AI.

Make your next attempt count. Sharpen your prompt, shape how AI works with you, and check what changed in your own work.

On your laptop · Claude Code or Codex · No account needed

THE NEXT ATTEMPT

A small change.
A check that matters.

01

Choose and sharpen

Pick a problem. Make your goal, prompt and acceptance check clear.

02

Agree and try

Review a scoped working agreement, then make a deliberate next attempt.

03

Keep the evidence

Check your decision and outcome. Better, worse, or unchanged: keep the lesson.

For every kind of knowledge workCodeAnalysis & financeResearch & strategyMath & scienceWriting & designOperations & educationGeneral tasks

Curated public work

See what good work
can look like.

Real work, published in the open. Find a useful technique in a notebook, guide, report or project, then try it in your own work.

Browse 28 curated public examples

Publisher names credit the original sources. They do not imply Better Loop membership or a shared assessment.

Curated public exampleCode

Code & software

Iterate on a tolerant XML parser

Ronacher published a tolerant XML parser after earlier parsing approaches failed. He describes directing the initial implementation, reviewing it, requesting further fixes and an extensive test suite, and releasing the resulting library.

Published byArmin Ronacher

Curated public exampleNotebook

Analysis & finance

Rolling market-exposure estimates with statsmodels

The notebook applies a 60-month rolling CAPM to technology-industry excess returns using Ken French’s factor and industry data. It displays coefficient tables, confidence-interval plots, and an expanding-window example.

Published bystatsmodels developers

Curated public exampleGuide

Research & strategy

Mapping grocery proximity in Chicago

The guide projects Chicago community areas and grocery locations into a metric coordinate system, buffers stores by one kilometre, and displays intersection and difference maps. It defines coverage geometrically.

Published byGeoPandas developers

Curated public exampleNotebook

Math & science

Partial pooling for Minnesota radon measurements

The notebook joins Minnesota household and county data, fits pooled, unpooled and hierarchical radon models, and displays posterior diagnostics and comparisons. The partial-pooling plots show how estimates change with county sample size.

Published byChris Fonnesbeck and contributors / PyMC

Curated public exampleNotebook

Writing & design

Ground a short caption in the pictured product

OpenAI works through furniture images from an Amazon product dataset, generating descriptions and then concise captions. Printed examples show how the product title identifies the intended item when an image contains several objects.

Published byOpenAI Cookbook

Curated public exampleLesson

Operations & education

Teaching plots with Gapminder data

Software Carpentry’s lesson reads Gapminder country data and displays labeled GDP time-series and scatter plots. It shows table transposition, legends and file export, alongside learner exercises and accessibility guidance.

Published bySoftware Carpentry / The Carpentries

Find a place to begin

A skill to practise.
A task to make your own.

Reconcile two totals. Challenge an answer. Give a draft a clearer brief. Start with a familiar situation and one check that matters.

Explore the practice boards
DelegationChoose the decisions that need you.DescriptionMake the result easier to recognise.DiscernmentTest the answer before you trust it.

Authored practice briefs. Choose routine, moderate or complex work; no participant counts or earned scores are implied.

Better Loop project runs

Real attempts. Useful lessons.

Start with our public evaluations: a repair that worked, guidance that didn’t improve the checked result, and host behavior with limits. These are project case studies, separate from community contributions.

Better Loop project runActual guidance comparison

More guidance. Same checked result.

Would extra guidance help identify when an approval remains valid?

Frozen case checks
12/1212/12
Reported model tokens
34,61734,981

Both attempts passed 12/12 cases and 3/3 scope checks. Guidance used 364 more reported tokens (about 1.05%).

A move to try

A more elaborate prompt needs to earn its place. Keep the same acceptance check when you try it.

Harness
Claude Code 2.1.270
Model evidence
Claude Opus 5 + Haiku 4.5 auxiliary
Task complexity
Moderate · rubric-estimated
Checks, conditions and limits

Quality: No measured quality gain on the frozen deterministic checks.

Resources: 34,617 → 34,981 reported tokens. All reported main and auxiliary model entries, including cache use, counted once. Shared overhead, cash billing and human effort unknown.

Comparison: One task, one pair, baseline first. Same frozen public code, structured-answer task and bl-approval-binding-judge-0.1 evaluator. Reported model entries include the auxiliary classifier.

Difficulty: Rubric-estimated in the frozen registration.

Exact reported model entries: claude-opus-5[1m], claude-haiku-4-5-20251001. These historical labels do not assert current availability.

Not counterbalanced or independently held out. No full-skill, human-improvement, speedup or savings claim.

Inspect the recorded measurements
Better Loop project runActual public software repair

Fix the claim. Keep the honest unknowns.

Could a validator reject unsupported human-action claims and still accept missing evidence?

Frozen regression checks
16/2424/24

Frozen checks passed 16/24 before and 24/24 after. Eight known failures were fixed; no case regressed.

A move to try

Write a test for the mistake and a control for what must keep working. Missing evidence can be the correct answer.

Harness
Node 22.10.0 · deterministic validator checks
Model evidence
No model invoked
Task complexity
Not rated
Checks, conditions and limits

Quality: All 24 decisions, six positive controls and ten negative controls met the prespecified floor.

Resources: No models invoked. Coding effort, model resources and cash cost were not measured. Execution duration is not a before/after speed comparison.

Comparison: Retrospective comparison of evidence draft.1 and draft.2 on the same known authored probes. Evaluator bl-evidence-attribution-repair-0.2; before then after.

Difficulty: No difficulty rating recorded for this repair evaluation.

Known, non-blinded cases. The initial infrastructure attempt evaluated no cases; the corrected retry is not replication. This measures a software repair, not a person's ability or model efficacy.

Inspect the recorded measurements

These different tasks, versions and evaluators cannot rank models, hosts or people. Public project measurements do not enter community cohorts, count as challenge completions or establish improved human ability.

Explore Codex, Claude and software project runs

Learning from real work

Real attempts.
Something to take forward.

Read what changed, what was checked, and what happened. A useful negative result can save you an unhelpful next step.

writing designSelf-reported

Retain a format miss after adding a working agreement

Keep an unsuccessful attempt visible. Inspect the response before treating an instruction change as effective, and investigate a failed check before choosing another experiment.

Reported outcomeNo change reported
quality rubric index100 to 100

This task comparison · baseline 100
Reported · rounded to five points
higher is better

Check categories
Not supplied
Comparison
controlled paired
Quality floor
met
What changed in the approach?

The task and source text stayed the same. Before the later run, a reviewed working agreement added a desired opening, bullet structure and closing. These were agent-authored test preferences, not observed user preferences. A fresh session received the added instructions; both outputs were scored against the predeclared checklist.

Evidence limits: One before-and-after pair on a public writing task. Source-fact checks were reviewed manually, without blinding. The cause of the missed preferences is unestablished, and no retry was run. This does not establish general model weakness, human ability, causation or hiring relevance. Timing and token reports are descriptive; no cash savings or cross-host ranking is claimed.

creating content · Self-reported evidence · limited coverage

See the change and the evidence
writing designSelf-reported

Check preference uptake after adding a working agreement

Make project preferences testable, start a new session after applying them, and inspect the next response. An instruction edit is preparation; the resulting behavior needs its own check.

Reported outcomeMixed result
quality rubric index100 to 400

This task comparison · baseline 100
Reported · rounded to five points
higher is better

Check categories
Not supplied
Comparison
controlled paired
Quality floor
met
What changed in the approach?

The task and source text stayed the same. Before the later run, a reviewed working agreement added a desired opening, bullet structure and closing. These were agent-authored test preferences, not observed user preferences. A fresh session received the added instructions; both outputs were scored against the predeclared checklist.

Evidence limits: One before-and-after pair on a public writing task. Source-fact checks were reviewed manually, without blinding. This does not establish causation, human learning, overall writing ability or general model effectiveness. Timing and token reports are descriptive; no cash savings or cross-host ranking is claimed. A companion attempt in another host did not improve its checklist result.

creating content · Self-reported evidence · limited coverage

See the change and the evidence
softwareSelf-reported

Check both valid and invalid records after a repair

When changing a validation rule, test examples that must be accepted as well as examples that must be rejected. Keep previously passing controls and retain unsuccessful attempts.

Reported outcomeImprovement reported
quality rubric index100 to 150

This task comparison · baseline 100
Reported · rounded to five points
higher is better

Check categories
Not supplied
Comparison
controlled paired
Quality floor
met
What changed in the approach?

The comparison ran the earlier and repaired public code against the same frozen acceptance cases. Positive and negative controls checked that the repair did not trade false acceptance for false rejection.

Evidence limits: This is a retrospective comparison of actual public code on known authored probes, not an unseen evaluation. The normalized index describes passed checks on that fixed set. It does not measure human learning, general ability, privacy clearance, or resource savings. Model and human effort were not measured.

diagnosing error · Self-reported evidence · limited coverage

See the change and the evidence

Your practice comes first

Get the benefit.
Keep your work yours.

Keep assessments, before/after work and your return checkpoint on your laptop. Private coaching stays useful even if you never publish.

Raw work stays out of Better Loop. Your configured model provider may process the evidence you select.

Build your practice →
What counts as evidence of improvement?

A later human decision, a clear check and an honest comparison of the work you selected. Copying prompts or applying preferences is preparation, not measured improvement. Results may be negative or inconclusive. Public evidence is self-reported; it does not establish general ability or guaranteed gains. See the measurement approach.