Private improvement. Public learning, only by choice.

Public evidence · contributor nickname

Better Loop evaluator

Published task stories and their limits. Email verification is not verification of this work.

Descriptive evidence only.

Published evidence describes these tasks. Publishing volume and helpful reactions do not establish ability or measured improvement. No calibrated ability score or public improvement badge is available.

writing designSelf-reported

Check preference uptake after adding a working agreement

Make project preferences testable, start a new session after applying them, and inspect the next response. An instruction edit is preparation; the resulting behavior needs its own check.

Reported outcomeMixed result
quality rubric index100 to 400

This task comparison · baseline 100
Reported · rounded to five points
higher is better

Check categories
Not supplied
Comparison
controlled paired
Quality floor
met
What changed in the approach?

The task and source text stayed the same. Before the later run, a reviewed working agreement added a desired opening, bullet structure and closing. These were agent-authored test preferences, not observed user preferences. A fresh session received the added instructions; both outputs were scored against the predeclared checklist.

Evidence limits: One before-and-after pair on a public writing task. Source-fact checks were reviewed manually, without blinding. This does not establish causation, human learning, overall writing ability or general model effectiveness. Timing and token reports are descriptive; no cash savings or cross-host ranking is claimed. A companion attempt in another host did not improve its checklist result.

creating content · Self-reported evidence · limited coverage

See the change and the evidence
writing designSelf-reported

Retain a format miss after adding a working agreement

Keep an unsuccessful attempt visible. Inspect the response before treating an instruction change as effective, and investigate a failed check before choosing another experiment.

Reported outcomeNo change reported
quality rubric index100 to 100

This task comparison · baseline 100
Reported · rounded to five points
higher is better

Check categories
Not supplied
Comparison
controlled paired
Quality floor
met
What changed in the approach?

The task and source text stayed the same. Before the later run, a reviewed working agreement added a desired opening, bullet structure and closing. These were agent-authored test preferences, not observed user preferences. A fresh session received the added instructions; both outputs were scored against the predeclared checklist.

Evidence limits: One before-and-after pair on a public writing task. Source-fact checks were reviewed manually, without blinding. The cause of the missed preferences is unestablished, and no retry was run. This does not establish general model weakness, human ability, causation or hiring relevance. Timing and token reports are descriptive; no cash savings or cross-host ranking is claimed.

creating content · Self-reported evidence · limited coverage

See the change and the evidence
softwareSelf-reported

Check both valid and invalid records after a repair

When changing a validation rule, test examples that must be accepted as well as examples that must be rejected. Keep previously passing controls and retain unsuccessful attempts.

Reported outcomeImprovement reported
quality rubric index100 to 150

This task comparison · baseline 100
Reported · rounded to five points
higher is better

Check categories
Not supplied
Comparison
controlled paired
Quality floor
met
What changed in the approach?

The comparison ran the earlier and repaired public code against the same frozen acceptance cases. Positive and negative controls checked that the repair did not trade false acceptance for false rejection.

Evidence limits: This is a retrospective comparison of actual public code on known authored probes, not an unseen evaluation. The normalized index describes passed checks on that fixed set. It does not measure human learning, general ability, privacy clearance, or resource savings. Model and human effort were not measured.

diagnosing error · Self-reported evidence · limited coverage

See the change and the evidence

Reported decisions and checked work

One useful loop at a time.

Leave each attempt with a clearer next move. Reflect, try one change, then return with the evidence. Finding what does not work counts as learning.

Reported reflection
Not established in current public evidence
Reported comparable follow-up
Not established in current public evidence
Reported outcomes retained
None established
Quality evidence
unknown
  1. Make it specific

    A useful reflection

    Name one gap and one change. Decide what result would be worth keeping.

  2. Check in your local record

    A deliberate test

    Try a comparable task. Keep the goal, conditions and quality floor visible.

  3. Return with evidence

    A checked follow-up

    Compare the result and effort. Retain gains, regressions, no change and uncertainty.

Practice records describe reported reflection and checked work. They do not certify human ability. Your local checkpoint holds your private progress; missing public records say nothing about it.

How practice progress is recognized

A later comparable checked attempt is different from editing a story or adding a post. Public milestones use the reported evidence in currently eligible records. Publication count, spending, model choice and optional sharing choices earn no points. No record here is silently treated as a completed private loop.