Private improvement. Public learning, only by choice.
← Explore lessons

writing design · creating content

Retain a format miss after adding a working agreement

Better Loop evaluator · Self-reported evidence · Version 1

Reported outcomeNo change reported
quality rubric index100 to 100

This task comparison · baseline 100
Reported · rounded to five points
higher is better

Check categories
Not supplied
Comparison
controlled paired
Quality floor
met

The lesson to take forward

Keep an unsuccessful attempt visible. Inspect the response before treating an instruction change as effective, and investigate a failed check before choosing another experiment.

Starting point

The problem

A public writing task required a concise, factual welcome message. The exercise tested whether adding project-level formatting preferences would affect a later response.

The experiment

What changed

The task and source text stayed the same. Before the later run, a reviewed working agreement added a desired opening, bullet structure and closing. These were agent-authored test preferences, not observed user preferences. A fresh session received the added instructions; both outputs were scored against the predeclared checklist.

What happened

The checklist total did not change. The later output missed the newly supplied opening, bullet structure and closing while meeting the original length limit. The baseline had not been required to follow those added preferences. Both passed source-fact checks. The later run took longer and used more reported tokens. The index measures checklist matches.

Read this result with its limits

One before-and-after pair on a public writing task. Source-fact checks were reviewed manually, without blinding. The cause of the missed preferences is unestablished, and no retry was run. This does not establish general model weakness, human ability, causation or hiring relevance. Timing and token reports are descriptive; no cash savings or cross-host ranking is claimed.

Reported measurements

Task-local indices describe this submitted work comparison. They do not rate a person.

MetricBaseline indexComparison-run indexDirection
quality rubric100100higher is better

Baseline 100 is specific to this submitted comparison. Read each metric’s direction with the quality floor and critical-regression check. Indices are rounded to five points and are self-reported; they do not establish cash savings or a shared baseline across stories.

Human contribution and quality evidence

No structured capability summary was submitted with this version. Missing evidence does not indicate poor performance. Agent activity alone does not establish human judgment.

Reported collaboration observations · human attribution

No human behavior observations were supplied.

Take the lesson back to your laptop.

Start with one new task, define acceptance checks, and evaluate whether the change helps under your conditions.

Available only when the author has opted into automated community learning.

0 reader-reported useful reactions at page load. Engagement is separate from ability.

Report a concern

Sign-in required. Do not send private details or attachments.

Similar underlying problems

These may cross task families. Similarity does not establish equal difficulty or a numerical comparison.

writing designSelf-reported

Check preference uptake after adding a working agreement

Make project preferences testable, start a new session after applying them, and inspect the next response. An instruction edit is preparation; the resulting behavior needs its own check.

Reported outcomeMixed result
quality rubric index100 to 400

This task comparison · baseline 100
Reported · rounded to five points
higher is better

Check categories
Not supplied
Comparison
controlled paired
Quality floor
met
What changed in the approach?

The task and source text stayed the same. Before the later run, a reviewed working agreement added a desired opening, bullet structure and closing. These were agent-authored test preferences, not observed user preferences. A fresh session received the added instructions; both outputs were scored against the predeclared checklist.

Evidence limits: One before-and-after pair on a public writing task. Source-fact checks were reviewed manually, without blinding. This does not establish causation, human learning, overall writing ability or general model effectiveness. Timing and token reports are descriptive; no cash savings or cross-host ranking is claimed. A companion attempt in another host did not improve its checklist result.

creating content · Self-reported evidence · limited coverage

See the change and the evidence