The three status labels
Every research page carries one of three statuses, and it is a required field in the schema rather than a convention. A study with no measurements and a study with measurements have to be distinguishable at a glance — in the listing, in the metadata, and on the page — because the failure mode of a "research" section is publishing a plausible methodology with invented numbers under it.
| Status | What it means | Count |
|---|---|---|
| Measured | The experiment ran here and the numbers are ours. Requires a sample description and a recorded environment — the schema refuses the page without them, because a number with no stated sample is unfalsifiable. | 2 |
| Harness | The apparatus is published and reproducible. There are no results, and the page says so in its first paragraph rather than in a footnote. | 3 |
| Planned | Designed, not built. | 0 |
3 of the 5 studies here carry no results. That is not a gap being apologised for. Running some of them requires paid accounts for products this site does not have, and publishing the apparatus without the numbers is the honest version of that situation — anyone with the access can run it and get a figure that means something.
What every study must state
These are schema fields, not editorial guidelines. A page missing one does not publish; the build fails.
- The question, phrased so it could come out either way. "Does Copilot improve productivity" is not a research question; it is a conclusion looking for support.
- At least two limitations. Minimum two, enforced. A study with no stated limitations has not been thought about, and the limitations sections here are written to be the part a critic would have written.
- Reproduction instructions concrete enough to follow — the commands, not a description of the approach.
- Three separate dates: published, updated, and last technically verified. They are different facts and conflating them is how a page comes to look fresher than it is.
- Named sources with access dates, weighted toward primary documentation.
Rules the measurements follow
Measured and interpreted are separated
A number and what we think it means are different claims with different reliability, and they are never in the same sentence. Where a study draws a conclusion, the conclusion is labelled as one.
A negative result is a result
Studies here have found that nothing interesting happened, and said so. A research section that only ever publishes findings is a research section that stopped publishing the other kind, and once a reader suspects that, none of the findings are worth much.
Nothing is inferred that could be measured, and nothing is invented that cannot
Where a value cannot be obtained reliably — the agent benchmark's intervention count is the clearest case, since no script can detect that a person nudged an agent — it is either entered by a human and labelled as such, or recorded as null. The scorer refuses to sum null as zero, because "not measured" and "zero" are different claims and only one of them is flattering.
Versions and dates are part of the result
Everything measured here is measured against software that changes weekly. A result names the versions it was taken against and the date it was taken, and the harnesses refuse to write a record that does not.
The data
3 of the 5 studies publish machine-readable data alongside the page, and the maintained datasets at /data/ are downloadable as JSON and CSV with their own sources and update history. Everything is licensed for reuse; the licence and the attribution it asks for are stated on each dataset page.
Reproducing or disputing a result
Every study's reproduction section is the actual procedure. If you run it and get something different, that is the most useful outcome either of us could produce — and it is more likely than not, because these are measurements of moving software.
Corrections go through the same route as everything else on this site: the corrections process. A confirmed error moves the updated date and the change is described on the page rather than made quietly.
What this is not
This is not an academic journal and does not pretend to be. There is no peer review, no institutional affiliation, and no claim to statistical power. What there is: published apparatus, stated limitations, real dates, and numbers that either exist or are explicitly absent.
Back to the research index · How code examples are tested · Editorial policy