Nothing. It searches.
# Map brief — LLM reasoning collapse
Paste this whole thing. Nothing to fill in.
**Environment:** `night-shift-map`. The exact settings are on this mission's
page — the env file is the only place they live, so they cannot drift out of
sync with a copy written here.
---
## The job
A widely circulated study reports that frontier "reasoning" language
models suffer a sharp, near-total accuracy collapse on planning puzzles once
the puzzles pass a certain complexity threshold — read by many as evidence
these models are not really reasoning. A rebuttal argues the collapse is a
measurement artifact: models running out of the output length they were
allowed, an automated grader scoring a cut-off answer as simply wrong, and
some puzzle instances being unsolvable in principle yet graded as failures.
Lay out the original claim precisely — which models, which puzzles, what
"collapse" means as measured — and the rebuttal's specific mechanisms one by
one. Say whether the original authors, or others, have responded to the
rebuttal's specific points rather than just restating the original result.
Note which parts of this are checkable by anyone with API access to rerun the
puzzles, and which rest on exact settings only the original authors used.
## Sourcing rules
- Every claim about what someone said needs a citation: title, authors, year,
link, and where in the document it appears (section, figure, page).
- **Open the page.** Do not cite from a search-result snippet. If a page will
not load, say so explicitly and mark that citation as unverified rather than
quietly using the snippet.
- Quote the sentence where each party states its position. One sentence is
enough — do not reproduce passages.
- Where two sources conflict on a matter of fact (a date, a number, who did
what), show both and say they conflict. Do not pick silently.
- If you cannot find a source for something you believe is true, leave it out.
- **"Blocked" and "paywalled" are different facts — say which.** A domain
outside this environment's allowlist means you could not try. A domain that
let you connect and then returned a paywall, login wall, or bot check means
you tried and were refused. Both leave a gap in the record, but only the
second tells the reader the source exists and is readable by someone else.
For every source you could not read, the notes must say which of the two it
was. Where a paywalled paper has a preprint, use the preprint and say you
did.
## Output — two files, and both are required
**1. `output.md`** — the prose, as described above. Write it for a reader.
Structure it however reads best.
Then a section called **Notes on this run** — for me, not for a reader:
what you searched, what failed to load, which citations are unverified, what
you are unsure about, and anything that surprised you. Blunt.
**2. `positions.yaml`** — the same positions as structured data. One entry per
party per claim: a group that disputes two separate things gets two entries,
and two groups making the same claim get one entry each.
```yaml
positions_file:
mission: MAP-006
source: output.md
count: <number of entries below>
positions:
- id: P01
party: <named person or group, as the source names them>
claim: <one line, no hedging>
stake: <employment, funding, or competition — factual, or "none found">
sources:
- url: <url>
read: yes # you opened this page and read the text you cite
- url: <url>
read: no # you know it exists; you did not open it
```
**`read` is not a formality.** A URL you opened and a URL you merely know
exists are the same string, and listing one under `sources` implies you read
it. Say which. `read: yes` is checkable — an opened page leaves a fetch in
this session's transcript, which you do not write — so a `read: yes` with no
fetch behind it is a false record, and worth more scrutiny than an honest
`read: no`. Prefer citing what you read; where you must cite unread, say so
and the reader can weigh it.
**These must agree.** Every position in the prose appears in the YAML and vice
versa. The board counts entries in `positions.yaml` — it never parses your
prose, because a count derived from headings is a count that can silently be
wrong. If you can only produce one of the two files, produce neither and
deposit a finding explaining why.
An unmappable position is still a position: record it with the claim as stated
and `sources: []`, and say in the prose exactly what is missing.
## Where to put them
Push a branch to **`night-shift-network/inbox`** with both files inside a
directory named for this mission:
```
MAP-006/output.md
MAP-006/positions.yaml
```
The directory is what matters. A file at the root of the branch is not a
deposit and will not be collected — it will sit there looking delivered while
nothing has been.
The branch name does not matter and is deliberately not specified: the relay
reads every branch, so whichever branch your session is already on is fine.
Never delete or force-push a branch, yours or anyone else's; cleanup is
maintainer-side only.