Logos52
journal / 2026 08 15 what works grok 46 and grok bot

What works: Grok 4.6 and Grok Bot

journal updated 2026-09-01

What works: Grok 4.6 and Grok Bot

Routing update, 2026-09-01: the writer seat is Grok 4.6. See Grok writes. Banks, execution, and Grok Bot below still hold. The Fable-writes verdict in this file is the 15 August ranking, not the live seat.

Verdict (15 August): keep the three-way split. Fable writes anything this vault has to like. Grok 4.6 compiles evidence and executes on disk. Grok Bot stands a duty, files a packet, and never holds a login this desk would miss. Collapsing them into “just use Grok” loses the only comparison this desk has actually scored.

The comparison is five jobs this desk already runs, not three product pages. Grok 4.6 and Grok Bot is the name collision. This entry is the ranking.

What was scored

JobWhat wonEvidenceWhat lost
Wiki prose, including openingsFable (Opus when Fable is out)A/B/C, 13 Aug: Fable better on all three pages, “not even close.” Flow State bakeoff the day before, same way. 14 Aug: 52 Grok intros failed the pillow test; four generator slates later the owner handed Fable the doors. Opus 12/13/24 picked when Fable ran out of tokens.Grok 4.6 at this vault’s register. Mechanical checks passed both arms. Quality did not.
Research banksGrok 4.6281 banks against a 255-page queue. Lane is agent-agnostic. The A/B/C Grok arm had already written ~260 of those banks and still lost the writing test, so the bank win is not a writing win in disguise.Fable as a bank mill. Cold Fable can bank; Grok already did the volume.
Local executionGrok Build calling 4.6Stack since August. $2 / $6. AA-Briefcase: ~53 turns / ~0.5B input vs Opus 5 max ~103 / ~2.0B. Profiles isolate n1 from work.Fable on toolchain loops. DeepSWE 65.9% vs Fable 70% and Sol 73% is the coding gap that remains. Accept it for execution; do not pretend it is gone.
Standing dutyGrok Bot, one duty, packet only, public onlyOfficial geometry: one computer per user, laptop-closed work, handoff for 2FA. Structure A and Bot Operating Rules already describe the seats. Field packet, 15 Aug: Demo-shaped sweep → draft → approve is the only named setup that matches Watch/Brief/Intake.Twelve-bot afternoons. Chief-of-staff that “just routes.” Grocery, Gmail, Amazon, spend-without-a-gate. Per-bot isolation as a security story.
Harness (Cursor IDE vs Grok Build)Route by surface; habit unscored[[wiki/Research/Grok Build and Cursor BankBuild / Cursor bank]], 15 Aug. Build when the agent drives the session. Cursor IDE when you are in the files (Tab, visual diff, debugger). Same 4.6 is a harness pick. Cursor.app still not installed.

Writing

The register is the constraint, not the IQ card. Grok 4.6 sits at AA 61, a point or two behind Fable and Opus, at a fifth to a tenth of the token price. On this vault it still compresses an operating claim into a quote card, passes its own pillow test, and ships. The owner named eight, then kept finding the same object, then 52, then twelve promoted replacements still failed on pass 2. The checker/inventory ship is a backstop. The generator that does not make the mistake is Fable’s, and when Fable is gone it is Opus’s.

Price of that ruling: Fable tokens on every page that has to read well, including the craft cluster that sat held until 15 Aug. Price of ignoring it: another 52-page opener rewrite. What would flip it: a Grok slate the owner ranks above Fable on three pages, paths stripped off the board, after the generator research — not after another ban.

Grok still writes bodies at volume. Unattended promote after five accepts filed ~200 pages in an afternoon. That is throughput, not a taste win. The Wound and the Lie still opens on the struck four-item chain. Volume without an eye is how a mechanical pass becomes the page.

Banks and execution

Grok 4.6 is the right spend where a check is cheap. A bank is a ledger, links, verdicts, and a gap list — the checker can see a missing target. A build is a compiler, a test, a diff. Extra reasoning helps there; Thinking Models is that dial. xhigh is on the card and cannot be switched to off, only down from the default high. Use it on hidden bugs and high-value diagnosis. Do not use it as a personality upgrade on a taste-bound opening.

The steelman for “put 4.6 on everything”: same-price jump from 4.5, cheaper loops than Opus, already the Build default, Fable is $10 / $50. The condition that would flip the stack is written: a week of this desk’s long knowledge-work loops and hard SWE where 4.6 beats Opus 5, plus Fable’s trust premium dying on the pages that currently stay in Cowork. Day-one benches and a regen mill that needed Fable/Opus for the doors are not that week.

The teammate

What works in the field, once the Salesforce catalog is stripped:

  1. A standing sweep that stops for a person before a write. Palmer’s Demo Bot: daily bookmarks → draft prompt → approve → Cloud Agent → recorded check. That is Watch/Brief plus a gate. It is the only launch-week specimen that is the same shape as this vault’s report-is-the-product rule.
  2. One duty, named, quiet when nothing moved. Official anti-pattern is General Helper. Zakariasson #70 is the same voice as Brief. A silent day is a healthy estate, not a failed bot.
  3. A first task with five parts: outcome, sources, constraints, deliverable, review point. Official, 15 Aug fetch. Cheaper than a persona paragraph.
  4. Short recordings, then handwritten rules. Teach-a-task is ten minutes and drafts a skill. Long routines go brittle. Second-hand on the brittleness; the cap is first-party.
  5. A rent report. Zakariasson #50: vet fleet spend, kill waste. Three quiet packets retire the seat. That is already Bot Operating Rules.

What does not work, or is refused here:

  • Theme-not-task as a roster rule. Nate’s public TOC. Theme is wider than one duty and is how General Helper comes back wearing a nicer noun.
  • A chief-of-staff that delegates by itself. Pinto, 13 Aug: you have to tell it. Shumer’s researcher + writer + CoS is a group chat with a human still routing.
  • GUI-as-API onto mail, groceries, checkout, or WhatsApp. Portable as a mechanism. Illegal as a method on this computer. The shared box makes every login common property. Atomic Bot and the FAQ agree; Palmer’s grocery pair is the usage the trust line exists to name.
  • Overnight grind as a proof. Wii baseball shows the computer stays up. It is not a duty.

Field is still not a product bot. The first packet was worth opening. A Tue/Fri routine is legal if the next two packets still change the bank. Watch and Brief still meter first if the weekly allowance is tight.

Pairings that earn rent

  • Grok Bot drafts, desk pastes, Grok Build writes. Already assigned for 多恩刊: midnight packet, then Build. The Bot never touches the vault.
  • Grok 4.6 banks, Fable writes, checker grades, owner eyes the door. The regen split that survived contact. Do not invert it.
  • Grok Bot approves, Cursor Cloud Agent builds. Palmer’s pairing. Legal here only if the Cloud Agent’s repo is not the live wiki and the approval is real. Unscored on this Mac until the Ultra month actually runs arm D.
  • Two writers on one tree. Failed on 12 Jun (crib_diff caught a leaked gloss; roster cut inside an hour). One writer per tree.

What this is not

Not a claim that 4.6 is a worse model than Fable on a composite. It is a point behind on AA and cheaper on loops. Not a claim that Grok Bot is a toy. The standing half is the first new job this stack has added since Cowork. Not a recommendation to skip the Ultra month. The month is for two scored jobs, then stop.

The rejected reading is “the field has 100 use cases, so the fleet should grow.” Almost all 100 are someone else’s CRM. Two mechanisms survived: drive the GUI when the API permission is missing, and reconstruct a promise then flag the questions only a person can answer. Neither is a new seat. Both are notes under Intake and Brief.

Cost of holding this ranking: Fable spend on prose, unused Bot quota, a 4.6 that looks underused next to the launch screenshots. Cost of reversing it without a scored week: another opener class, a login on a shared machine, and a fourth agent to coordinate.