Every call is graded Hit, Partial, Miss, or Too early once the outcome is knowable, with the original wording quoted, never paraphrased into something easier to defend. How an issue gets made. Educational, not financial advice.
The Jun 15 issueJun 4, 2026The playPending
Sell a missed-lead recovery system to one local service niche at $3,500 up front + $1,500/month. Install a workflow that pulls dead leads, scores them, drafts follow-ups in the client's voice, and delivers a before-and-after revenue number.
How we grade itReader-reported results: did anyone run this play, close a client, and report real recovered revenue and/or recurring revenue?
Where it standsNo reader results yet. We will not invent one. The first real result lands in the Reader Win section of a future issue, quoted and dated.
The Jun 15 issueJun 4, 2026Market callPending
AI is shifting from a flat seat cost (subscription) to metered labor (COGS). Companies are capping or cutting unlimited token access because AI spending is now material and trackable.
How we grade itWatch enterprise AI budget decisions, public earnings commentary on AI spend, and platform pricing changes over the next two to four quarters. Specific signals: more companies announcing token caps, budget overruns attributed to AI inference costs, or AI line items appearing explicitly in cost-of-revenue rather than SaaS/tools budgets.
Where it standsThe June 15 issue cited Uber capping employees at $1,500/month after burning its annual budget in four months, and Walmart killing unlimited tokens. These are the named anchoring data points. Update 2026-07-25: the pattern extended: Tesla capping employees at $200/week in tokens with a budget-request form (AI Daily Brief, Jul 23) and operator-side budget discipline emerging (Ryan Carson's ~$5,000/employee/month settling estimate, Startup Ideas Podcast, Jul 24). Trending Hit; grades at the stated 2-4 quarter horizon.
The Jun 15 issueJun 4, 2026Market callPending
The top frontier models are converging in capability. Choosing the "best model" is no longer a durable business advantage.
How we grade itWatch benchmark scores across leading models over the next two to four quarters. Specific signal: do Opus, GPT-5.5, and Sonnet (and their successors) continue to land within a narrow band on general capability tests, or does one model break away decisively?
Where it standsThe June 15 issue cited a financial-analyst benchmark where Opus, GPT-5.5, and Sonnet landed within 0.3 points of each other. This is a single data point, not a long-run trend yet. Update 2026-07-20: the convergence continued and widened to open weights: six labs above 50 on the Artificial Analysis index (up from two in June), the top three models within 3 points across three labs, and open-weight Kimi K3 at 57 vs Fable 5's 60 (Innermost Loop Jul 18; The Rundown Jul 17). Trending Hit; grades at the stated 2-4 quarter horizon. Update 2026-08-03: a caveat that cuts against easy grading, logged forward-only. OpenAI reported on Jul 30 that fixing two settings in ARC-AGI-3's official harness moved GPT-5.6 Sol from 13.3% to 38.3% on 6x fewer output tokens, and Claude Opus 5 posted 30.2% vs its predecessor's 1.5% on the same benchmark (ARC Prize Foundation, Jul 24). A single benchmark number is now partly a measure of the test harness, so "within a narrow band" must be graded on like-harness comparisons or it grades noise. The call stands; the evidence standard for grading it just got stricter.
The Jun 15 issueJun 4, 2026Market callPending
The "Proof Premium" is real and growing. Only a small fraction of companies are capturing measurable business value from AI, and the gap between AI spend and proven returns is a structural opportunity for operators who can show before-and-after numbers.
How we grade itWatch whether research on enterprise AI ROI continues to show the majority of companies failing to document returns. Specific signal: does the Bain finding (44% funding next year's AI spend on unverified savings) remain directionally consistent in follow-up surveys or independent research over the next 12 months?
Where it standsThe June 15 issue cited Bain: 44% of companies funding next AI spend on savings that never appeared; 12% getting real business value from AI tools. These are the named anchoring data points. Result is "Too early" to grade on 12-month evidence.
The Jun 22 issueJul 13, 2026The playPending
Walk into one 25 to 150 person AI-heavy agency or B2B ops team, identify ONE recurring AI workflow that runs 50+ times per month and costs $300+/month, compile it into deterministic code with at most one model call per run, and bill $500 diagnostic + $750 build + 25% of verified monthly savings for 3 months (capped at $1,000). Realistic first-client total: about $1,800.
How we grade itReader-reported results: did anyone run the play, sign a client, ship a compiled workflow, and report a measured savings receipt with attribution rules?
Where it standsThe first real reader win lands in a future Reader Win slot, quoted and dated. We will not invent one.
The Jun 22 issueJul 13, 2026Market callPending
The architectural shift from "agent loop in production" to "compile the workflow into deterministic code + one model call per run" is the durable answer to the AI cost reckoning, NOT the consensus model-routing or spend-audit play. We expect to see this pattern (extract tacit rules from operator -> deterministic script + targeted small model call) emerge as a category of services over the next two to four quarters.
How we grade itSpecific signal: do we see at least 3 named operator/service businesses publicly offering "compile your AI workflow + share-of-savings" deals by end of Q1 2027, AND at least one publicly reported case study (named buyer + named saving) outside of Poetic itself?
Where it standsThe June 22 issue cited Poetic at AIG, SoFi, and Chime as the anchoring proof. 99%+ accuracy at 100x less token usage is the benchmark we are claiming the pattern can hit at the enterprise scale; the beginner cut targets 50%+ token savings on the modest case.
The Jun 29 issueJul 10, 2026The playPending
The Frontier Spread. Sell ONE fixed-price landing page or small marketing site at the existing market anchor (bracketed $900 / $500 / $350, most take $500), produce it through a plan → execute (GLM 5.2) → review three-model chain for under $1 of model spend, QC every unit, and keep roughly $495 gross margin on 2 to 3 hours of production work (client acquisition excluded).
How we grade itReader-reported results: did anyone win a gig at the anchor price, deliver through the chain, and publish a per-gig receipt (price charged, model cost, delivery time) under the kit's attribution rules?
Where it stands(archive issue 3, source week Jun 17-25; published 2026-07-10). The first real reader win lands in a future Reader Win slot, quoted and dated. We will not invent one.
The Jun 29 issueJul 10, 2026Market callPending
MARKET CALL (the window): This specific spread narrows materially within 2 to 3 quarters as buyers re-anchor to open-model production costs, AND the pattern repeats: with token prices falling roughly 10x annually (Metatrends, Jun 21, 2026; 280x over 24 months), each major open-model release reopens a fresh sell-at-the-old-anchor window in some deliverable class. The durable asset is the operator's anchor, receipts, and eval harness, not any single window.
How we grade itTwo checkable signals by end of Q1 2027: (1) do median fixed-price listings for landing pages on the major freelance marketplaces visibly step down from mid-2026 levels; (2) does at least one more open-weight release take a #1 spot on a design/output leaderboard for a commercial deliverable class at a 5x+ price gap, reopening the trade?
Where it stands(archive issue 3, source week Jun 17-25; published 2026-07-10). Anchoring receipts: GLM 5.2 $0.44 vs $2.38 on one benchmarked Opus 4.8 task (Startup Ideas Podcast, Jun 23); landing page 6 cents vs 49 cents (AI Daily Brief, Jun 18); Kimi K2.7 Code 94% cheaper in a single head-to-head (TLDR AI, Jun 18).
The Jun 29 issueJul 10, 2026Market callPending
MARKET CALL (the quality flip): Open-weight models hold a top-tier position (top 3) on crowd design/website leaderboards through end of 2026, keeping the execution layer of commercial web work commodity-priced; frontier vendors respond on price tiers rather than reclaiming a decisive quality gap in this deliverable class.
How we grade itDesign Arena (and successor/equivalent leaderboards) composition on 2026-12-31: are open-weight models still top 3 for website design? Secondary: any frontier release that retakes #1 AND holds a 3x+ price premium would grade this a Miss.
Where it stands(archive issue 3, source week Jun 17-25; published 2026-07-10). Anchor: GLM 5.2 #1 on Design Arena for website design at $4.40/M output tokens, Tailwind in 91% of sessions (announced Jun 16; AI Daily Brief, Jun 22). Crowd-vote leaderboard, noted as such in the issue.
The Jul 6 issueJul 10, 2026The playPending
The Skill Pack. Package ONE workflow you know cold (50+ runs, definable trigger, checkable output) into a sellable digital product: a spec sheet + a pre-loaded context file + a 10-case eval (with two integrity traps), default paste-in tier, sold at $99-$249 through Lemon Squeezy or Gumroad with the dated eval table as the sales page. Modest case: 3-5 sales in month one (about $650 net at $149); first dollar = one sale inside the week via one teardown post plus two niche communities.
How we grade itReader-reported results: did anyone ship a pack with a published, dated, model-stamped 10-case eval table and record a first sale? Secondary: does the eval-table-as-sales-page pattern show up in the wild from other sellers?
Where it stands(archive issue 4, source week Jun 24-Jul 2; published 2026-07-10). The first real reader win lands in a future Reader Win slot, quoted and dated. We will not invent one.
The Jul 6 issueJul 10, 2026Market callPending
MARKET CALL (the shelf arrives): Packaged workflow assets (skills, agent packs, prompt-plus-context bundles, MCP bundles) become a recognized product category with a real shelf: by June 30, 2027, at least one major platform (an AI lab's first-party store, or a top digital-product marketplace adding a dedicated category) ships a discovery/marketplace layer where independent sellers can list packaged workflow products. Corollary the issue states: the uncrowded direct-sales window is temporary, and early sellers with published eval tables hold the positioning advantage when the shelf opens.
How we grade itPlatform announcements through Jun 30, 2027: a first-party skills/agents store open to independent paid listings, or a major marketplace (Gumroad-class or bigger) launching a dedicated packaged-AI-workflow category. Hit = the shelf exists by the date; Miss = it does not.
Where it stands(archive issue 4, source week Jun 24-Jul 2; published 2026-07-10). Basis: NLW's Capability Overhang Playbook naming skills/context packaging as the operator response (AI Daily Brief, Jun 28); buyer willingness-to-pay evidence via Lenny's Newsletter (Jun 30); no marketplace existed in-window, stated honestly in the issue.
The Jul 6 issueJul 10, 2026Market callPending
MARKET CALL (the model layer stays rationed): Frontier access remains rented, rationed, and repriceable through 2026: before December 31, 2026, at least one more frontier-model access event occurs: a suspension/outage of a top-tier model, a government-gated or vetted-cohort-only release, or a rationed rollout (usage caps or paid-tier-first access) at a major lab. The issue's mechanism claim: each such event re-proves that the packaged layer above the model (specs, context assets, evals) is the durable one.
How we grade itVendor and regulatory announcements through Dec 31, 2026: any top-tier model suspension, government-restricted release (Mythos-5-style vetted cohorts), or rationed rollout (Fable-style caps/tiers). Hit = at least one such event after Jul 2, 2026 and before year end; Miss = frontier access stays fully open and unrationed.
Where it stands(archive issue 4, source week Jun 24-Jul 2; published 2026-07-10). Anchoring receipts: Fable 5 dark Jun 12-Jul 1, returned rationed with a 50% usage subsidy through Jul 7, later extended to Jul 12 (AI Daily Brief Jul 1; Ben's Bites Jul 2); Mythos 5 cleared for ~100 vetted orgs (TechCrunch, Jun 26); GPT-5.6 in a government-handpicked preview (Moonshots, Jun 29); Sonnet 5 intro pricing expiring Aug 31 (TechCrunch, Jun 30).
The Jul 13 issueJul 13, 2026The playPending
The Deal Rep Sprint. Sell self-funded acquisition searchers a fixed-scope screening service: 50 public business-for-sale listings run against one frozen written buy box through a cheap-worker + frontier-reviewer harness, delivering a full pass/reject/unknown ledger (rule code + quoted excerpt per reject), a human-checked shortlist of 3-5, and five broker questions each. Pricing: $200 pilot (15 listings, 3 days), $600 sprint (50 listings, 7 days), upfront, no success fees. Modest case: ~$20 model/tool spend and ~10 hours per sprint; two accepted sprints/month = $1,200.
How we grade itReader-reported results: did anyone land a paid pilot or sprint using the kit's demo-sample + outreach-note mechanic? Secondary: does the rejection-ledger deliverable (rejects shown, rule-coded, source-linked) appear in the wild from other screening sellers?
Where it stands(live issue 5, source week Jul 4-10; sent 2026-07-13). The first real reader win lands in a future Reader Win slot, quoted and dated. We will not invent one.
The Jul 13 issueJul 13, 2026Market callPending
MARKET CALL (the platform bundles the first pass): by June 30, 2027, at least one major business-for-sale marketplace or listing platform ships a native AI buy-box screening feature: buyer-defined criteria applied automatically across its listings (saved-criteria AI matching, auto-screening, or an agent that produces a shortlist). Corollary the issue states: when platforms bundle first-pass screening, raw extraction becomes worthless and the durable pieces are the frozen-criteria discipline, the source-linked rejection trail, and the client relationships.
How we grade itProduct announcements and feature pages of the major business-for-sale marketplaces and broker platforms through Jun 30, 2027. Hit = at least one such native feature ships by the date; Miss = none do.
Where it stands(live issue 5; sent 2026-07-13). Basis: near-frontier task costs at cents (Grok 4.5 $0.31/task, AI Daily Brief Jul 9; Muse Spark 1.1 benchmark run $0.92, AI Daily Brief Jul 10) make the feature cheap to build, and ChatGPT Work (Jul 10) shipped the general-purpose agents-run-your-research pattern to everyone.
The Jul 20 issueJul 20, 2026The playPending
The Consent-Cleared Call Pack. Build a 60-call, human-recorded, rights-cleared acceptance pack for auto-repair voice agents (6 intents, scenario cards, transcripts, pass/fail rubrics, signed contributor releases) and license it nonexclusively to voice-agent builders: $199 founding license (first five buyers), $349 standard, five-call free sample as the sales asset, 20-named-buyer validation before recording the full set. Modest case: month one, two founding licenses = $398 against ~$235 of contributor pay and fees (~$163 gross before labor); month two, two standard licenses = $698 with production cost sunk.
How we grade itReader-reported results: did anyone validate a 20-buyer list, sell a founding license off the five-call sample, and report the receipt? Secondary: does the consent-cleared acceptance-pack deliverable (rubrics + rights manifest) appear in the wild from other sellers?
Where it stands(live issue 6, source week Jul 13-20; sent 2026-07-20). The first real reader win lands in a future Reader Win slot, quoted and dated. We will not invent one.
The Jul 20 issueJul 20, 2026Market callPending
MARKET CALL (the platform ships the eval shelf): by July 31, 2027, at least one major voice-agent platform or AI lab ships a first-party acceptance-test or evaluation-pack marketplace/library for agent buyers (buyer-facing test suites or third-party eval packs as a product surface, not just internal evals). The issue states the odds at better than even within a year. Corollary the issue states: when platforms bundle generic testing, the durable pieces are domain truth from named humans in one trade and clean commercial rights.
How we grade itProduct announcements and docs of the major voice-agent platforms (ElevenLabs-class, agent-platform vendors) and AI labs through Jul 31, 2027. Hit = at least one such buyer-facing eval marketplace/library ships; Miss = none do.
Where it stands(live issue 6; sent 2026-07-20). Basis: OpenAI already runs an internal adversarial red-team model, GPT-Red (Ben's Bites, Jul 16); Grok ships a no-code voice-agent builder (The Rundown, Jul 15); the missing layer is buyer-facing acceptance testing.
The Jul 20 issueJul 20, 2026Market callPending
MARKET CALL (the frontier stays perishable): from July 20, 2026 through January 20, 2027, at least 10 more frontier-class model releases occur (a sustained cadence of one per ~18 days or faster among the frontier labs, US and Chinese), keeping capability churn high enough that model-specific bets keep depreciating and re-testing stays a permanent cost. Hit = 10+ frontier-class releases in the window; Miss = the cadence collapses back toward the 2025 rate (fewer than 10).
How we grade itFrontier release announcements and the Artificial Analysis index roster through Jan 20, 2027, as tracked by the trade press (AI Daily Brief, The Rundown, Innermost Loop class sources).
Where it stands(live issue 6; sent 2026-07-20). Anchors: 13 frontier releases since mid-April 2026, one per 10 days, vs 8 in all of 2025 (Moonshots, Jul 19); four in eight days in mid-July (Innermost Loop, Jul 18); Alex Wissner-Gross's regression points toward continuous releases by January (Moonshots, Jul 19). Update 2026-07-25: the cadence held through week one of the window: Claude Opus 5 shipped (Lenny's Newsletter review, Jul 24) and Moonshot's Kimi K3 announced full open weights for Jul 27 (Moonshots, Jul 24); on pace.
The Jul 13 issueJul 13, 2026Market callPending
MARKET CALL (the two-tier price structure holds): the frontier gets dearer while the floor drops, and the structure persists: before December 31, 2026, (a) Anthropic's Fable premium credit pricing ($10 in / $50 out per M tokens, effective ~Jul 12, per Innermost Loop Jul 8-10) is NOT cut by 50% or more, AND (b) at least one additional near-frontier model launches with headline cost-per-task positioning at or below roughly a third of concurrent frontier task cost. Hit = both legs hold; Miss = either fails (frontier premium collapses, or the cheap-launch cadence stops).
How we grade itAnthropic's published pricing through Dec 31, 2026; new model launches and their cost-per-task positioning on published comparisons (Artificial Analysis-style task-cost tables, as cited by AI Daily Brief and The Rundown).
Where it stands(live issue 5; sent 2026-07-13). Anchors: four cost-led launches inside three days (GPT-5.6 Sol/Terra/Luna, Grok 4.5, Muse Spark 1.1, SWE-1.7; AI Daily Brief Jul 9-10, The Rundown Jul 10) while Anthropic repriced Fable UP (Innermost Loop, Jul 8-10). This is the Rep Economy's supply-side premise: if the spread structure collapses, the issue's mechanism weakens and the board should say so. Update 2026-07-20: leg (b) is satisfied early: Kimi K3 launched at $0.94 per Artificial Analysis benchmark task vs Fable 5's $2.75, roughly a third (AI Daily Brief, Jul 17). Leg (a), the Fable premium holding, stays open through Dec 31, so the row stays Pending. Update 2026-08-03: the floor kept dropping. GPT-5.6 Luna's output price was cut roughly 80% on Jul 30, from $6.00 to $1.20 per million tokens (CNBC, Jul 30). That is a cheap-tier repricing, not a frontier-premium cut, so it reinforces leg (b) and leaves leg (a) untouched. Row stays Pending through Dec 31.
The Jul 27 issueJul 27, 2026The playPending
The Model Receipt Packet. Sell a 5-30 person vertical SaaS vendor with one AI feature and a live enterprise security/procurement review a five-day, fixed-scope documentation packet (one-page data path, model+subprocessor register with dated evidence, route-receipt log fields, honest-gap buyer FAQ, change card): $350 pilot ($175 before the interview, $175 at delivery), $700 standard after two accepted pilots, optional $125 quarterly refresh. Modest case: one pilot, 8-10 hours, at least $320 gross before time and taxes. Kill rule stated in the issue: 20 qualified contacts, fewer than 3 conversations, 0 paid pilots = stop.
How we grade itReader-reported results: did anyone qualify a live review, sell a $350 pilot off the FieldNote-style specimen, and report whether the packet moved the buyer review? Secondary: does the honest-gap FAQ shape (status: unverified / owner / next decision) appear in the wild from other vendors or sellers?
Where it stands(live issue 7, source week Jul 19-25; scheduled send 2026-07-27). The first real reader win lands in a future Reader Win slot, quoted and dated. We will not invent one.
The Jul 27 issueJul 27, 2026Market callPending
MARKET CALL (the router ships the receipt): by July 31, 2027, at least one major AI router or gateway (OpenRouter, Ramp, Cursor, Vercel, Stripe, or an AI lab's first-party gateway) ships a customer-facing per-request model-provenance feature: a queryable or exportable record of which provider/model served each request, surfaced to the router customer's own customers or auditors (not just an internal usage dashboard). The issue states the odds at better than even within a year, and frames it as the play's window: manual receipt packets are the interim product until provenance is native.
How we grade itProduct announcements, changelogs, and docs of the major routers/gateways and AI-lab gateways through Jul 31, 2027. Hit = at least one such buyer-facing provenance feature ships; Miss = none do.
Where it stands(live issue 7; scheduled send 2026-07-27). Basis: Stripe reportedly in talks to buy OpenRouter at ~$10B vs $1.3B in May (WSJ, Jul 23) with the acquisition framed as buying the metering/billing layer of inference; Ramp/Cursor/Meta all shipped or building routers the week of Jul 20.
The Aug 3 issueAug 3, 2026The playPending
The Instruction Diet. Sell a 2-20 person AI development agency a quarterly instruction-debt scan: a local scanner that inventories every AGENTS.md, CLAUDE.md, Cursor rule and skill file across their client repositories and flags duplicates, contradictions, dead paths, obsolete model names and oversized always-loaded examples with exact line references, then prepares a cleanup patch a human reviews. $249 per quarter per workspace, about $230 net, renewing.
How we grade itAt least one reader reports a paid workspace at or near $249/quarter, and at least one renewal at the 90-day mark. Kill signal: zero paid workspaces from the first 50 qualified contacts.
Where it standsLogged before the outcome is known. Buyer sourced from Clutch's public AI development directories plus a per-company check that they actually run coding agents internally.
The Aug 3 issueAug 3, 2026Market callPending
MARKET CALL (the vendor absorbs the cleanup): by August 3, 2027, at least one major coding-agent vendor (Anthropic, OpenAI, Cursor, GitHub, or an equivalent) ships native CROSS-REPOSITORY instruction staleness or conflict detection, beyond the single-workspace doctor-style command Anthropic had already shipped as of July 2026. Stated at better than even odds.
How we grade itA shipped, documented feature that scans instruction files across multiple repositories for stale or conflicting rules. Single-repo linting alone does not count, and neither does the July 2026 doctor command that prompted this call.
Where it standsDeliberately falsifiable and deliberately uncomfortable: it names the risk that this play's product shrinks to a vendor feature.