Command Palette
Search for a command to run...
Scores are model reads on a synthetic news corpus — not validated against real prices. Aftermath verdicts come from corpus follow-ups, not market data.

Market-impact playbook — running notes

Living doc. The heuristics + worked examples we use to tag stock-market impact of news, compiled as we go. Pairs with the study (market_impact_study.md) and the judge prompts (tools/market_impact/).

The core heuristic (keep it simple and direct)

For ANY event, ask two plain questions first:

  1. Who benefits — whose near-term earnings/demand get better?
  2. Who suffers — whose get worse?

Then follow the chain to the second-order effects. You do not need a ticker named in the story — extrapolate it. Name both sides in tickers where you can.

Two force-multipliers on top of the core heuristic

  • Weight outsized moves in small / niche pure-plays. A broad scare barely moves a mega-cap but can materially move a focused name. A niche beneficiary getting an outsized boost is exactly the high-value, non-obvious tag we want (e.g. a produce scare → Scotts/AeroGarden). Don't just tag the obvious big loser; find the small winner.
  • Classify the scope of the event — it changes the ticker list:
    • one-stock story — idiosyncratic to a single brand (Chipotle's own 2015 E. coli; an IBM miss). Big move in that name, limited spillover.
    • sub-sector story — hits a category (a supplier/produce scare hits the whole salad category; a memory shortage hits all memory makers).
    • macro / index story — Fed path, oil supply shock, a broad AI selloff.
    • The same theme can be any of the three: Chipotle-2015 was one-stock; an unsourced multi-state produce outbreak is sub-sector.

The forward-looking test (the real magnitude question)

Don't ask "is this surprising" (too abstract). Ask: if this is true, and if it continues/scales, is it a big deal? Can it get a lot bigger (or smaller)? The signal is in the trajectory and the ceiling, not the level.

  • New mechanism that can compound = high signal (a cost curve dropping, a capability improving, a competitor/substitute that just entered and can take share). It changes future earnings.
  • Milestones are the opposite — "hit 1M subscribers / $1B revenue / record cap" is backward-looking proof of stuff already done; no forward implication → discount.

Failures > launches; negative (real) > positive (PR).

  • A product failure / delay / recall / outage is more meaningful and more reliable than a launch. Launches are company PR/self-promo → discount. Failures are true and forced → they carry real forward-earnings info.
  • Asymmetry: a company announcing its own good news → weight down (PR risk). Something bad that leaks, or a delay that's forced → weight up (harder to fake). A competitor's unexpected move changes the incumbent's future earnings even though the incumbent said nothing. (v6.)

Seed stories — small/boring now, big weeks later (the inverse of milestones)

The tiny, novel, low-author story on a new mechanism often out-predicts the huge-number headline. Track these; some of today's a=3 breach or a=1 supply note is the seed of a real move.

seed (date, authors) payoff weeks later the move
"DeepSeek V4 delayed, shifting to Huawei chips" — Apr 18, 1 author Kimi K3 "$850B US tech selloff" (Jul) NVDA / closed-lab repricing
"SK Hynix may cut Nvidia HBM4 shipments 20–30%" — Apr 14, 1 author memory shortage → KOSPI circuit breakers (Jul) MU/SOXX/EWY
"Goldman sees data-center power +220% by 2030" — Apr 14, 1 author AI-power bottleneck; Oracle NM delay; NY moratorium (Jul) CEG/VST/GEV (nuclear/IPP/electrical)
"Cyclospora, seek source, 843 cases" — Jul 11, 6 authors FDA confirms Taylor Farms; Taco Bell/Walmart pull — 1 week YUM / SG down, SMG up
"Spirit Airlines faces liquidation within days" — Apr 15, 3 authors Spirit shuts down (May) airline survivor capacity (UAL/DAL/ULCC)
"Eli Lilly's Retatrutide shows ~30% weight loss" (Phase 3) — Apr–Jun, 4–6 authors Novo −75% from peak; Lilly ~60% GLP-1 share; CagriSema fails to beat tirzepatide NVO down / LLY up — the competitor's trial data was the real signal, not Novo's own "Wegovy lifts outlook" PR (a=10)

Canonical "competitor move > incumbent PR" case (the retatrutide/Novo lesson): the flag-worthy story was a competitor's trial result (Lilly retatrutide ~30% weight loss vs Wegovy ~15%), quiet at 4–6 authors — reliable (trial data), forward-loaded (structural share threat). The trap was Novo's louder, PR-flavored "Wegovy lifts outlook" (a=10) and the minor "China generic delay." When judging what hurts an incumbent, look for the competitor's advance and the incumbent's own trial failures, not the incumbent's positive guidance.

Worked examples (event → winners / losers → precedent)

event winners (up) losers (down) scope precedent / note
Oil / gas price spike energy XLE, refiners (VLO/MPC/PSX), "stay-at-home" mild airlines (JETS; a marginal one like Spirit can fail), grocers (freight cost), consumer discretionary macro Hormuz weeks 2026; 2022 fuel spike
Fresh-produce contamination scare Scotts SMG / AeroGarden (grow-your-own), maybe frozen/canned salad-forward restaurants (Sweetgreen SG, Chipotle CMG, Cava); a named brand (Taco Bell → YUM); suppliers if named (DOLE/FDP, Taylor Farms is private) sub-sector (grocers ~neutral) Chipotle 2015: CMG −5.5% (~$1.75B), comps to −20%, −40% over the episode — but that was one-stock (brand-specific). Cyclospora is milder than E. coli, so the move is smaller absent a viral cycle. Diversification dilutes a named loser: the Jul 2026 outbreak was sourced to Taylor Farms lettuce served at Taco Bell — but the public loser YUM has KFC/Pizza Hut too, so it moves far less than a pure-play Taco Bell would. Always ask: is the named public co. a pure-play or diversified?
Open-weight frontier model (DeepSeek/Kimi/Z.ai/Inkling) China tech (KWEB) sometimes NVDA, closed-lab proxies (MSFT/GOOGL/AMZN for OpenAI/Anthropic), SOXX, QQQ macro/thesis "DeepSeek moment" — repriced NVDA ~$600B–$1T in a day, Jan 2025
Memory / HBM shortage memory makers (MU, SK Hynix, Samsung), SOXX device/PC/auto OEMs paying more sub-sector SK Hynix CEO "shortage past 2030"; Micron $50B guide
Fed turns hawkish (hikes on the table) banks (XLF), USD (UUP) SPY, QQQ, TLT (bonds), GLD, rate-sensitive/crypto macro recurring 2026 theme — voters repeatedly floating hikes on oil-driven inflation
Big-cap enterprise-IT miss memory/hardware (capex shift) the name (IBM −25%), software peers (IGV, MSFT) one-stock → sub-sector IBM Jul 2026, read as "AI eating services"
AI-capex doubt from an authority NVDA, SOXX, QQQ, hyperscalers; watch hyperscaler credit spreads macro/thesis Fed Chair Warsh "AI funding may dry up"
M&A bid target up acquirer often down on deal cost one-stock Uber–Delivery Hero

What NOT to over-tag (from calibration)

  • Routine product/benchmark announcements from Google/Apple/OpenAI/Anthropic/Meta — near-zero. They announce constantly.
  • Financing / IPO / valuation numbers by themselves — a $Nbn raise is not a capability release; usually doesn't move the public proxies. (The model moves markets, not the funding round.)
  • War rhetoric / threats with no actual supply/facility/chokepoint hit — looks big, usually already priced.
  • Single-name earnings with no sector spillover — that's small, not moderate.
  • Big-number / milestone / ranking-optics stories — the trap a low-effort read over-values. "$1T wiped/added," "X overtakes Y as most valuable company" (Apple/Nvidia/Google swapping #1), "shares fall below IPO price" (SpaceX below $135), "record valuation" with no new driver. The huge number writes its own headline but no company's near-term earnings/demand changed — it's mark-to-market, positioning, or ranking optics. THE TEST: did earnings/demand for anything actually change, or is this just a price level / ranking / big number? If the latter, discount hard. A real driver attached to the number (a miss, a supply shock, a capability release) is the mover; the number alone is not. Round-number milestones are the purest case: "crosses 100M subscribers / 1B users / $Nbn run-rate" — 99M vs 100M vs 101M is the same business; the milestone is optics. What matters is the rate of change and the surprise vs. expectations (a subscriber miss or a guide cut moves the stock; crossing a round number does not). (v5 RULE 5.)

Judge prompt lineage

  • v1 — baseline; failed by importing the day's loudest theme (cluster-bleed) and under-rating structural AI stories.
  • v2 — score the story's OWN event; magnitude can be structural (open-weight model = major).
  • v3 — single-name earnings → small; discount war-rhetoric + funding size.
  • v4 — leads with the CORE HEURISTIC above (who benefits/suffers, follow the chain, weight niche pure-plays). Current.

Under-tagging: significance vs near-term price move (the magnitude hole)

Magnitude currently measures one thing — expected near-term price move — and the judge lets that stand in for how significant the development is. Two axes are being mashed together:

  • How major is the development.
  • How large/confident is the near-term price move.

A development can be major with a muted, slow, or unmeasurable price reaction, and the judge rates those small, which throws away that they matter. Confirmed live calibration cases:

  • Google "Frozen V2" inference chip, 6–10x more efficient than TPUs (Jul 20). Judge rated small/22 with the RIGHT tickers and direction — GOOGL up, NVDA down, AVGO (Broadcom, the ASIC partner) up — the custom-silicon-erodes-Nvidia thesis, correctly extrapolated. User: rank it "a bit higher than small." The 2028 timeline and report/leak framing justify a muted price tag, but as a development it's a real data point in the biggest question hanging over Nvidia. Structural/roadmap developments with distant timelines should not collapse to small.
  • FDA reverses Taylor Farms lettuce test to false positive (Jul 20). Rated small, YUM only, "no sector spillover" — under-tagged on BREADTH. A produce-contamination reversal is a category event (YUM up, salad chains recover, SMG gives back the scare bid), same as the outbreak on the way down. Reversal of a 30-state, 1,600-case recall is a major development even if the price relief is muted (outbreak still live).
  • Domino's Q2 (Jul 20). Judge tagged DPZ down calling it an "EPS miss." Wrong framing: revenue beat, EPS ~in-line and +6.8% YoY; the only soft number was same-store comps (+0.1% vs +0.6%). Direction should be mixed. Also the cluster had 4 tweets, no price reaction — the model inferred direction off a mixed scorecard and got it wrong. Watch: small clusters with no reaction tweets → low-confidence direction.

Fix direction TBD (a separate "development significance" field vs the price-move magnitude, or stop treating a muted price move as low significance).

Judge prompt lineage

  • v1 — baseline; failed by importing the day's loudest theme (cluster-bleed) and under-rating structural AI stories.
  • v2 — score the story's OWN event; magnitude can be structural (open-weight model = major).
  • v3 — single-name earnings → small; discount war-rhetoric + funding size.
  • v4 — leads with the CORE HEURISTIC above (who benefits/suffers, follow the chain, weight niche pure-plays).
  • v6 — adds forward-looking test, failures>launches, milestone/round-number discount. Current (used for the Jul 20 and May 18 300-story runs).

Ranker is the coverage bottleneck (measured May 18 week)

Judging the top 300 vs the top 40 of the May 18–25 week: the top-40 caught 9 movers; the full 300 found 69, so 60 movers (87%) sat below the top-40 cut, including 11 of the 15 big ones. The misses were systematic — the prescore ranks by dollar-figures and chip/AI keywords, so it floats IPO/funding stories the judge then rates no move, and buries macro/rates stories the judge scores highest ("Fed hike bets harden" scored 78 but ranked 226/300; a record 30-yr Treasury yield ranked 187; a Japan bond blowup ranked 250s). The eligibility+ranking stage, not the judge, is where coverage fails. Under-weighted categories: macro/rates/Fed, M&A, and non-tech sectors (healthcare, retail, housing, industrials, insurance) — the keyword list is tech/macro-shaped.

14-week deep run (Apr–Jul 2026) — the ranker-fix evidence

Judged the top 300 stories/week for all 14 weeks (4,199 judgments: 306 big, 681 medium, 713 small), bucketed each week into major-movers / single-stock / small-but-novel, then ran an aftermath pass (corpus follow-up 1–2 weeks later + web where the ticker is real). Report: docs/market_impact_weekly.html (clickable overview + per-week tabs). Generators: tools/market_impact/build_weekly_report.py, find_followups.py, build_aftermath_prompts.py.

  • Aftermath materialization: of 247 bucketed stories with any follow-up signal, 52% materialized (128), 24% faded, 25% mixed; 82 had no follow-up. Web confirmed a few real ones (Amazon–Globalstar $11.6B deal → GSAT surged; Micron NAND rally).
  • 64 "should have ranked higher" wins — small/medium stories the live ranker put below the top-40 cut that our judge flagged AND that materialized 1–2 weeks later. This is the concrete payoff: the non-obvious movers the current ranking misses. Confirmed archetypes:
    • FDA cyclospora in Taylor Farms lettuce (ranked #78) → real recall; YUM/CMG/SG all fell. The user's lettuce case, validated end to end.
    • FDA proposes excluding Novo/Lilly weight-loss drugs (moderate, buried) → NVO +9% weeks later.
    • Lululemon names ex-Nike exec as CEO (small, #265) → LULU −13% next day.
    • India raises gold/silver import tariffs (#99) → spot silver −5%.
    • Uber / Delivery Hero M&A (buried) → takeover advanced.
  • Takeaway for ranking: feed the judge's magnitude/score back into the story ranker so novel, off-mega-theme, forward-looking stories (product failures, recalls, regulatory proposals, custom-silicon, M&A approaches) get lifted — not just earnings and the loud oil/chip themes. The ranker currently spends its top slots on IPO/funding dollar-figures the judge rates no move.

Bellwether contagion — when a big name breaks, three rings move (and the move itself is the signal)

A large move in a bellwether (Tesla, SpaceX, Nvidia) reads through in three rings, and you weight them by the REASON for the move:

  1. Supply chain / thesis inputs — Tesla down on weak EV demand → batteries & lithium (Panasonic, LG Energy, CATL, Albemarle, SQM), EV power semis (ON, Wolfspeed, STMicro), charging (ChargePoint, EVgo), EV peers (Rivian, Lucid).
  2. Sector comps — SpaceX down → the listed space set re-rates (Rocket Lab, AST SpaceMobile, Planet, Intuitive Machines, Redwire).
  3. The complex + sentiment — Tesla/SpaceX both breaking = "the growth/innovation trade is being repriced." This ring has zero direct link — the move itself is the read. Severity scales with SPECULATIVENESS: hardest hit are profitless/story stocks and ARKK (Cathie Wood's ETF concentrates exactly those, and Tesla is a top holding → forced selling); crypto too (both hold BTC; risk-off hits COIN/MSTR). NVIDIA gets a SMALLER hit — real earnings anchor it, so it's less of a pure-sentiment casualty than the speculative long tail.

The fork: weak-earnings/demand cause → ring 1 (supply chain) hardest hit. General-disenchantment/multiple-compression cause → ring 3 (contagion) dominates, and defensives/value/utilities catch the rotation.

The other side (who's UP): robotaxi-threatened incumbents get relief (Uber/Lyft when Tesla autonomy looks less imminent), Starlink rivals (Viasat/EchoStar, competitive relief though contagion may drag first), legacy autos on share-shift IF the weakness is Tesla-specific execution.

Read the story sequence, not the tally — handle conflicting updates

The ranker sorts by how much coverage a story got (author/tweet volume). That silently over-weights the older story: a crash plays out over three days and piles up authors, while the update that reverses it is fresh and thinly covered, so it ranks below the thing it just contradicted. Volume is popularity, not recency. This is how the outlook went stale on Korea: it led with "-40% crash" (heavily covered, older) and buried "Samsung +18%, futures limit-up" (the newer reversal). User: "you had info that Korean stocks already rallied but you didn't include it in your logic and over-valued older information" and "look at news in sequence, not just scramble them and take 'the biggest story' since often the older story has more authors than an update."

The rules:

  1. Order events in time before forming a view. Sort the relevant stories by timestamp and read them as a thread, not a bag. The last real update on a name wins over an earlier, louder one.
  2. Don't just sort — weight by recency. Ordering alone isn't enough; a newer update should actively outweigh an older story on the same name, even when the older one has far more authors. Recency and sequence position are inputs to the score, not just the display order.
  3. Discount repetitions of the same story. Many authors piling onto one event is popularity, not N independent events. Collapse a cluster of "Korea crashed" retellings toward a single data point so raw author-count stops inflating the older story over its own update.
  4. When two stories on the same name conflict, resolve by recency and state the arc. Don't average "crashed" and "rallied" into mush, and don't pick the louder one. Say what happened in order: fell hard first, then reversed, here's where it ended. The direction of the latest update is the current read; the earlier move is context.
  5. For the current-prediction / outlook, pull EVERY story on a stock we hold a view on — even small ones. If we have a view on MU/Samsung/Korea, a low-author update on that exact name must be included regardless of its impact score. The very_unlikely/milestone filter is for building the movers digest, not for reading the outlook: a thinly-covered "+8% futures limit-up" is exactly the update that flips the near-term read, and the volume ranker is what buries it. Query by ticker/entity for our watched names, don't rely on the top-N cut.

This is the flip side of the price-milestone rule: there, coverage over-credited a non-event; here, coverage over-credits the older of two real events. Both come from treating volume as importance.

Note when the story IS the reaction — the move is already digested. A headline of the form "shares fell N% on weak guidance" / "stock drops on the miss" is reporting a price move that already printed. The drop is in the price. This is a real driver (unlike a milestone, which has none), so the view isn't worthless — but the only thing left to trade is continuation: whether the driver keeps pushing after the initial gap. That is a weaker, second-hand read than a fresh catalyst that hasn't fully played out. Three tiers to keep straight: (1) fresh catalyst, still developing — best (a Hormuz closure that's ongoing, a shortage getting confirmed); (2) reaction story, move digested — useful only as a continuation bet (Corning −14% on the guide, the megacap earnings gaps, an airline guide-down that already sold off); (3) milestone / price level, no driver — worth nothing forward. When a negative view rests on a tier-2 story, say so — it's still useful, but flag that the first move is already gone.

Open threads / to watch

  • Cyclospora produce outbreak (31 states, 3,000+ cases, source being traced) — if the /debug update names lettuce, apply the produce-scare row: SG/CMG down, SMG/AeroGarden up. Watch for the source-identification version. Update Jul 20: FDA reversed the Taylor Farms lettuce test to a false positive; treat the reversal as a category event, not YUM-only.
  • Encode the winners/losers + niche-outsized reasoning into the ranking boost and the story-text "market read" line.
  • Accuracy is still unmeasured — every result so far is model-vs-model (headline-only vs with-tweets), never scored against real prices. The one validation that would make "better" more than an assumption: pull actual price moves for the named tickers around each story's timestamp.

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive