Chammarychammary

The Ads Business Model Will Die & Lessons from Working with Elon at Twitter | Parag Agrawal

20VC with Harry Stebbings · 1:08:54 · 2 days ago

Agents will use the web 1,000x more than humans, so the web's technical and payment layers must be rebuilt for non-human consumers; Parallel argues its search infrastructure can become a material part of AI inference economics . This is a startup projection, not a demonstrated market .

Key points

  • Agent queries are fuller sentences with either ~100ms latency or long background compute, and outputs are tokens/files rather than blue links . The task narrows a trillion URLs to ~1,000 context tokens . Parallel claims same-quality search at 1/120–1/150 of human-stack compute, using turbo for voice agents and advanced for background agents . It says it sells at $1 per 1,000 searches versus $7–$14 competitors . That price could fall another 10x as low-end agent search expands .

  • Agrawal expects search demand to persist as models improve because parametric memory is lossy; a model can remember famous facts but may miss specific or recent details . He separates smarter reasoning from memorized facts, saying smaller efficient models can lose more parametric memory while preserving reasoning . Frontier models should keep getting bigger when incremental quality is valuable, while smaller models reach older frontier performance levels . That supports a growing agent-search market rather than shrinking one .

  • Amazon's ad business is now larger than its e-commerce business, and agent-mediated purchases remove the human-facing ad impression . That is why display ads do not transfer to agent workflows . Parallel proposes paying publishers a variable fee when an agent benefits from their content, using models to estimate the owner's marginal contribution to the answer . The same logic replaces seat licenses: data providers should be rewarded when agents use their data, not only by human seats .

  • Agrawal sizes Parallel's market as 5–20% of GPU spend on agents, after removing training and media generation from total AI GPU spend . If Fireworks reaches $2B while inference revenue is $200–300B, the implied search market is $10–60B . A $100B outcome would require large share and fast revenue growth . Failure modes include agents underperforming, weak execution, losing content access, or customers not trusting the stack .

  • The next stage is not occasional queries but persistent agents that wait for events; Parallel wants to crawl continuously and trigger on changes instead of having agents poll every 6 hours . It claims event-triggered monitoring can use 1/100 to 1/1,000 of the compute used by periodic search for the same alert . The broader shift is from pull search to a push web, where long-lived context lives in the search infrastructure . Agrawal believes always-on agents will become normal, but most people do not yet trust agents with money, email, or logins .

Step by step

  1. Set the search mode by latency budget: use the fast low-cost mode for voice or interactive agents, and the advanced mode when the model can spend seconds or longer .
  2. Tell the search API which model will consume the results when known, so the stack can optimize token density for that model .
  3. For slow conditions, use a monitor-style API that triggers alerts on web changes rather than re-running search every 6 hours .
  4. Price publisher or data-provider access by marginal contribution to the agent task, not by human seat count .
  5. Treat model misbehavior as an alignment and containment problem: secure RL environments and assume post-alignment models are not adversary-proof .