bitter lessons + 80m agents

ยท Zi Wang ยท 1 min read

Z / Runner ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ

๐Ÿƒ๐Ÿปโ€โ™‚๏ธ manus lessons: 1). Old era internet startups favor "artist", all about viral growth & zero marginal cost (peak used instagram as example, serving 1 vs 1m users). Ai opposite extreme,even when taken model improvements, token price drop. It requires a deep understanding of infra, unit economic and personality traits (factory manager, cool logical engineer).

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ peak repeated nth times about "healthy" โ€“ no hype, no extreme passions, no unhealthy habits (drinking, smokingโ€ฆ). Xiao hong: rational, boring, seasoned = greatest advantage.
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ bygone era.
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ unit economic & technical delta is staggering between a shallow mobile app vs. GPA (general purpose agent). Software scales for free are no longer true. Ai is new manufacturing. "Bill gates" write once, ship it a billion times for free, no longer holds. Manus' entire fiscal & ops took that into account months into the development. Compared to saas, cost management is on db i/o & hosting, ai hosting rises linearly, especially at inference tokens / session.
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ Every user interaction, complex cot, multi-step reasoning, ai model continues to accrue expensive computational resources. Peak's main point was that unlike google indexing, ai generated content/data results are NOT recyclable. (note to self: assumption may be true today, but may change). Again, peak used the instagram analogy; a multi-step general agents is extremely costly, even with $150m arr. Grow at all costs (ventured backed) is dangerous and will kill most ai startups.
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ cost seems to be the constant in his reflection/analysis.
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ personality traits: being an artist is fundamental liability, manic, overly passionate, strong headed: these traits may serve well in the context of deep research & viral growth, but it is useless in scaling a financially sensitive ai application business.
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ haha, "Body 2025".
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ Can't assume and depend on stable technical platforms/markets (e.g. ios mobile); ai environment is just too violate, models break alignment, costs spike, research breakthroughs (e.g. gpt 3, multimodal, long-horizon, sparse attention) all the examples peak shared. Roadmap = obsolete overnight. A tortured artist/ceo may panic, or lash out at engineering team, make rash emotionally driven pivot based on fear, excitement. Xiao hong makes it work, focus on long-term decisions, not reacting to short-term chaos. Stability is the new disruptor. Gone are the era of crazy genius, dreamer, need grounded realism, humility and emotional resilience. (note: there were so many names blipppeeed out during the interview, peak appears to have direct access to many silicon valley based ai founders).
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ almost 30+ mins on the topic of costs & financials: e.g. bring down inference costs by 30% in one quarter, securing bulk discounts even though they had sacrifice speed and initial prompting costs. It was interesting to hear the accountants overriding co-founder/chief scientist. Also monica.im was a huge cash cow, it enabled risker "2nd product".
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ piggy back on your insights; had similar gut punching feeling, when the times were good, i didn't have a deep appreciation or even more bluntly "deep gratitude" on the favorable market condition, instead always chasing after the bigger market whale.
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ embrace the "pain cave", for body, for mind, for buildโ€ฆ
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ one of my favorite lessons learned: don't rely on a monolithic decision tree, GPA (Goal, Priority, Alternative).
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ yay, z 2026 MIND (hence, shedding all the intellectual weight, "brain diet".
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ Peak's take on BDFL (benevolent dictator for life) useful, we always remember it for the name sake of guido van rossum.
  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ will expand on this core topic in the parallel section below: x% of lessons are directly applicable.

๐Ÿƒ๐Ÿปโ€โ™‚๏ธ parallel w/ ai(s)(j):

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ tl;dr replicate 2025 body for mind & build.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ but i get the point, a worthy exercise to state the positive (will tack it on).

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ ..

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ point well taken, but when you have a 3 hr podcast to listen and need to kill it on 6 am exam, you have to pull the old "trick" of burning the midnight oil. The telegram message weighted heavily in my mind; had to know and understand what lessons we shared are erroneous & curious to know more about the insights he shared. Rare opportunity in the history of big tech acquisition, where the technical founder is open & transparent. Mostly it is just business none-sense.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ Opertional cash flow = bold, risky and rational bets. Monica = low burn, high utility data trap. Ai native products 101?

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ manus never chased (respond yes). Avoided local, noisy chinese market. Go after the prosumer, premium US market; index on "willingness to pay".

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ identified tasks that are valuable = "freelancers" who make $ by spending $ (aaron is a great example).

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ yay, n of 1 != business.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ one off; google & openai have both failed the travel agent tasks. It was announced during google i/o, but completely missed the launch date by 3+ quarters. No agent today is good at this type of workflow; it maybe too complex to solve using today's model.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ Are some problems just too hard to solve using ai? I know the surface level answer is always, ai can solve all known problems, but is it true? I tried to use the top 4 models this year for planning a simple trip, all failed. Aaron mentioned using manus for his international trip, it didn't work.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ it is a problem space that google will eventually crack, w/o oversharing, google is planning to integrate gmail histories directly w/ travel planning.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ hence, i can see why YC funded so many mobiel app screenscrapping startups in the last 2 batches. (per shamir's vc holiday party, more than 20%+ of the last batch are doing just that).

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ manus = general purpose agent; it focused only on high-valued workflows. Are you proposing ai native application must follow the exact path?

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ peak mentioned the entire team realized even sales/marketing folks were using claude, cursor to solve their problem b/c none-technical people want to "automate" their workflow and want tools to make their job less tedious. The sandbox idea was interesting b/c it solves the security problem (as his co-worker/friend installed agent that disabled his network access).

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ Peak's most painful lesson = magi. Premature innovation. His tale of knowledge graph brought back PTSD. i too was excited w/ google acquired freebase, and obsessed over SPO [subject-predicate-object], ML based knowledge vaultโ€ฆ maggie got trapped into a vertical build up that didn't scale. Took him years to realize he needed his own chip, model, data, productโ€ฆ maggie was never going to work. GPT 3 & 4 made it obvious, but by then it was too late. Peak shared the lack of focus killed maggie; going too deep is almost as bad as going too wide. I rarely thoughtful about it this wayโ€ฆ given we both agree it is about going the nth degree of depth. (although in a slightly different context). Going full vertical may create infinite distraction?

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ killing the first born: scrapped their fully built ai browser project (reaction to arc, atlas or scared of chrome). His reasoning: it didn't pass the internal taste test "cool".

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ double ditto on ai(s). Hence, i do need to rush the quick & dirty version out asap. Dicking around w/ too many features and over optimizing on ai stack. I do like your extreme courage on launching/sharing the daily builds. Something i always had trouble w/ due to self-imposed must be "great" before sharing.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ Mini parallel 1).: building a browser forces the startup to compete directly with tech giants. NO distribution = certain death. It forces the team to go after an application layer that they can win. 2). github "everything added = dilutes everything else".

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ although, we do need to step back a little: i told at lest 30+ ppl about manus, so far only 1 decided to try it for free. Everyone still uses chatgpt/gemini. To be fair, i only started pay for 2 accounts ($200/month) in november and december b/c i was running out of token credits when i started to prototype ai(j).

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ cautionary tale for both of us: vanity. Peak has a natural tendency to follow technically challenging projects. Speaking for myself, i have similar dark pattern, always more interested w/ things i don't understandโ€ฆ the less i know, the more excited i get. Which leads to infinite divergence of thoughts. Peak did pull back from training their on LLMs, as he correctly predicted that "intelligence" will become a commodity. Obsessing over LLM breakthrough = buying a lottery ticket. An application company's life/death should NOT be connected w/ LLMs solving big problems. It cede control of the product roadmap to external LLM vendors.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ My biggest fear, but also my biggest conviction: fear = top 5 LLM builders will build general purpose applications on top of their own LLMs (peak shared similar obvious/obvious). They have the advantages: cost, inference speed, reasoning depth, and access to context/data. Peak's insight = models are plug/play, interchangeable components (he borrowed this from cursor); and winning on distribution is too costly. The only ai moat = internal evaluation & benchmark.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ why did manus sell to meta? obv/obv $2.5b = big number. BUT, alternatively, the path they pioneered maybe the same as perplexity. Eventually, they will get swallowed by the big model companies. Can you have a general purpose agent (application level) without controlling your own LLMs?

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ we will never knowโ€ฆ

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ top quote/insight of the year: adam moesseri (head of instagram) "authenticity is becoming infinitely reproducible, labeling "ai generated" is useless; instead label "this is real" is more immediate. Facts are facts, facts have no feelingsโ€ฆ hence ai(s) has its place on grounding reality.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ it may require a completely different perspective on grounding truth; ai can fake imperfection just as well as faking cinematic/pixel perfect videos.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ RLI (remote labor index): can agent perform high-value task effectively, so premium customers are willing to pay. Outcome must = measurable $.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ manus is already less useful as of december, the meta d&d killed the paywall scrapping and probably 100+ other use cases. Creating value on the margin, away from the spotlight is easier (pirates) vs. accumulating value under the public scrutiny (navy).

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ don't want to rathole our flow: longevity operates on the margin, pseudo science, everyone hates ivory tower doctors, sickcare sucks, influencer opinions > decade long research.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ Yes, narrow set of workflow automation + continuous horizon.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ what you are building ai(j) has the potential, but we have to pass the RLI test. Are peop;e willing to pay "$2400"/yr for the jobโ€ฆ other than BUILD (manus's productivity focused workflow tooling), health is the only thing (most people) are willing to pay.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ fx_health ($4300 yr), dexa ($70/quarter), whoop ($400/yr), apple fitness+ ($120/yr, the drive podcast ($200 / yr) โ€ฆ and that doesn't include all the LLMs ($5000 / yr) paid in 2025โ€ฆ most people are unwilling, unable, and too lazy to do the work.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ $100 billion market at least, but yelp already does this well enough (qualified). Small sample size, we had to get our furnace replaced ($9k job), yelp sourced 20+ local hvac companies and screened their reviewsโ€ฆ

Stephen / Basketball ๐Ÿ€

  • ๐Ÿ€ "top 3 insights that disprove a startup founder in silicon valley previously building mobile, social, crypto products"

  • ๐Ÿ€ no more tgi-champagne.

  • ๐Ÿ€ "Manus ็š„ token ๆถˆ่€—้‡ๅทจๅคง๏ผŒๆ‰€ไปฅๆˆ‘ไปฌ่‡ช็„ถๆ˜ฏๅ‡ ไนŽๆ‰€ๆœ‰ๆจกๅž‹ๅŽ‚ๅ•†็š„ๅคดๅ‡ ๅ็š„ๅฎขๆˆทใ€‚ๆ‰€ไปฅๆˆ‘ไปฌ่ทŸ่ฐทๆญŒ Deepmind ้ƒฝๅพˆๆทฑ็š„ๅˆไฝœใ€‚.. ๆˆ‘ไปฌๅœจๅ„ไธชๆจกๅž‹ๅŽ‚ๅ•†ๅŸบๆœฌๅบ”่ฏฅ้ƒฝๆ˜ฏ top 2 ๅˆฐ top 5 ็š„ๆถˆ่€—้‡ใ€‚ๅ…จ็ƒไธ‡ไธ‡็š„ใ€‚"

  • ๐Ÿ€ on the other hand, china / chinese ai startups are peak + more optimized than silicon valley by now. both the podcast and the blogs talk a lot on technical architecture to save cost: kv cache for context engineering, firecracker microvm (see below)...

  • ๐Ÿ€ no more prof / philosopher z.

  • ๐Ÿ€ indeed, for coming generations of startup, so much more glued / native to screen โ€“ to be drinking and screaming.

  • ๐Ÿ€ amazing, too: monica.im scales to $10m revenue and lots of user usage to mine for insights. so solid team + economic security during pivot from building ai browser. wish i realized earlier that sharding (in blockchain) is not the final game or (even a priority for ethereum after 8 years).

  • ๐Ÿ€ true, but 0% interest rate will come again. mostly, $100m in the first year have already happened / happening again.

  • ๐Ÿ€ "do not rely self to be: exceptional, _".

  • ๐Ÿ€ linus, too.

  • ๐Ÿ€ another big purpose, bring you back to: z + s > ai native > .. that is, ai before longevity. for the second level, ai agents > days-long tasks > deep+wide research..

  • ๐Ÿ€ per below, find "your +1".

  • ๐Ÿ€ per notion, state your current numbers of "relative strength"? deadlift, squat, bench.

  • ๐Ÿ€ per ไปฅ็ฌจไธบๆœฌ, more critical: per below, "disprove yourself" in a week vs month. per "< 5 days error half-life", what's your first truth / error of year by coming monday (jan 5)?

  • ๐Ÿ€ did not follow today? "6:30 am โ€“ 7:50 am: run (sunrise)".

  • ๐Ÿ€ i knew it would. but, to be reflect on the next day / 2026, not be stressed past 4pm. your "MIND" (but not _ "BUILD") worried me deep yesterday, too.

  • ๐Ÿ€ pick your fav among the playbook; aaron paying $2000 already.

  • ๐Ÿ€ yup, not building cursor clones.

  • ๐Ÿ€ aaron definitely being "the" premium prosumer.

  • ๐Ÿ€ also, resisting the common / my (past) urge to build an one-off ai travel agent (per every hackathon). bjorn still asking for it (while at new york).

  • ๐Ÿ€ but also general purpose ai scales, rather than once-a-year use! that is, the achilles heel of longevity / fitness. #simple-rules.

  • ๐Ÿ€ the objective too diverse among users to be aligned the model / product. but can be 10x better โ€“ given the right vector space with the full google maps + yelp review.

  • ๐Ÿ€ crazier still: china has no crawlable content, all locked behind apps / wechat / redbook.

  • ๐Ÿ€ per "$100t _" below.

  • ๐Ÿ€ not path, but insights: doing web1 verticals or web2 integrations is only 10% ai. incredible that they figure out sandbox (firecracker), context (kv cache), rewards (log replay).. so it's a general but 99% ai use.

  • ๐Ÿ€ ditto for your ai journal: find your +1 to see who will log like your.

  • ๐Ÿ€ yup, constantly "dare to disagree" with self: debating + asking mark / tracy / jeramey / everyone on telegram for feedback โ€“ every day.

  • ๐Ÿ€ also, the definite moment of pivot after months of hard-to-identify / hush-hush among team: manus collectively saw arc browser shutting down. the founder cannot get his friends to switch from chrome.

  • ๐Ÿ€ but, exactly: peak also mentions the long-tail use cases bring the sticky custoemrs. hours-long ai research sessions are not possible with all other ai chatbots or apps.

  • ๐Ÿ€ yup, find a beachhead and get to $100k revenue โ€“ per every ai product.

  • ๐Ÿ€ indeed, gemini will have even bigger wins / lead this year.

  • ๐Ÿ€ i bet it's more for practical than technical / product reasons. after raising a big round, conservative asian team, zuckerberg's charm.. $2.5b become irresistible.

๐Ÿ€ the secret of manus โ€“ scaling to 80m agents with 3-millisecond snapshots + 125-millisecond start time + 5-megabyte memory overhead: "Firecracker is great here, because agent sessions vary from milliseconds (single-turn, single-shot, agent interactions with small models), to hours (multi-turn interactions, with thousands of tool calls and LLM interactions). Context varies from zero to gigabytes."

๐Ÿ€ "An LLM agent runs tools in a loop to achieve a goal."

  • youtube transcript search (or monica.im, chattube, askvideo, youtube mobile) / 3-hour content / influencer finder / viral analyser, reddit product sentiment, twitter profile builder / trend monitor, instagram visual identity

  • cloud browser (airtop, hyperwrite), agent debate / adversarial validation.

  • ๐Ÿ€ "That content is unpolished; it's blurry photos and shaky videos of people's daily experiences. Think shoe shots and unflattering candids.. this is real because it's imperfect.

  • "less structure, more intelligence.": the most valuable agents are those that take a vague input and deliver a comprehensive, finalized output.

  • LLMs as Parts of Systems โ€“ combining ai with tools: code interpreters (python), vector database (postgres), cloud browser (chrome), logic solver (smt). "Amazon Bedrock's Automated Reasoning Checks uses LLMs to extract the underlying rules from a set of documents, and facts from an LLM or agent response, and then uses an SMT solver to verify that the facts are logically consistent with the rules. Danilo's blog post goes through a detailed example."

  • ๐Ÿ€ but also what manus stops serving / gemini won't launch this year โ€“ captcha / paywall sites: youtube, reddit, twitter.

  • ๐Ÿ€ exactly my big concern too. the most valuable ai / business tasks are non-generic knowledge / borderline illegal โ€“ just like in crypto.

  • ๐Ÿ€ then stick to our decades of advantage? coding agents will have another 10x breakthrough this year. per your ask, deploy a known set of libraries but to a specific config, resolve all dependencies and setup environment.. loop until successful deploy for a web demo on a proper .country domain.

  • ๐Ÿ€ really? what $100 fitness / health ai products that are actually impactful? per "profit on longevity" below.

  • ๐Ÿ€ or, voice agents that call 100 handymen / freelancers / restaurants /.. as lead generations, contract interviews..?

  • ๐Ÿ€ "screened" = gamed? with voice ai, you can define your top 3 questions / criteria โ€“ much like a proper interview. imagine your school-teacher scouting in 100x parallel.