1m errors + savoir faire
Β· Zi Wang Β· 1 min read
Z / Runner ππ»ββοΈ
ππ»ββοΈ Solving a Million-Step LLM Task with Zero Errors (arxiv), shared by shamir; some of the concepts are worthwhile to explore w/ general agent. Arch': max agentic decomposition + first to ahead by k-votingβ¦ the ability to execute 1M+ steps; extreme sub-task assignments, multi-sampling + multi-agent votingβ¦ this is probably why manus exceeds its peers. Instead of cot w/ small number of substeps, have a large sample of steps w/ quality control (red flagging).
-
ππ»ββοΈ yeah, for extreme long reasoning (think multi-day sessions), the ability to reduce error rate is so valuable. As the paper stated, errors do compound and get more expensive further down the cot. -
ππ»ββοΈ towers of hanoi (only focus the execution, apply templated strategies (don't rely on runtime). -
ππ»ββοΈ voting rule: sample candidate actions repeatedly until one candidate is ahead by k vs. other agents (runtime testing). -
ππ»ββοΈ voting has it is place, many ai compute/reasoning tasks don't need 99.99999999999β¦ sampling maybe a cheaper and faster. -
ππ»ββοΈ don't know the where research ends and production grade deployment starts, it feels very expensive to "over" decompose tasksβ¦ compute is cheap but not free. -
ππ»ββοΈ clever way to split LLM reasoning vs. tool execution; for industrial grade domain applications, this is may have similar benefits as lightweight SMT w/o having to formalize everything. -
ππ»ββοΈ the math and logic are right, but remain skeptical if the big ai leap in 2026 will be formal verification. (we had similar pattern in crypto for nearly 5+ years, it is still just making its way to the mainstage). -
ππ»ββοΈ common dyi, functionality != usability. Just b/c anyone can put these devices and ai slab things together, it doesn't mean it will work well. But i guess the larger point about learning vs. doing, making progress vs. being perfect.. -
ππ»ββοΈ speaking of hermes, 1 of the few times, yo & I listened to a podcast together for 3+ hrs (link).
Stephen / Basketball π
π but smt together!
-
π "[1-million-step] decomposition of a task into subtasks, each of which can be tackled by focused microagents. [the modularity] allows error correction.. through an efficient multi-agent voting scheme" -
π "per the paper, top use cases for general ai agents": digital content monetization, software development. -
π i bet voting is the wrong approach. formal verification is the only way for absolute and 0%. -
π "team and vc behind cognizant.com": nah, just a research lab at ut austin. -
π "compare the paper to formal verification and tools based on type theories": only (1-1e-9)**1e9 = 37% correct even with 9 nines.
π hermes of agents: hand craft, member lore, service bundle, "savoir-faire" (know-how) as deep training + end-to-end integrations.. vs monkey-see-monkey-do by a demi god.
π exactly my link and spin-off thoughts here. (vs porsche).