?? + days long

ยท Zi Wang ยท 1 min read

Z / Runner ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ

๐Ÿƒ๐Ÿปโ€โ™‚๏ธ ?? (struggle to finish 2026 goals).

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ active session seems to work really well, assume pubmed lacks aggressive anti-bot control.

  • ๐Ÿƒ๐Ÿปโ€โ™‚๏ธ it is always a manual process, almost a quick & dirty RLHF, the evaluation is human judgment, which won't scale. I thought about creating a multi-rounded loops and maybe adding 1 or 2 layers of decision-trees (ai generated) to automate/standardize the "good vs. better vs. best" judgement.

Stephen / Basketball ๐Ÿ€

๐Ÿ€ just show the work of the day.

๐Ÿ€ #days-long cloud browser sessions โ€“ the only sustainable and scalable way of useful agent tasks over weeks: monitor markets, manage campaigns, reply customers, track competitors. as the god of web โˆž, ai must materialize at will (1000 instances in any region), impersonate as humans (authenticate + captcha + 2fa + proxy ip), portal with visuals (screenshots + streaming + replay), resurrect itself (serialize a session on event triggers).

  • $1000(0)/month services: context window checkpoints.., shared failure lessons, heavy chromium machines, mobile phone farm, static residential ip, sandboxed credentials security, credible fingerprint history.
  • karpathy on summoning gods: decouple the brain (language model) from the body (cloud browser) from the sense (multimodal understanding) from the soul (vector store).
  • existing: deepmind mariner, openai operator, anthropic computer use, perplexity research, manus browser, multion agent (agi inc), lindy ai.
  • infra: neko, cloudflare browser isolation, browserbox; browserbase, steel.dev, hyperbeam, adept act-1, skyvern; geelark, coronium, multilogin, iproxy, bright data, dolphin-anty; gologin, stytch; memgpt, langgpt. firecracker microvm.
  • benchmarks: android, web, computer. agent-q course: tree search, self critique, direct preference.
  • launched! 13.52.133.111 using port=8080 user=1 password=DTL2vIHaxQnJaENSVnEA. note: http, not https.
  • ๐Ÿ€ full browser ui, can take login auth.

๐Ÿ€ #ai-debate "running the same prompt against gemini, claude, chatgpt.. then manually reading / diffing through all responses? or copy-pasting into n^2 sessions so they debate with each other? or.. can a perplexity-like meta layer be able to syntheize / validate / finalize what you are doing manually?

  • ๐Ÿ€ decision tree -> goal hierarchy.