110m time + $9 analyzer

Β· Zi Wang Β· 1 min read

Z / Runner πŸƒπŸ»β€β™‚οΈ

πŸƒπŸ»β€β™‚οΈ response to "#video-analyzer" (willingness to pay $9/ ai agent/action): i already run high info throughout per YT. baseline my consumption patterns & usage (per my YT history): avg ~110 mins/day of YT, mostly 1.5x - 2x, audio-first.

  • πŸƒπŸ»β€β™‚οΈ mostly ambient, passive info-tainment (most podcasts are ~= distributed via yt)... i rarely watch anything w/ full attention on yt, that's reserved for reading & working. frame seeking is less than 10%/ rewatching happens mostly on shorts. As you noted, most yt is research driven for me.

πŸƒπŸ»β€β™‚οΈ Given i have decent recall on what i listen, my "digestion" of info is average or slightly above the norm. What i will pay for is a superhuman edge. Meaning, make me shine, where I can "watch 1,000 videos" per session not via summary but get decisive insight to take action

  • πŸƒπŸ»β€β™‚οΈ short answer yes, long answer, given the compression/decompression of yt videos are managed by general agent, then it is safe to assume your workflow is flexible for all content types.

πŸƒπŸ»β€β™‚οΈ Time & memory (understanding/insights) still my hard-cap. Speed alone doesn't bend that curve anymore. Meaning, does our agent gives me structural compression, or is it marginal convenience (manus delivers that).

  • Give a specific example, Blinkist (i sub for about 6+ yrs); no longer useful to me. Taking a 30+ audiobook and collaps into 30 mins was worth $5-10/month. Faster consumption (2000s - 2022)β†’ order of magnitude reduction in cognitive cost (not just time).

  • πŸƒπŸ»β€β™‚οΈ to be more blunt & wielding a bigger hammer, subs to "the information", ben thompson, peter attia, huberman lab… should all down to zero if there is a generalizable agent that can help 10,000 longevity startups…

  • Single-video analysis doesn't clear that bar.

  • As per manus session, i run a manual loop before/after watching/listening for useful videos: grab transcript, summarize, jot notes, form mental model, collect/analyze comments (there are a ton of valuable meta data)... it's too linear & local (problem statement).

  • πŸƒπŸ»β€β™‚οΈ YES

  • What I want: i know real signal isn't inside one video, it is across multiple videos on the subject. Treat collection of YT playlist with reinforcing & contracting ideas that evolve over w/m/y (hence recurrence). This is where I struggle the most, limited time, cognition and no agent workflow captures it.

πŸƒπŸ»β€β™‚οΈ zoom out: frame by frame analysis; you & I are aligned, per video/frame, downsample, keyframes, fast to deep thinking… this may be the necessary plumbing to build the above system.

πŸƒπŸ»β€β™‚οΈ No sure if you have thought about corpus-first vs. video-first? For example, if we treat "subscription" as the base unit of "corpus" and the ai agent can auto my workflow, then I would pay $10/momnth. Ingest everything new across all my subss, compress across time & creators, surface meta/hidden data, highlight contradictions, themes, deltas and give me an output that's denser and decision/action friendly format.

  • πŸƒπŸ»β€β™‚οΈ frame-level perception might be the necessary infra… but paying $10/ usage needs to meet product/value boundary.

Stephen / Basketball πŸ€

  • πŸ€ 110 mins after speedup?! 95% passive listening (of podcasts)? how much watching with full attention? how much actively searching, frame seeking, rewatching.. as research / study?

  • πŸ€ a superhuman edge on.. longevity? do you still want / need an edge on running / product / design / kids..? is there an edge for entertainment videos like nba / books / ..?

  • πŸ€ not bad for 10 * 12 * 6 = $720.

  • πŸ€ 1000 videos + life-long context.

  • πŸ€ also organization + knowledge across sub (agent) sessions: project folders, prompt search, tag labels, copy/paste context.. are just terrible for cognitive overload.

  • πŸ€ exactly! any human can watch only 1% of 1% of 1% = 500 hours per year of the 500m-hour total youtube videos.

  • πŸ€ 100% aligned. note: videos not only on youtube, but on phone (lazy) + google/apple cloud (all consumers) + amazon s3 (professional). e.g. my basketball + emma dance practice videos.

  • πŸ€ this opportunity window probably only lasts 12 months: photo analysis of key video frames are perfect, but raw video understanding is completely failing.

  • πŸ€ very possible if <1000 video crawling per day per user. but must be priced at $99 (prosumer) if not at least $40 (like manus) per month.

πŸ€ #video-analyzer z, how many videos and what tasks will you pay $9 each for ai (agents) to complete today?

  • break an hour-long video into 60*60 = 3600 frames (1 second each).
  • run gemini 3 flash on each low-res frame (70 tokens) for $0.13 total to pick 1% key frames.
  • run gemini deep think on 36 high-res frames (258 input + 5000 thinking tokens) for $1.50 total.
  • or, run gemini deep think on a merged video of the key frames for $0.40 total.
  • πŸ€ video ai is / will remain expensive. on-demand pricing vs subscription fee will be key.