110m time + $9 analyzer
Β· Zi Wang Β· 1 min read
Z / Runner ππ»ββοΈ
ππ»ββοΈ response to "#video-analyzer" (willingness to pay $9/ ai agent/action): i already run high info throughout per YT. baseline my consumption patterns & usage (per my YT history): avg ~110 mins/day of YT, mostly 1.5x - 2x, audio-first.
ππ»ββοΈ mostly ambient, passive info-tainment (most podcasts are ~= distributed via yt)... i rarely watch anything w/ full attention on yt, that's reserved for reading & working. frame seeking is less than 10%/ rewatching happens mostly on shorts. As you noted, most yt is research driven for me.
ππ»ββοΈ Given i have decent recall on what i listen, my "digestion" of info is average or slightly above the norm. What i will pay for is a superhuman edge. Meaning, make me shine, where I can "watch 1,000 videos" per session not via summary but get decisive insight to take action
ππ»ββοΈ short answer yes, long answer, given the compression/decompression of yt videos are managed by general agent, then it is safe to assume your workflow is flexible for all content types.
ππ»ββοΈ Time & memory (understanding/insights) still my hard-cap. Speed alone doesn't bend that curve anymore. Meaning, does our agent gives me structural compression, or is it marginal convenience (manus delivers that).
-
Give a specific example, Blinkist (i sub for about 6+ yrs); no longer useful to me. Taking a 30+ audiobook and collaps into 30 mins was worth $5-10/month. Faster consumption (2000s - 2022)β order of magnitude reduction in cognitive cost (not just time). -
ππ»ββοΈ to be more blunt & wielding a bigger hammer, subs to "the information", ben thompson, peter attia, huberman labβ¦ should all down to zero if there is a generalizable agent that can help 10,000 longevity startupsβ¦ -
Single-video analysis doesn't clear that bar. -
As per manus session, i run a manual loop before/after watching/listening for useful videos: grab transcript, summarize, jot notes, form mental model, collect/analyze comments (there are a ton of valuable meta data)... it's too linear & local (problem statement). -
ππ»ββοΈ YES -
What I want: i know real signal isn't inside one video, it is across multiple videos on the subject. Treat collection of YT playlist with reinforcing & contracting ideas that evolve over w/m/y (hence recurrence). This is where I struggle the most, limited time, cognition and no agent workflow captures it.
ππ»ββοΈ zoom out: frame by frame analysis; you & I are aligned, per video/frame, downsample, keyframes, fast to deep thinkingβ¦ this may be the necessary plumbing to build the above system.
ππ»ββοΈ No sure if you have thought about corpus-first vs. video-first? For example, if we treat "subscription" as the base unit of "corpus" and the ai agent can auto my workflow, then I would pay $10/momnth. Ingest everything new across all my subss, compress across time & creators, surface meta/hidden data, highlight contradictions, themes, deltas and give me an output that's denser and decision/action friendly format.
ππ»ββοΈ frame-level perception might be the necessary infraβ¦ but paying $10/ usage needs to meet product/value boundary.
Stephen / Basketball π
-
π 110 mins after speedup?! 95% passive listening (of podcasts)? how much watching with full attention? how much actively searching, frame seeking, rewatching.. as research / study? -
π a superhuman edge on.. longevity? do you still want / need an edge on running / product / design / kids..? is there an edge for entertainment videos like nba / books / ..? -
π not bad for 10 * 12 * 6 = $720. -
π 1000 videos + life-long context. -
π also organization + knowledge across sub (agent) sessions: project folders, prompt search, tag labels, copy/paste context.. are just terrible for cognitive overload. -
π exactly! any human can watch only 1% of 1% of 1% = 500 hours per year of the 500m-hour total youtube videos. -
π 100% aligned. note: videos not only on youtube, but on phone (lazy) + google/apple cloud (all consumers) + amazon s3 (professional). e.g. my basketball + emma dance practice videos. -
π this opportunity window probably only lasts 12 months: photo analysis of key video frames are perfect, but raw video understanding is completely failing. -
π very possible if <1000 video crawling per day per user. but must be priced at $99 (prosumer) if not at least $40 (like manus) per month.
π #video-analyzer z, how many videos and what tasks will you pay $9 each for ai (agents) to complete today?
break an hour-long video into 60*60 = 3600 frames (1 second each).run gemini 3 flash on each low-res frame (70 tokens) for $0.13 total to pick 1% key frames.run gemini deep think on 36 high-res frames (258 input + 5000 thinking tokens) for $1.50 total.or, run gemini deep think on a merged video of the key frames for $0.40 total.π video ai is / will remain expensive. on-demand pricing vs subscription fee will be key.