i gave two AI setups the same four hours


i gave two AI setups the same four hours

i built BuilderBench around work i actually use AI for: writing, editing video, designing pages, making motion graphics, using apps, researching decisions and fixing workflows.

the first pilot compared Opus 5.5 with Claude and GPT-6 Sol with Codex. same frozen task pack. four hours of offline work, plus a live founder conversation. Extra High reasoning and Standard speed.

i scored the anonymous work, then locked my ratings before revealing the names.

Opus 5.5: 68.26/100.
GPT-6 Sol: 45.55/100.

Opus led on seven of the eight task families. Sol led on browser and computer use, and scored 90 on intake automation.

the cost figures need care: US$49.67 for Opus in provider-reported API-equivalent estimates; US$12.75 for Sol from partial recorded usage.

these are estimates, not bills. i used existing subscriptions. actual extra spend is unknown. the methods and coverage differ, so these numbers do not give us a reliable cost ratio.

this is one owner-scored v1 pilot, testing each model together with its harness and tools. it is not a universal ranking. complete normal-speed watched-and-listened review of both full video outputs was not established.

v2 changes the tasks and inputs. its scores will be labelled separately.

see every task score, the setup and the cost limits

Lenny

Lenny’s AI Builders

i’m Lennox. i sift through AI Twitter and share one thing i tried with ChatGPT. for people making products, content and useful systems. five emails a week, with a weekly option.

Read more from Lenny’s AI Builders

i've got 2 for the price of 1 for you: LAB-0005 and LAB-0006. the headlines for 0005: a big maths claim from OpenAI. a model price fight. Grok 4.7 on a real job. what agents might do to app checkout. and a local image model i want to test. how to use GPT-6 Astra to create brand design assets (spoiler alert: the local image model is wild. it does *anything*. without guardrails - scary stuff if put in the wrong hands) check out episode 0005 here: headlines for 0006: Opus 5.5 drops 1.5 hours...

7 ai updates that changed how i build this week: we've entered the token scarcity era use a handoff.md for long-running agent tasks check out CUA (Computer Use Agent) if you use Claude agent portability is the new meta; don't get locked in to one app ... ... ... [this one's a fkn doozy!] i'm not gonna make it that easy for you ;) want the last 3 updates? you know what to do! watch LAB 0004 on YouTube speak soon x Lenny

Jev is a new type of AI based on System 1 thinking. wtf does that mean? Daniel Kahneman coined this idea of System 1 and System 2 thinking. basically, S1 makes decisions super fast; S2 thinks slowly about a problem. think: difference between reacting after feeling your hand on a stove vs using mental maths to calculate 69 x 43. since the end of 2024, AI models have been getting better (much, much better) at S2 "long and slow" thinking. that left a gap for S1. until Jev. that's what i dive...