On Sunday I rebuilt WarmPath, my job-search app, in 2 hours. Building it took 3 weeks in June. The kids went to bed on time.
A rewrite like that used to be the kind of project I’d plan for months as an engineering manager (the “small rewrite” that was actually a six-month migration). Now it takes an afternoon and some careful thought up front. Friends who don’t write code keep asking me what that means for their work, so here’s the rule of thumb I use: AI gets very good, very fast, at any work where a computer can check the answer.
Software is the clearest case. Most code comes with tests, small programs that say pass or fail, so the computer can grade its own homework. The labs train models on exactly that. DeepSeek’s R1 paper rewards the model when a math answer matches or when code passes predefined test cases, and the approach now has a name: reinforcement learning with verifiable rewards. The same check works on the job. In my own Bakeoff runs, the frontier models wrote tests without being asked (GPT-6 Astra on 12 of 14 runs).
Even without tests, AI can catch more errors than you expect by running the software on live data. WarmPath’s first full run turned up a job listing where “Pay Type: Salary” was being shown as the pay. An automated code review caught a place where WarmPath could crash that could have sent the daily email twice. As I wrote in August, experience gives me a good hypothesis and AI lets me test it. More and more, AI does the testing.
My favorite example comes from retro video games. A Super Nintendo cartridge like Star Fox holds only machine code, the bare minimum the console needs, with none of the names or notes a person would use to understand it. Turning that into a PC program is called recompilation, and for decades almost nobody did it because the reading was too tedious. The workaround was emulation - pretending to be the old console. That changed this year, with coding agents checking recompiled games against their emulated versions. Star Fox got a SNES recompilation on GitHub in July, built with AI, and Matthew Stanley has built recompilers for the NES, SNES, Genesis, PlayStation, and more since February, work he says “historically taken teams of developers years”. The tedious part of rewriting the code in a modern way is no longer the obstacle to recompilation.
So what’s left for people? Deciding what “correct” means: clear goals, clear requirements, and the edge cases a few trial runs won’t show you. The reason I had to rebuild WarmPath was that my initial goals were muddled - I wanted to build an agent and I also wanted to find relevant jobs. Together these goals grew the code to about 11,000 lines, which made adding a new user a big job. Adding a user shouldn’t be that hard, so I cut the scope back to ranking relevant jobs with AI. The rebuild is about 2,000 lines of code.
Which field is next? My bet is anything where the check can be automated:
- Math. A computer can check a proof line by line. Google DeepMind scored 35 of 42 at the 2025 International Mathematical Olympiad, a gold medal score on problems written for high school students which blew my mind at the time. Fourteen months later, OpenAI says an internal model solved the Navier-Stokes Millennium Prize problem, a question about how fluids move that had been open for about 90 years, and a computer checked the proof in 17 hours. It’s OpenAI’s own claim and very new, but high school problems to a 90-year-old open problem looks like acceleration.
- Lab science with robots. GPT-5 ran 36,000 experiments in Ginkgo’s robotic lab over six months and cut the cost of making a protein by 40% below the best published result.
- Weather. Every forecast gets checked against what actually happened. Google’s GenCast beat the best operational forecast on 97.2% of 1,320 targets back in 2024.
Are the models just memorizing the test? Partly. OpenAI stopped reporting SWE-bench Verified, a coding benchmark it helped build, after finding that every frontier model it tested could reproduce some of the answers. But scores are also climbing on tests the labs haven’t published. ARC-AGI-2 launched in March 2025 with plain language models at 0%, and the best result on its held-back set is now 95%. I think the models are getting generally smarter, not just better at studying. I don’t know how far that goes.
What in your work has a clear right answer a computer could check? That’s where I’d start.