Method
The story, then the build
Every project starts with a written case for why it should exist. AI models write most of the code from there, and I am responsible for the decisions, the testing and the call to ship.
The method in one view
I trained as a journalist and spent years in content and research before building products. I still work like a reporter. I check the first account against a second source, reconcile evidence that disagrees, and follow a feature request back to the problem underneath it.
The written case has three parts, shown below. Its last question, how we will know, becomes the test plan for the build.
Start with
Customer evidence
The reason
Write the story
-
Situation
What is true now, for whom, and what it costs while it stays true.
-
Complication
What has kept it that way, and why earlier fixes did not stick.
-
Resolution
What will be different for the customer, and how we will know.
The behavior
Build against it
-
01
Draft
The model produces the first pass. I read every line before it lands.
-
02
Evaluate
I decide what gets tested, starting with the failures that would stay silent.
-
03
Harden
I test hostile inputs, missing data, accessibility, reversibility and runtime fallbacks.
-
04
Attack
A separate adversarial pass hunts for poisoned state, lost edits and interactions that fail under real use.
Ship gate
The reason and the behavior both have to survive review.
Where judgment stays human
Before Segment Atlas, I had never decoded a binary ride file, parsed a geospatial raster or built a 3D scene. Models got me through that quickly. They did not decide that recorded results stay primary, that estimates carry confidence labels, or that an oversized route is refused before any data is fetched. I made those calls.
Stage
I own
The model contributes
Frame
I own
Which problem is worth solving, and the evidence behind it.
The model contributes
Research breadth, alternate framings and counterexamples.
Design
I own
Architecture, tradeoffs, data contracts and the decisions users will feel.
The model contributes
Implementation options, first drafts and refactors.
Evaluate
I own
Failure inventory and test criteria.
The model contributes
Test scaffolding, fuzz cases and review breadth.
Ship
I own
Trust UX, reversibility and the final release decision.
The model contributes
Documentation and cleanup.
Where it went wrong
The mistakes that matter look plausible. These four come from Segment Atlas and the internal AI suite.
Mistake
What happened
What changed
Math that looks right
Segment Atlas
What happened
Averaging wind directions of 350° and 10° gives 180°, due south. A model drafts that confidently, and a quick review passes it.
What changed
Wind is interpolated as vectors, and a test pins that exact case to north.
Read the test ↗A theory that measurement disproved
Internal AI suite
What happened
Heavy reports failed intermittently. My first theory, too many calls at once, pointed to a fix that would have made it worse.
What changed
I measured the retry behavior first. The real limit was the cost of each request, so heavy requests now split by brand. A 63-part report runs in 27 seconds.
Tests that passed when they should not have
Internal AI suite
What happened
Twice, a regression check passed because its test data had been built from files that were already trimmed.
What changed
Test inputs are now padded to full size, and each check is compared against live output before its results count.
A number too wrong to ship
Internal AI suite
What happened
One brand’s monthly figure came back about 130 times higher than the client’s delivered file.
What changed
I left the brand out with a visible note. It was confirmed and added back five days later through configuration, with no code change.
Check the work
The Segment Atlas code is public, tests included. The case studies show the same method on enterprise work, and two more are available on request.