Test the journey, not the click.
Most CRO is a redesign in disguise: someone's taste, a best-practice checklist, and a test that calls a winner on traffic that was half bots. We do the other kind - tests your data asked for, judged on the whole journey they produce.
This is not a menu item - it's where your Foundation report leads. When the monthly analysis finds a specific funnel leak and prices it, we scope the testing programme as a fixed-price engagement and run it on your data. It starts with DarkField Foundation.
Behavioural A/B - the journey decides, not the click
How CRO is usually done.
Testing isn't the weak point - what gets tested, and how winners get called, is.
Opinion-driven backlogs
A brainstorm produces a list of hunches - button colours, hero swaps - and the roadmap is whoever argued best. Your own data never got a vote on what to test first.
Dirty-traffic verdicts
Bots, staff sessions and out-of-market visitors all vote in most A/B tests. Enough noise and a coin flip looks significant - teams ship “winners” that never show up in revenue.
Day-one scoring
The variant that squeezes one more same-session conversion wins, even when it builds visitors who never come back. Nobody measures what the test did to the customer, only to the click.
Testing opportunities come from the data
The monthly analysis surfaces where visitors hesitate, drop off or loop back - the funnel step that leaks, the pages viewers abandon, the path returners take that first-timers can't find. Those become the test queue, ranked by expected impact. No guessing which button to recolour; the data nominates the experiments.

Behavioural A/B testing
Two variants can tie on day-one conversion while one quietly builds visitors who explore deeper, return more and spend more over time. Because every session is scored on the 49-signal engagement scale, we measure the whole journey each variant produces - depth, returns, score trajectory, predicted LTV - and ship the variant that compounds, not the one that flattered a single session.
Verdicts you can trust
Tests run on cleaned traffic - bots, staff sessions and out-of-market visitors stripped out before anyone votes - so a winner is a real winner and a tie is called a tie. Behavioural metrics also move earlier than purchases, which means verdicts arrive on less traffic than purchase-only testing needs. Every verdict is written up in the monthly report, including the losses.
Tooling that fits your stack
We use a proven testing tool where it fits, or build lightweight custom tests where it doesn't. Either way the results land in your warehouse, measured in the same monthly report as everything else.
What you actually get
- A test queue built from your own data, ranked by expected impact - refreshed every monthly report
- Tests run on cleaned traffic: bots, staff and out-of-market sessions can't vote
- Whole-journey verdicts - depth, return rate, engagement score, predicted LTV - not just day-one conversion
- Every test written up honestly in the monthly report: wins, losses and ties alike
- Tooling that fits your stack - a proven testing tool where it fits, lightweight custom where it doesn't
Scoped from your Foundation report - either your team ships the winning variants from a precise spec, or we execute the testing programme ourselves.
Same day-1 conversion, different business
In a behavioural test, variant B tied on conversion but its visitors explored deeper and scored far higher - so they converted again and spent more. Traditional testing would have called it a tie.
Engagement scoring
The 49-signal score is the measuring stick for every test - a leading indicator of revenue, available days before the sale.
You're probably wondering.
01Do we have enough traffic to A/B test?
If you're in our fit range (roughly 50k+ visits a month), usually yes - and behavioural metrics help, because engagement moves on far more sessions than purchases do, so verdicts need less traffic than purchase-only tests. Where volume is genuinely thin for a question, we say so and test a bigger swing instead of pretending a small one reached significance.
02Isn't CRO just “make the button bigger”?
The checklist version is. The version that moves revenue works on the journey: the sample report's month (an ecommerce store), for instance, turns on the product-view→cart step doubling - that's page-level work on real friction, found in the funnel, not in a heuristics blog post.
03Who implements the winning variant?
Your choice. With Foundation alone: your developers, from a precise spec - what to change, where, and what it's expected to earn. As a custom engagement, we execute the testing programme ourselves, end to end. Either way the next report measures what shipping it actually did.
Find your next test in the data.
It starts with Foundation - the report builds the test queue from your own funnel before anything gets redesigned.