A consistent method for turning behavioural data into measurable gains.
How I built an evidence-led optimisation framework that transformed customer insight into continuous product improvements across multiple retail brands.
The problem was never a shortage of ideas.
Stakeholders proposed improvements every week. Product wanted one thing. Marketing wanted another. Merchandising had its own list. None of them were wrong — that was the problem. Every idea sounded reasonable in a meeting.
What didn’t exist was a shared way to answer five questions before any of them got built: Which problems actually matter? Which ideas deserve investment? Which opportunities are commercially worthwhile? Which changes should be prioritised? Which assumptions still need validating?
Without that, prioritisation became a negotiation, not a decision.
The challenge was never a lack of ideas. It was confidence — a repeatable way to know which ideas deserved investment.
Evidence artefact — the hypothesis backlog. A dense, unordered queue of competing proposals, before a shared framework existed to arbitrate between them.
One loop, run continuously.
Research produced evidence. Evidence produced a hypothesis. Hypotheses became prototypes. Prototypes became experiments. Experiments produced new evidence — which fed the next hypothesis.
The individual experiment was never the point. The loop was.
Not every idea earned the same amount of evidence before it shipped. Some were resolved quickly — fixed and released, no formal test required. Others carried real uncertainty, and those went through the full loop: research, prototype, test, result. The judgement was knowing which was which.
Not every idea needed the same weight of evidence. Quick fixes and tested hypotheses were treated as two different categories from the start.
What the loop actually looked like, once.
The next six beats are one real project — the Harveys Swatch Selector — and they map exactly onto the loop above.
The plan, stated before anything was built.
Because we saw users in moderated research struggle to change colour options on Harveys’ mobile product page, we expected repositioning the selector would improve ease of navigation. We measured this using conversion rate and swatch interaction — named before the test ran, not chosen afterward to flatter the result.
The problem, found in the lab.
Colour swatches sat below the product image, in a separate scrollable panel. Changing colour didn’t visibly update the image without scrolling — users scrolled up to check it had worked, then back down to keep choosing. Small friction, repeated constantly, on the page carrying the most commercial weight on the site.
Solved cheaply, before it got expensive.
A sketch, tested against paper with colleagues who weren’t designers. Then a wireframe, carrying the same structure into a rapid screen prototype. Each stage cost almost nothing and ruled out the ideas that wouldn’t have survived contact with a real customer.
Only once the direction was confident.
Full visual design, validated once more through guerrilla research on a working prototype — before a single line of production code. Only then, a formal A/B test: a 50/50 traffic split between the redesigned control and a bolder variation, run against real customers.
Reported at its real, honest value.
+13.2% increase in conversion rate. +69.2% increase in interaction. An 84.2% probability of beating control, calculated using Bayesian statistics and reported without rounding up. Predicted return: £388,467 over six months.
The test didn’t find the problem. A person, watching another person’s hands hesitate, found the problem. The test told us what it was worth.
The learning didn’t come from the test. It came from watching someone scroll.
Not once. Repeatedly.
The Swatch Selector wasn’t a special case. It was the same loop, run against a different problem, across six years and two brands.
Finance Calculator
- Problem
- Users struggled to understand and use the finance calculator on mobile product pages.
- Evidence
- Moderated research flagged confusion at a specific interaction point on the mobile PDP.
- Approach
- Reformatted the calculator; measured against three metrics, not one — Add-to-Basket rate, conversion rate, and finance application rate.
- Outcome
- +30.4% conversion rate, +31% increase in transactions and +14.1% increase in number of finance applications
- Learning
- Measuring three outcomes instead of one prevented a narrow win from masking a broader problem.
Checkout Optimisation
- Problem
- The checkout funnel lost 18% of all traffic at Payment Details — worse on mobile (22%) than desktop (13%).
- Evidence
- The three payment methods behaved nothing alike — 67% of users chose Finance, but PayPal, chosen by only 6%, had the highest drop-off at 42%.
- Approach
- The funnel numbers were only one input. Heuristic evaluation, competitor benchmarking, session recordings and service blueprinting were triangulated before validating improvements through a 2017–18 programme of labs and A/B tests.
- Outcome
- +27% checkout conversion, measured over a four-week A/B test ahead of full rollout.
- Learning
- Analytics highlighted the symptoms. Research uncovered the causes. The checkout wasn’t one problem — it was three different journeys wearing one name.
Search & Navigation
- Problem
- Customers relied on search constantly, but it sat visually hidden in the header — friction for the people with the clearest purchase intent.
- Evidence
- Exposing it lifted sessions with search and progression to PDP, but also reduced use of the store finder and softened conversion among searchers — a sign the variation reached lower-intent users too, not just more of the same ones.
- Approach
- No new feature, no copy change — the search bar itself was untouched. Made visible, then measured across the full funnel, not just search usage in isolation.
- Outcome
- +54% uplift in sessions with search · +51% uplift in progression to PDP
- Learning
- The biggest single number in the whole programme came from making an existing tool visible, not building a new one — though it changed who used it, not just how many.
SleepPRO Kiosk
- Problem
- The supplier-built kiosk had too many steps and little signposting — in-store testing showed customers couldn’t navigate it unassisted.
- Evidence
- Moderated in-store sessions with real customers surfaced exactly where people got stuck, screen by screen.
- Approach
- Sketched a simpler flow, built a rapid screen prototype, tested it with colleagues who’d never seen it before, reiterated.
- Outcome
- +25% store conversion · +60% customer usage
- Learning
- A supplier’s default interface isn’t a fixed constraint. It’s just another UI, and it responds to the same process as anything built in-house.
None of it ran on instinct.
Every hypothesis in this programme started as evidence, not opinion. Moderated usability labs, guerrilla research, customer interviews, in-store observation, surveys. Behavioural analytics, funnel analysis, exit analysis, device segmentation, A/B and multivariate testing.
Neither alone was enough. Analytics could tell you where people left. Only a lab session could tell you why.
Real respondent behaviour, not a designer’s summary of it — the payment-method split that shaped the checkout redesign.
Act V — What Compounds
Every experiment became something the next one could use.
A test that failed wasn’t deleted. It was archived, separately, deliberately — its own category, not folded into the wins. Over six years, that discipline built something larger than any individual page redesign: a living record of what had already been tried, what had already been learned, and what was still worth asking.
The programme didn’t just improve the product. It got faster at improving the product, year over year, because less of it had to be rediscovered.
The evidence didn’t just prove individual decisions. It made every decision after it faster and better informed.
120 cards · 13 lists · 177 attachments. Including a dedicated, standing archive of losing and inconclusive tests — kept, not discarded.
The work that never touched a screen.
The KPI taxonomy behind this programme didn’t stop at the website. Alongside online conversion, every initiative was tagged against telesales support and store footfall — evidence that the thinking here was never scoped to a browser window. A change to a mobile PDP was expected to show up in a call centre’s numbers and a store’s footfall, not just a funnel report.
The research hub itself is a leadership artefact, not just a research archive. Choosing to build shared, structured, searchable organisational memory — rather than letting six years of research live in inboxes and forgotten decks — is itself a decision only someone thinking beyond their own deliverables makes.
Reflection
What actually compounded.
The biggest lesson wasn’t that one experiment improved conversion. It was that evidence-led decision-making compounds — each result making the next hypothesis sharper, each failure archived rather than repeated.
Research, prototyping, and experimentation were never separate disciplines here. They were the same instrument, used at different speeds — a sketch reduces uncertainty for the cost of an afternoon; a guerrilla test, for the cost of a day; an A/B test, for the cost of a build.
By building a repeatable framework, optimisation stopped being a sequence of projects and became an organisational capability. That’s still true of how I work now.
See the same discipline applied to a longer, higher-stakes discovery programme.