How to run a device intelligence bake-off

Person waving a chequered racing flag against a blue sky.

We sell device intelligence, so treat this accordingly. It is also the most useful thing we can publish, because most evaluations in this category are run badly and the people running them know it.

What follows is how we would run one, including the parts that could go against us.

Run both. Replace nothing

The instinct is to swap one system for another and compare the before and after. Do not.

Traffic changes. Attack patterns change. A quiet fortnight followed by a busy one will produce a result that has nothing to do with either vendor. Sequential comparisons in an adversarial environment are close to meaningless.

Run both systems on the same traffic, at the same time, on one flow, with no enforcement from the new one. Nothing switched off, nothing rerouted, no customer experiencing anything different. This is usually called shadow mode and it is the lowest-risk version of the exercise.

The output you want is not a score from either vendor. It is a set of cases where the two systems disagreed.

Pick the flow for label speed, not importance

This is the single most useful paragraph here and it is the one most often skipped.

An evaluation is only as good as your ability to find out who was right. That depends entirely on how quickly the truth arrives in the flow you chose.

Bot registration and promotional abuse resolve in days. You know within a week whether an account was farmed. Duplicate signups, incentive claims and referral abuse all produce answers fast.

Chargebacks take 30 to 120 days. Sometimes longer. If you evaluate on payment fraud with a 30-day window, the window closes before the labels arrive, and you finish the exercise with a large set of disagreements and no way to settle any of them.

So the instinct to test on the most valuable flow is usually wrong. Test on the flow that answers fastest, then extend to the valuable one once you know how the systems behave.

If payment fraud is the only flow that matters commercially, plan for 120 days and say so at the start, rather than discovering it at day 29.

Agree what “right” means before you start

Both systems will produce outputs. Neither output is the truth.

The reference has to be yours. Confirmed fraud cases. Chargebacks that were upheld. Accounts you banned and did not have to unban. Promotional claims you reversed. Things that resolved, in your data, where somebody eventually established what actually happened.

Write down which of these counts as ground truth before the run begins, and who adjudicates the ambiguous ones.

This matters more than it sounds. Without an agreed reference, a detection rate and a false-positive rate are indistinguishable. “System A flagged more” and “System A was wrong more often” are the same observation until something external tells you which. Every argument you will have at the end of the evaluation traces back to this paragraph.

Write the success criteria down first

Decide, in writing, before anything runs:

What result would make you switch. What result would make you keep what you have. What result would be inconclusive, and what you would do about it.

Goalposts move once the first interesting result lands, and they usually move in the direction of whoever is most invested. A document written before the run is the only defence against that, and it takes twenty minutes.

Measure the disagreements

Where two systems agree, they tell you nothing new. Agreement is the large, boring majority of the traffic and it should be set aside immediately.

The disagreement set is the experiment.

For every case where the systems differed, you want to know what happened afterwards. Did the account get banned. Did the payment charge back. Did the promotional claim get reversed. Did nothing happen at all, which is itself an answer.

Count four outcomes. Cases where the incumbent was right, cases where the challenger was right, cases where both were wrong, and cases that never resolved. That fourth number is often the largest and it is worth reporting rather than hiding.

Ask both vendors for the full metric set

Including us, and including whatever you already run.

Coverage. On what proportion of sessions does the system return a usable answer at all.

Abstention. How often does it decline to answer, and on what kind of traffic. This is the number almost nobody publishes, and it is the one that makes every other number readable. A high detection rate on the 60 per cent of traffic a system is confident about is a different product from a high detection rate on all of it.

False positives and false negatives. Both, or neither. A detection rate on its own is improvable by flagging more. A false-positive rate on its own is improvable by flagging less. Either number alone can be moved by degrading the product.

Precision and recall, with the denominators stated.

If a vendor can give you some of these and not others, ask which ones are missing and why. The pattern of what is available is informative.

What thirty days cannot tell you

Say this out loud at the start so nobody over-reads the result.

Rare events. If the thing you care about happens twice a quarter, a month of traffic contains no useful information about it. The evaluation will be silent on your most expensive risk.

Anything with a long label lag. Covered above, and it is the most common way these exercises fail.

Behaviour under an attack neither system has seen. Both systems are being measured on the traffic that happened to arrive. A coordinated attack in month four is the actual test and no evaluation can run it in advance.

Integration quality at scale. A pilot on one flow tells you almost nothing about what happens when the same system is in eight flows with real latency budgets.

An evaluation narrows uncertainty. It does not remove it, and a vendor implying otherwise is telling you something about themselves.

A workable shape

For a first exercise, this is roughly what we would suggest.

One flow, chosen for label speed. Thirty days of parallel running with no enforcement. An agreed reference written down in week zero, and success criteria written down with it. Weekly check-ins that look at volume and coverage only, deliberately not at who is winning. Then a single analysis at the end, on the disagreement set, with the four outcome counts and the full metric set from both systems.

Then a decision that is allowed to be “neither, for now”. That is a legitimate result and it is more common than vendors admit.

What we do

One section, and it is the only part of this post that is about us.

We would rather run alongside what you have than replace it. One flow, no enforcement, and your outcomes decide who was right. We will give you the disagreement set, the abstention rate we ran at, and the cases where we returned nothing at all.

If what you already run wins, that is a useful thing to have established, and you will have established it with your own data rather than anyone’s marketing.