Index / Detection and tuning

Behavioural Analytics: What It Can and Cannot Tell You

kind
Explainer
domain
Analytics
stage
Detection and tuning
read
3 min
assumes
No prior programme in place

User behaviour analytics is sold as detection of intent. It detects deviation from a baseline, which is a different and narrower thing.

Behavioural analytics is the standard answer to the limits of content inspection: instead of examining what is moving, examine whether the person's activity is unusual.

It genuinely adds something. It is also consistently oversold, in a way that produces investigations of innocent people.

What it actually does

Builds a statistical profile of normal activity — for an individual, and for their peer group — then scores deviation from it.

Typical inputs: volume of data accessed, systems touched, time of day, location, file operations, printing, use of removable media, failed access attempts.

Typical output: a risk score, sometimes with contributing factors listed.

What it is good at

Compromised accounts. This is the strongest use. An attacker using stolen credentials behaves differently from the legitimate user — different hours, different systems, different volumes. Behavioural deviation catches this well, and it is a category content inspection cannot address at all.

Volume anomalies. Someone downloading far more than they ever have. Crude and effective.

Access broadening. A person starting to touch systems outside their normal role, gradually. Hard to see any other way.

Correlating weak signals. Individually unremarkable events that together form a pattern: a resignation, followed by increased access, followed by removable media use.

What it is poor at

Intent. A score is not a state of mind. A high score means unusual, and unusual has many innocent explanations: a new project, covering for a colleague, a deadline, a role change, a reorganisation.

Baselines for people who change. New starters have no baseline. People who move roles invalidate theirs. Anyone whose work is genuinely varied looks anomalous constantly.

Low-and-slow activity. Someone taking a small amount regularly, within their normal patterns, is invisible to a system looking for deviation. This is also how a knowledgeable insider would behave, which limits the value precisely where it is most needed.

Small populations. Peer-group comparison requires a meaningful peer group. In a team of three, everyone is an outlier.

The problem with scores

A number between 0 and 100 conveys precision that is not there, and it invites treatment as evidence.

A score of 85 does not mean an 85% probability of wrongdoing. It means the person's activity was unusual relative to a model with assumptions you did not set, weighted in ways you cannot inspect, against a baseline that may be stale.

Use the contributing factors, not the score. "Accessed 12 systems outside normal role, transferred 3GB to removable media, three days after resignation" is something a human can evaluate. "Risk score 85" is not.

Never present a score as a finding to HR. It will be challenged and it will not survive.

Using it well

As a prioritisation layer, not a detector. It ranks what a human looks at first. It does not decide anything.

Combined with context that has meaning. A resignation date. A performance process. A role change. These external facts turn an anomaly into something interpretable.

With a stated review period for baselines, so a stale model does not generate months of noise after a reorganisation.

With awareness of its bias. Models trained on historical activity encode historical patterns. Someone with an unusual working pattern — different hours for caring responsibilities, a disability accommodation, a different time zone — will score higher permanently, for reasons that have nothing to do with risk.

That last point deserves more attention than it gets. A programme that repeatedly flags the same person because their legitimate working pattern is atypical is discriminating, whether or not anyone intended it. Check the distribution of your alerts against who is generating them.