The full evaluation.
Method, protocol, results, and what was missed. Every figure here comes from a public dataset, so none of it has to be taken on trust.
If you want the argument rather than the arithmetic, structural detection is explained here.
How the numbers were produced.
The dataset
CICIDS2017, Canadian Institute for Cybersecurity, University of New Brunswick. Five days of capture from a purpose-built testbed, replayed from the labelled bidirectional flow CSVs: 1.15 GB across 8 files, 17 internal hosts on a single /24, and 2.8 million flows.
It is a public benchmark, so every figure on this page can be reproduced independently rather than taken on trust. We chose it over UNSW-NB15 for one decisive reason: it retains source and destination addresses and real timestamps. UNSW-NB15 strips the address columns, so it cannot exercise a detector whose premise is relationships between identities.
unb.ca/cic/datasets/ids-2017.htmlFrozen baseline
The baseline is warmed once on two benign days and then frozen. Every attack day is replayed independently against that identical state, so no day can contaminate another and no attack informs the detection of any other.
Benign control day
A day containing no attacks is replayed against its own baseline on every run. Its output is the false positive figure: 3 cases. Without a benign control, a recall number says nothing about whether the system is simply noisy.
Hosts, not just episodes
The measure that matters is whether every attacked host got reported, because that has no threshold in it. A case is one investigation. An observation is a lone departure that had nothing to correlate with, reported a tier lower. Episodes sit alongside: contiguous attack traffic is one attack, ended by a 25 minute gap.
Day by day
| Replay | Episodes | Caught | Recall | Time to detect | Hosts in a case | Cases |
|---|---|---|---|---|---|---|
| Monday (benign control) | 0 | – | – | – | – | 3 |
| Tuesday: brute force | 2 | 2 | 100% | <1 min | 2 / 2 | 12 |
| Wednesday: DoS | 4 | 4 | 100% | <1 min | 3 / 3 | 12 |
| Thursday: web attacks | 1 | 1 | 100% | <1 min | 2 / 2 | 9 |
| Thursday: infiltration | 1 | 1 | 100% | 12 min | 1 / 1 | 24 |
| Friday: botnet C2 | 1 | 1 | 100% | 14 min | 6 / 7 | 6 |
| Friday: port scan | 3 | 3 | 100% | <1 min | 2 / 2 | 8 |
| Friday: DDoS | 1 | 1 | 100% | <1 min | 2 / 2 | 11 |
| All attack replays | 13 | 13 | 100% | <1 min | 18 / 1919 / 19 surfaced | 85 |
Every attacked host, reported
Every attack was detected, and every attacked host was reported. Brute force, DoS, web attacks, infiltration, botnet C2, port scan and DDoS were all found, with no signatures and no prior knowledge of any of them.
Eighteen of the nineteen arrived as cases. The nineteenth departed on its own, and a case is built from two connected nodes, so it is emitted as an observation: one tier below a case, one severity down, outside the analyst queue. Detected and unreported is the one outcome we are not willing to ship.
Observations do not alert, so the queue is unchanged at 3 cases on the benign control day. They sit in the same place as cases, one query away, which is where you want them when a case names a host and you go looking. Most of the time an intrusion moves more than one thing at once and correlates by itself. The exception is the one worth catching: something beaconing quietly and doing nothing else.
By attack class
The same detector caught every class, with no tuning between them and no signature anywhere. That includes a C2 channel whose only distinguishing feature was that one host stopped talking to 3,000 external peers and started talking to one.
Port scan
3 / 3
<1 min
DoS
4 / 4
<1 min
Brute force
2 / 2
<1 min
Web attacks
1 / 1
<1 min
Infiltration
1 / 1
12 min
DDoS
1 / 1
<1 min
C2 beaconing
1 / 1
14 min
What one benchmark can and cannot show.
Stated plainly, because a reviewer will establish it anyway and it is better read here than discovered later.
It is one dataset
Sixteen estate endpoints and 2.8 million flows, captured on a purpose-built testbed. A real estate has proxies, backup windows, vulnerability scanners and CI runners, all structurally noisy in ways a benchmark is not.
It is three quarters attack days
Seven of the eight replay segments carry attack traffic by construction. That is why the week-wide routing figure is a floor rather than an expectation, and why the benign control day is the closer analogue to a real week.
It cannot exercise everything
CICIDS2017 carries no process telemetry, so the local-only attack classes are untested here. Identity and lateral movement handling is built but this dataset cannot reach it.
It is a replay, not an operation
Deterministic replay against a frozen baseline is the right way to compare runs, and it is not the same as running continuously on a live estate for a quarter.
Warm-up is reasoned, not measured
The planning figure of two to four weeks comes from reasoning about weekly and monthly periodicity in a real estate, not from a measurement on this data.
The dataset is old
CICIDS2017 is from 2017 and is heavily used in academic work. It is public and reproducible, which is why it is here, and it is not a substitute for current traffic.
Reproduce it, or test it on your own data.
CICIDS2017 is public and the method above is stated in full. The more useful test is your own estate: run SIET alongside your existing alerting, switch nothing off, and compare cases raised, alerts missed by each, and indexed volume.