SIH 26157 · Supervisory analytics
Supervisoryanalyticsfor SOCassessment.
Analyse SOC evidence.Find what needs supervisory attention.
Raw SOC evidence
- INGEST
- ANALYSE
- PRIORITISE
- REVIEW
Press Play, or rearrange it yourself: move through the evidence and press and hold.Press Play, or press and hold the evidence to rearrange it.
Who needs attention? · live demo roster
- 01 CSE-ANOM42/100
- 02 CSE-EXEC40/100
- 03 CSE-NEG40/100
- 04 CSE-PEER29/100
- 05 ACME-BANK26/100
- 06 ACME-BANK-WEB26/100
- 07 CSE-HEALTHY19/100
System identity
What SAT-SA is, and is not.
- A periodic, offline, evidence-driven supervisory analytics system.
- Human-supervised: agents observe and recommend, a human examiner decides. Terminal authority, append-only audit.
- Post-quantum trusted: every run and finding signed with ML-DSA-65, hash-chained evidence ledger, deterministic content digests.
- Air-gapped by design: zero network calls anywhere in the pipeline.
- Not a SIEM, and not a SOC replacement.
- Not a real-time monitor or a continuous telemetry collector.
- Not a national monitoring system.
- Not an autonomous supervisory authority: a human is always the terminal decision-maker.
See it work
This is not a mockup. This is the finding.
Execution Gap: Ack Without Investigation
ACME-BANK · execution_gap.ack_without_investigation
4 acknowledged-and-closed alert(s) had fewer than 2 investigation step(s) recorded across their linked case(s). Closure followed acknowledgement, but with no traceable investigation behind it.
The analytical rule that fired, the thresholds it compared against, and the observed value.
- Observed statistic
- 4.0000
- Effect (deviation)
- 1.0000
- Threshold
- 2.0000
- Confidence · analytical_support
- 0.900
- Confidence · evidence_completeness
- 1.000
- Confidence · overall
- 0.900
- Confidence · peer_confidence
- n/a
By the numbers
Engineering, at a glance.
5,372 statements · 405 missed
satsa/ + evaluation/
Supervisory agents
docs/CLAIMS.md, README.md. Reproduce with `pytest tests/ -q` and `coverage report`.
Architecture
From submission to supervisory intelligence.
Every SAT-SA workflow terminates at the human supervisor's recorded decision, and at the cryptographic integrity layer that proves the evidence was not tampered with.
01
Ingest & normalize
Normalize CSE submissions.
02
Analyze
Detect operational signals.
03
Correlate & score risk
Fuse signals into explainable risk.
04
Prioritize & recommend
Focus supervisory review.
05
Human review
The examiner records the decision.
06
TRUST-SAT verification
Prove evidence integrity.

The 32-agent model
32 agents, one supervisory fabric.
Unchanged: they govern the ML platform itself.
- Data
- Performance
- Security
- QuantumSecurity
- RedTeam
- Governance
- IncidentResponse
- Optimization
- TrainingOptimization
Orchestrated by SAT-SA's own Observe → Reason → Act → Verify → Learn engine.
- EntityAssetResolution
- Ingestion
- Normalization
- ExecutionGap
- NegativeSpace
- WorkflowReconstruction
- Anomaly
- PeerBenchmark
- CoverageGap
- Drift
- CrossEntityInsights
- CaseSimilarity
- EvidenceCompleteness
- CorrelationSignalFusion
- EntityRiskScoring
- Prioritization
- Recommendation
- ReviewWorkflow
- TrustProvenance
- EvidenceAssembly
- MetaAudit
- ReportGeneration
- Validation

docs/AGENT_INVENTORY.md
Cryptographic integrity
TRUST-SAT is the foundation.
NIST FIPS 204, over the canonical SHA3-256 digest
canonical JSON: sort_keys, ensure_ascii=False, no whitespace
append-only; deletion and forged insertion both detected
terminal decision-maker; recommendation ≠ decision

The product
Real pages, real demo data.




Validation
Per-layer, never one accuracy number.
10 deterministic scenarios, generated outside the analytical pipeline, compared after the fact.
A hand-crafted submission with non-canonical column names ingested with 0 rejections, real findings produced.
3.75× lift reviewing the top 10% of the queue: an empirical, 500-trial measurement, not a claim about real analyst time saved.
Measured at 5 / 10 / 25 / 50 CSE. 100 / 1000 CSE targets are pending.
12 controlled scenarios, schema-compatible with CIC-IDS2017 / Splunk BOTS, run on hand-built sample rows, not the actual downloaded datasets.
The instrument is built (per-finding YES/NO/UNLABELED); real reviewer responses are not yet collected.
docs/CLAIMS.md: every row there cites a test file or an explicit pending/unverified label.
Honesty discipline
What this doesn't claim.
- Risk weights are a starting hypothesis pending domain-expert calibration, never a claimed-final model.
- PQC is pure-Python (dilithium-py / kyber-py), not side-channel hardened; a liboqs migration is the stated path.
- No PKCS#11 hardware token supports ML-DSA / ML-KEM yet, industry-wide. The platform fails closed and self-certifies rather than pretending otherwise.
- SQLite's single-writer boundary is documented, not hidden.
- Public-benchmark adapters are schema-compatible with CIC-IDS2017 / Splunk BOTS but have not been run against the real downloaded dataset files.