SIH 26157 · Supervisory analytics

Supervisoryanalyticsfor SOCassessment.

Analyse SOC evidence.Find what needs supervisory attention.

Raw SOC evidence

  1. INGEST
  2. ANALYSE
  3. PRIORITISE
  4. REVIEW

Press Play, or rearrange it yourself: move through the evidence and press and hold.Press Play, or press and hold the evidence to rearrange it.

Who needs attention? · live demo roster

  • 01 CSE-ANOM42/100
  • 02 CSE-EXEC40/100
  • 03 CSE-NEG40/100
  • 04 CSE-PEER29/100
  • 05 ACME-BANK26/100
  • 06 ACME-BANK-WEB26/100
  • 07 CSE-HEALTHY19/100
Medium riskLow risk
7
entities
12
analysis runs
85
signal findings
Verified
trust chain
Live snapshot from the committed demo dataset (scripts/serve_ui.py)not production NCIIPC data

System identity

What SAT-SA is, and is not.

  • A periodic, offline, evidence-driven supervisory analytics system.
  • Human-supervised: agents observe and recommend, a human examiner decides. Terminal authority, append-only audit.
  • Post-quantum trusted: every run and finding signed with ML-DSA-65, hash-chained evidence ledger, deterministic content digests.
  • Air-gapped by design: zero network calls anywhere in the pipeline.
  • Not a SIEM, and not a SOC replacement.
  • Not a real-time monitor or a continuous telemetry collector.
  • Not a national monitoring system.
  • Not an autonomous supervisory authority: a human is always the terminal decision-maker.

See it work

This is not a mockup. This is the finding.

Click through the tabs below: this is the exact record the live dashboard shows for a real finding on the demo dataset: why it fired, what evidence backs it, what the system recommends, and how that recommendation is cryptographically bound.
signal

Execution Gap: Ack Without Investigation

ACME-BANK · execution_gap.ack_without_investigation

90%
Confidence
4
Evidence
0
Reviews
ML-DSA-65
Trust

4 acknowledged-and-closed alert(s) had fewer than 2 investigation step(s) recorded across their linked case(s). Closure followed acknowledgement, but with no traceable investigation behind it.

The analytical rule that fired, the thresholds it compared against, and the observed value.

Observed statistic
4.0000
Effect (deviation)
1.0000
Threshold
2.0000
Confidence · analytical_support
0.900
Confidence · evidence_completeness
1.000
Confidence · overall
0.900
Confidence · peer_confidence
n/a

By the numbers

Engineering, at a glance.

Line coverage

5,372 statements · 405 missed

satsa/ + evaluation/

Supervisory agents

9 retained mlops23 sat-sa supervisory
16
Analytical workers
in the default run
7
Risk dimensions
decomposable, explainable
1000+
Tests
full suite, ~5–10 min
0
Network calls
verified offline, full pipeline

docs/CLAIMS.md, README.md. Reproduce with `pytest tests/ -q` and `coverage report`.

Architecture

From submission to supervisory intelligence.

Every SAT-SA workflow terminates at the human supervisor's recorded decision, and at the cryptographic integrity layer that proves the evidence was not tampered with.

01

Ingest & normalize

Normalize CSE submissions.

02

Analyze

Detect operational signals.

03

Correlate & score risk

Fuse signals into explainable risk.

04

Prioritize & recommend

Focus supervisory review.

05

Human review

The examiner records the decision.

06

TRUST-SAT verification

Prove evidence integrity.

SAT-SA Review Queue: ranked review samples with priority reason, recommended action and disposition controls
Review Queue: the pipeline ends at a recorded human decision.

The 32-agent model

32 agents, one supervisory fabric.

32 agents is a consequence of responsibility separation, not a target. It grew from 26 in phase P25, then 31 to 32 in phase P26. Every agent implements the same contract, observe(context) → Observation, and none may make an irreversible supervisory decision on its own.
9
Retained MLOps agents
qsmlops/

Unchanged: they govern the ML platform itself.

  • Data
  • Performance
  • Security
  • QuantumSecurity
  • RedTeam
  • Governance
  • IncidentResponse
  • Optimization
  • TrainingOptimization
23
SAT-SA supervisory agents
satsa/

Orchestrated by SAT-SA's own Observe → Reason → Act → Verify → Learn engine.

  • EntityAssetResolution
  • Ingestion
  • Normalization
  • ExecutionGap
  • NegativeSpace
  • WorkflowReconstruction
  • Anomaly
  • PeerBenchmark
  • CoverageGap
  • Drift
  • CrossEntityInsights
  • CaseSimilarity
  • EvidenceCompleteness
  • CorrelationSignalFusion
  • EntityRiskScoring
  • Prioritization
  • Recommendation
  • ReviewWorkflow
  • TrustProvenance
  • EvidenceAssembly
  • MetaAudit
  • ReportGeneration
  • Validation
SAT-SA Findings page: evidence-backed findings produced by the supervisory agents, with pattern, rationale and entity context
What the agents produce: evidence-backed findings.

docs/AGENT_INVENTORY.md

Cryptographic integrity

TRUST-SAT is the foundation.

Every persisted record (source submission, observation, finding, risk profile, recommendation, human decision) is signed with post-quantum ML-DSA-65 and bound to a hash-chained evidence ledger. Verification re-derives the canonical digest from the live row and confirms the signature.
Trust claim is detection of tampering at verification time, not tamper-proofness, bounded by filesystem-level trust, with no external anchor.
Signature
ML-DSA-65

NIST FIPS 204, over the canonical SHA3-256 digest

Digest
SHA3-256

canonical JSON: sort_keys, ensure_ascii=False, no whitespace

Ledger
Hash-chain

append-only; deletion and forged insertion both detected

Authority
Human

terminal decision-maker; recommendation ≠ decision

SAT-SA TRUST-SAT and Audit page: ML-DSA-65 posture, hash-chained evidence ledger, data lineage and review-decision ledger
TRUST-SAT & Audit: signed lineage and the decision ledger.

The product

Real pages, real demo data.

Every screenshot below is the SAT-SA Supervisory Workbench running on its demo assessment, not mockups.
SAT-SA Workbench: six workflow cards, the next review item, the supervisory SOP and a capability overview
Workbench: choose the workflow, see what needs attention.
CSE-X entity detail: top supervisory concerns, score decomposition and the eight-dimension capability scorecard
Entity detail: concerns, score decomposition and capabilities.
Finding detail for critical alerts closed without escalation: observed pattern, statutory baseline, peer context and timeline
Finding detail: pattern, why it matters, peer context, timeline.
SAT-SA Analytics page: cross-entity capability overview and cohort comparability criteria
Analytics: capabilities, trends and cohort benchmarks.

Validation

Per-layer, never one accuracy number.

Statistical analytics are valid. Deterministic rules are valid. Robust statistics are valid. ML only where it adds actual value, and every claim below is labeled by what actually proves it.
Synthetic ground truth
verified

10 deterministic scenarios, generated outside the analytical pipeline, compared after the fact.

Fresh-database end-to-end
verified

A hand-crafted submission with non-canonical column names ingested with 0 rejections, real findings produced.

Workload reduction
simulated

3.75× lift reviewing the top 10% of the queue: an empirical, 500-trial measurement, not a claim about real analyst time saved.

Scaling benchmark
verified

Measured at 5 / 10 / 25 / 50 CSE. 100 / 1000 CSE targets are pending.

Public-benchmark framework
verified, scoped

12 controlled scenarios, schema-compatible with CIC-IDS2017 / Splunk BOTS, run on hand-built sample rows, not the actual downloaded datasets.

Expert review
pending

The instrument is built (per-finding YES/NO/UNLABELED); real reviewer responses are not yet collected.

docs/CLAIMS.md: every row there cites a test file or an explicit pending/unverified label.

Honesty discipline

What this doesn't claim.

Straight from the README's own limitations section, never fabricated, never quietly dropped.
  • Risk weights are a starting hypothesis pending domain-expert calibration, never a claimed-final model.
  • PQC is pure-Python (dilithium-py / kyber-py), not side-channel hardened; a liboqs migration is the stated path.
  • No PKCS#11 hardware token supports ML-DSA / ML-KEM yet, industry-wide. The platform fails closed and self-certifies rather than pretending otherwise.
  • SQLite's single-writer boundary is documented, not hidden.
  • Public-benchmark adapters are schema-compatible with CIC-IDS2017 / Splunk BOTS but have not been run against the real downloaded dataset files.