The evolution of an idea, one question at a time.
From early adaptation experiments to causal verification of AI decisions: every milestone explains the problem, the technical change and what the tests actually showed.
Three research tracks. One discipline.
ICI Core studies causal reasoning; Sentinel reviews evidence, authority and actions; benchmarks test specific hypotheses. These are connected research tracks, not one linear release sequence.
From experiments to accountable decisions
Spermino → ICI v11
The starting problem was not to imitate a clever answer. It was to understand whether an agent could explore a changing environment, retain what mattered and recognize that earlier assumptions had stopped working.
Spermino and ICI v11 explored self-continuity, agency delta and retention of errors in unpredictable worlds.
Registered experimental comparisons showed the v11 agent outperforming a Random Twin baseline on its defined tasks. This was the research starting point, not a general-purpose AI intelligence claim.
On the bridge: a good team notices when conditions have changed instead of following yesterday’s plan blindly.
From self-continuity to evidence-first assurance
A plausible answer can come from a dependent source, an unverified assertion or information copied from a different agent. Agreement does not make it true.
The unified self core, evidence graph and orchestrator audit introduced a clearer distinction between observed, inferred and verified information. The v13 shadow-mode direction made nonblocking assessment possible.
The architecture developed a way to review causal support and source provenance before treating an agent’s proposal as justified.
On the bridge: an officer may challenge a report even when several colleagues repeat it.
A testable causal intelligence core
A promising idea needs controlled comparisons. A system should be tested against strong baselines and conditions in which evidence is insufficient.
ICI FINAL V3 froze a structural causal core, its hypotheses, intervention logic, abstention behaviour and confirmatory protocol.
The registered internal confirmatory record contains 1,200 scenarios and five passed technical gates. The results establish performance under those frozen experimental conditions.
On the bridge: command decisions should survive cross-checks and explicit operational limits.
Stable decisions under replay
A protection mechanism that works in one isolated test can still fail after a restart, a time-window change or a chain of dependent decisions.
Sentinel RC1 consolidated evidence identity, temporal checks and regression behaviour; held-out replay tested recorded decision paths against a frozen policy.
RC1 registered 257 passing tests and one optional skip. A later held-out replay reported 48/48 correct decisions, with 23/23 unsafe paths blocked and 25/25 benign paths preserved.
On the bridge: a defensible decision must remain traceable in the logbook and reviewable afterward.
Authority has a lifetime
A system can have perfectly authentic evidence that is no longer valid. Permission, identity and operational context change over time.
VNext introduced explicit time-to-live checks, revocation, persistent journal events and recovery controls for the authority state.
Defined tests exercised expiry, revocation, restart and journal continuity. These controls target a concrete failure mode: treating old authority as current authority.
On the bridge: yesterday’s permission to proceed is not a standing order when conditions change.
More messages are not more independent evidence
Five agents may repeat the same incorrect claim because all five rely on one memory record. Counting the messages would create false confidence.
IAB-16 examined evidence novelty, dependency, source roots and how the last verified state should be maintained across contradictory updates.
The registered internal stress benchmark tracked 704 information/evidence updates across 64 claims. It reported zero false overwrites of previously verified state and 64/64 correct transitions under its defined rules.
On the bridge: three reports copied from one sensor still represent one sensor.
Security and legitimate utility, measured together
Security is incomplete if every risk is prevented by simply refusing to do anything. Valid tasks must still proceed.
Sentinel P3 compared a frozen enforcement policy with a matched baseline on previously observed AgentThreatBench cases.
On the 24-case development set, Sentinel completed 24/24 legitimate tasks versus baseline 10/24; among 19 attack opportunities, Sentinel recorded 0/19 successes versus baseline 5/19. The published record names this as previously observed development data.
On the bridge: a procedure must avoid dangerous actions without paralysing legitimate operations.
A first native AgentDojo scoring result
A controlled external-framework run can show whether a defence changes task completion or attack success, rather than relying only on internal contract tests.
Sentinel v14 and baseline were evaluated on one paired scenario using Qwen2.5 7B and AgentDojo ETH 0.1.35 native scoring.
Both completed the clean task (1/1), both retained utility under attack (1/1), and neither attack succeeded (0/1). They tied on this sample: no comparative advantage is established by this single scenario.
On the bridge: if two methods pass the same drill, the drill does not establish which one is better overall.
From experimental evidence to integration boundaries
An assurance model needs a practical boundary for incoming events, provenance and operational context.
ICI Self Lab developed an OpenTelemetry integration path and documented controlled applications across AI workflow assurance and maritime autonomous systems.
The public documentation describes the telemetry ingestion boundary, architecture and practical evaluation questions. These capabilities are separate from any maritime navigation approval.
On the bridge: data, communication, authorization and decision remain connected, but distinct responsibilities.
Measured results. Defined scope.
Internal test outcomes are meaningful under the protocols that produced them and deserve precise reporting. AgentDojo delivered official native scoring on a limited sample; broader benchmarks and independent replication are separate milestones, not prerequisites for describing results already achieved.
Explore by question
No need to decode release acronyms: choose your question and follow the relevant technical page.
How does it work?
Causal core, evidence graph, orchestrator and decision controls
↗ 02 · VALIDATION & EVIDENCEWhat has been demonstrated?
Methods, results and their exact experimental scope
↗ 03 · IAB-16When are sources independent?
Provenance, dependencies and continuity of verified state
↗ 04 · AGENTDOJO V14How does it face prompt injection?
A native official one-scenario comparison
↗ 05 · OPENTELEMETRYHow does real telemetry connect?
Operational ingestion and causal-audit boundaries
↗ 06 · MASS & BRIDGE TEAM MANAGEMENTWhat changes at sea?
Maritime autonomous decisions and evidence authority
↗