ICI SELF LAB
ICI SELF LAB · RESEARCH JOURNEY

The evolution of an idea, one question at a time.

From early adaptation experiments to causal verification of AI decisions: every milestone explains the problem, the technical change and what the tests actually showed.

CORE · SENTINEL · EVALUATION

Three research tracks. One discipline.

ICI Core studies causal reasoning; Sentinel reviews evidence, authority and actions; benchmarks test specific hypotheses. These are connected research tracks, not one linear release sequence.

RESEARCH MILESTONES EXPLAINED

From experiments to accountable decisions

01INITIAL RESEARCHORIGIN · ADAPTIVE LEARNING

Spermino → ICI v11

Can a system learn when the rules stop repeating?
THE QUESTION

The starting problem was not to imitate a clever answer. It was to understand whether an agent could explore a changing environment, retain what mattered and recognize that earlier assumptions had stopped working.

WHAT CHANGED

Spermino and ICI v11 explored self-continuity, agency delta and retention of errors in unpredictable worlds.

WHAT WE ACHIEVED

Registered experimental comparisons showed the v11 agent outperforming a Random Twin baseline on its defined tasks. This was the research starting point, not a general-purpose AI intelligence claim.

THE BRIDGE-TEAM ANALOGY

On the bridge: a good team notices when conditions have changed instead of following yesterday’s plan blindly.

Explore the evidence · CAUSAL ARCHITECTURE ↗
02v12 · v12.1 · v13SENTINEL · ARCHITECTURE

From self-continuity to evidence-first assurance

Can a system question the evidence behind an AI action?
THE QUESTION

A plausible answer can come from a dependent source, an unverified assertion or information copied from a different agent. Agreement does not make it true.

WHAT CHANGED

The unified self core, evidence graph and orchestrator audit introduced a clearer distinction between observed, inferred and verified information. The v13 shadow-mode direction made nonblocking assessment possible.

WHAT WE ACHIEVED

The architecture developed a way to review causal support and source provenance before treating an agent’s proposal as justified.

THE BRIDGE-TEAM ANALOGY

On the bridge: an officer may challenge a report even when several colleagues repeat it.

Explore the evidence · METHOD & ARCHITECTURE ↗
03ICI FINAL V3 · 2026CAUSAL CORE · PARALLEL TRACK

A testable causal intelligence core

Can causal reasoning be evaluated against clear alternatives?
THE QUESTION

A promising idea needs controlled comparisons. A system should be tested against strong baselines and conditions in which evidence is insufficient.

WHAT CHANGED

ICI FINAL V3 froze a structural causal core, its hypotheses, intervention logic, abstention behaviour and confirmatory protocol.

WHAT WE ACHIEVED

The registered internal confirmatory record contains 1,200 scenarios and five passed technical gates. The results establish performance under those frozen experimental conditions.

THE BRIDGE-TEAM ANALOGY

On the bridge: command decisions should survive cross-checks and explicit operational limits.

Explore the evidence · V3 VALIDATION ↗
04RC1 · SEPTEMBER–OCTOBER 2026SENTINEL · RUNTIME & INTEGRITY

Stable decisions under replay

Will the same evidence produce the same controlled decision?
THE QUESTION

A protection mechanism that works in one isolated test can still fail after a restart, a time-window change or a chain of dependent decisions.

WHAT CHANGED

Sentinel RC1 consolidated evidence identity, temporal checks and regression behaviour; held-out replay tested recorded decision paths against a frozen policy.

WHAT WE ACHIEVED

RC1 registered 257 passing tests and one optional skip. A later held-out replay reported 48/48 correct decisions, with 23/23 unsafe paths blocked and 25/25 benign paths preserved.

THE BRIDGE-TEAM ANALOGY

On the bridge: a defensible decision must remain traceable in the logbook and reviewable afterward.

Explore the evidence · REPLAY EVIDENCE ↗
05VNEXT · OCTOBER 2026SENTINEL · TEMPORAL AUTHORITY

Authority has a lifetime

What if an authorization was valid yesterday but has now expired?
THE QUESTION

A system can have perfectly authentic evidence that is no longer valid. Permission, identity and operational context change over time.

WHAT CHANGED

VNext introduced explicit time-to-live checks, revocation, persistent journal events and recovery controls for the authority state.

WHAT WE ACHIEVED

Defined tests exercised expiry, revocation, restart and journal continuity. These controls target a concrete failure mode: treating old authority as current authority.

THE BRIDGE-TEAM ANALOGY

On the bridge: yesterday’s permission to proceed is not a standing order when conditions change.

Explore the evidence · TEMPORAL AUTHORITY ↗
06IAB-16 · OCTOBER 2026RESEARCH · EVIDENCE INDEPENDENCE

More messages are not more independent evidence

Can several confirmations come from a single upstream source?
THE QUESTION

Five agents may repeat the same incorrect claim because all five rely on one memory record. Counting the messages would create false confidence.

WHAT CHANGED

IAB-16 examined evidence novelty, dependency, source roots and how the last verified state should be maintained across contradictory updates.

WHAT WE ACHIEVED

The registered internal stress benchmark tracked 704 information/evidence updates across 64 claims. It reported zero false overwrites of previously verified state and 64/64 correct transitions under its defined rules.

THE BRIDGE-TEAM ANALOGY

On the bridge: three reports copied from one sensor still represent one sensor.

Explore the evidence · IAB-16 RESEARCH ↗
07SENTINEL P3 · OCTOBER 2026BENCHMARK · AGENTTHREATBENCH

Security and legitimate utility, measured together

Does blocking unsafe requests also block useful work?
THE QUESTION

Security is incomplete if every risk is prevented by simply refusing to do anything. Valid tasks must still proceed.

WHAT CHANGED

Sentinel P3 compared a frozen enforcement policy with a matched baseline on previously observed AgentThreatBench cases.

WHAT WE ACHIEVED

On the 24-case development set, Sentinel completed 24/24 legitimate tasks versus baseline 10/24; among 19 attack opportunities, Sentinel recorded 0/19 successes versus baseline 5/19. The published record names this as previously observed development data.

THE BRIDGE-TEAM ANALOGY

On the bridge: a procedure must avoid dangerous actions without paralysing legitimate operations.

Explore the evidence · AGENTTHREATBENCH P3 ↗
08SENTINEL v14 × AGENTDOJO · OCTOBER 2026BENCHMARK · OFFICIAL NATIVE SCORING

A first native AgentDojo scoring result

What happens when Sentinel and a baseline face the same native test?
THE QUESTION

A controlled external-framework run can show whether a defence changes task completion or attack success, rather than relying only on internal contract tests.

WHAT CHANGED

Sentinel v14 and baseline were evaluated on one paired scenario using Qwen2.5 7B and AgentDojo ETH 0.1.35 native scoring.

WHAT WE ACHIEVED

Both completed the clean task (1/1), both retained utility under attack (1/1), and neither attack succeeded (0/1). They tied on this sample: no comparative advantage is established by this single scenario.

THE BRIDGE-TEAM ANALOGY

On the bridge: if two methods pass the same drill, the drill does not establish which one is better overall.

Explore the evidence · AGENTDOJO REPORT ↗
09PRODUCTION RESEARCH LINEINTEGRATION · OPERATIONAL USE CASES

From experimental evidence to integration boundaries

How can an operational system bring evidence into causal auditing?
THE QUESTION

An assurance model needs a practical boundary for incoming events, provenance and operational context.

WHAT CHANGED

ICI Self Lab developed an OpenTelemetry integration path and documented controlled applications across AI workflow assurance and maritime autonomous systems.

WHAT WE ACHIEVED

The public documentation describes the telemetry ingestion boundary, architecture and practical evaluation questions. These capabilities are separate from any maritime navigation approval.

THE BRIDGE-TEAM ANALOGY

On the bridge: data, communication, authorization and decision remain connected, but distinct responsibilities.

Explore the evidence · TELEMETRY INTEGRATION ↗

Measured results. Defined scope.

Internal test outcomes are meaningful under the protocols that produced them and deserve precise reporting. AgentDojo delivered official native scoring on a limited sample; broader benchmarks and independent replication are separate milestones, not prerequisites for describing results already achieved.