CausalPilot
Evidence-gated experimentation with deterministic causal tools and selective action rights.
Research question
Under limited experiment and human-attention budgets, what evidence should an agent need before it is allowed to act? CausalPilot separates a language model’s proposal from the deterministic gates that own diagnosis, probing, human escalation, deployment, pause, rollback, and abstention rights.
Approach
The project combines a FastAPI state machine, deterministic statistical tools, explicit model/causal/data/environment risk signals, and a Next.js control room. Its checked-in benchmark isolates hidden simulator truth from policies and stores the seeds, raw decisions, intervals, fixtures, and configuration fingerprints needed to audit the reported run.
Evidence and boundaries
The current artifact is a controlled-simulation engineering benchmark, not a causal research conclusion or live business-lift claim. Its named “LLM” benchmark policies are decision-rule proxies; the canonical run did not call model APIs. The checked-in results do not establish matched-coverage superiority, calibrated risk probabilities, real-model planning quality, or safe enterprise deployment.
Artifacts and provenance
- Live deterministic replay
- Source code
- Technical report
- Claim boundaries
- Displayed screenshot source
- Code, report, and generated project media are distributed under the repository’s MIT license.