Skip to main content
AI & Operations

NEXUS — Autonomous Homelab SRE

An evidence-first autonomous SRE agent built around four hard rules — evidence over vibes, allowlisted typed tools, approval enforced outside the model, and explainable reasoning. Investigates incidents the way a careful operator would, verified live against PostgreSQL and Redis.

Engineering ProjectOngoing engineering projectTags: LLM Agents, SRE, Evidence-Based AI
Context

What was happening.

An AI SRE for real infrastructure that has to show its work: it observes the live system, collects evidence with typed tools, forms hypotheses, tests them, and only then explains a root cause — and it refuses to sound confident when it cannot prove something.

The problem

What needed to change.

  • Most “AI ops” demos ask an LLM to guess and let confidence stand in for correctness.
  • Without structured tools and permissions, agents either do nothing useful or have an alarming amount of power — arbitrary bash, rm -rf — with no audit trail.
  • Explainability gets sacrificed: conclusions arrive without observable evidence or a reasoning path you can verify.
The approach

How it was done.

Evidence over vibes

Confidence is derived from observations, contradictions and required-evidence coverage — never invented by the model. When NEXUS cannot verify something, the honest output is “I could not verify disk usage because the host did not respond” — not “disk usage is normal.”

Confidence derived from observations, never invented
Allowlist, not denylist

No arbitrary shell. Every capability is a typed tool with a permission class (READ_ONLY, LOW_RISK, REQUIRES_APPROVAL, FORBIDDEN). The model cannot bypass the permission layer because authorization is enforced outside the LLM.

Authorization is enforced outside the LLM — it cannot bypass the permission layer
Approval before risk

REQUIRES_APPROVAL actions cannot run until a human says yes. Default mode is READ_ONLY — no destructive operations until the permission system is explicitly enabled.

Default mode is READ_ONLY — destructive ops must be explicitly enabled
Explainable reasoning

The console shows reasoning summaries and the evidence collected at every step — never hidden chain-of-thought.

Reasoning summaries and evidence at every step — never hidden chain of thought
Investigation lifecycle

classify → context → evidence → hypothesis → test → root cause → remediation → approval → execute → verify → close. When evidence is insufficient the reasoning loop iterates instead of inventing a root cause just to finish.

When evidence is insufficient the loop iterates instead of inventing a root cause
Phase 1 — verified foundation

FastAPI application factory with a live /api/v1/health database check, PostgreSQL 16 persistence (SQLAlchemy 2.0 async, Alembic migrations), Redis, structured logging with structlog, a sandbox Docker Compose stack, and CI running ruff, strict mypy and pytest against real PostgreSQL and Redis — with graceful degradation when a dependency is down.

CI runs ruff, strict mypy and pytest against real PostgreSQL and Redis
Measured results

The numbers that came out.

READ_ONLY
Default mode

No destructive ops until the permission system is enabled.

4
Permission classes

READ_ONLY, LOW_RISK, REQUIRES_APPROVAL, FORBIDDEN.

Complete
Phase 1

Verified against live PostgreSQL 16 + Redis 7.

Evidence-backed
Root causes

Hypotheses tested against the live system before claiming anything.

Outcomes

Where it landed.

  • An agent architecture where sounding confident is never confused with being correct — the core fix for the AI-ops hype problem.
  • Safety enforced structurally: READ_ONLY by default, four permission classes, human approval outside the model, and a complete audit trail.
  • Phase 1 foundation verified against live PostgreSQL and Redis; subsequent phases — discovery, typed tools, LangGraph agent, evaluation framework — documented and scheduled in verifiable milestones.
  • An honest feature ledger: everything is marked REAL, SIMULATED, MOCK or NOT IMPLEMENTED — nothing is claimed before it is tested.
Stack & tools
PythonFastAPILangGraphPostgreSQL 16Redis 7SQLAlchemy 2.0 (async)AlembicDockerstructlogpydantic-settings
View source on GitHub
Start a conversation

Your situation probably looks different — but the method doesn't.

Understand the work, decide the approach, deliver and measure. A free conversation establishes whether it applies to your problem.

or email joseph.gitau.c@gmail.com