✨
CoPilot — AI-Assisted Troubleshooting
Ask operational questions, run AI Analyze on problem pods, and get root cause, evidence, confidence, risk, commands, YAML fixes, and next steps from your local LLM.
🔍
Pilot — SRE Troubleshooting Browser
Browse Deployments, StatefulSets, DaemonSets, Jobs, Services, Ingresses, ConfigMaps, Secrets, and PVCs across namespaces with live YAML and log viewers.
🤖
AutoPilot — AI Fixer Self-Healing
Policy-gated remediation in off, dry-run, or active mode. Confidence floors, namespace blocklists, action allow-lists, cooldowns, hourly caps, and a pause kill switch keep automation auditable.
📊
Cluster Resource Gauges
Live CPU, memory, and storage usage with per-StorageClass breakdown. Longhorn-aware for accurate physical disk readings, not just bound PVC totals.
📱
Native iOS Troubleshooting
A SwiftUI mobile app connects to the same KubePilot APIs for health cards, pod search, pod detail tabs, live logs, events, RCA reports, Face ID, widgets, and Siri/App Intents.
📋
Runbooks & YAML Workflows
7 pre-built diagnostic workflows ship in the box. Drop YAML runbooks into a directory for fsnotify-driven hot reload — user runbooks override builtins by ID.
🌐
LAN, WAN & Tunnel Node IPs
Node views classify private LAN/VPC addresses, public WAN addresses, and WireGuard or flannel tunnel endpoints using Kubernetes addresses plus k3s and flannel annotations.
📜
Pod Detail Workbench
Open a pod to inspect overview, containers, events, logs, sanitized YAML, metrics context, restart counts, uptime, and AI analysis without jumping between kubectl commands.
🔗
Pod & Service Port-Forwarding
Open a tunnel into the cluster directly from the UI and reach the forwarded port through the dashboard reverse-proxy. Sessions are listed, cancellable, and auditable.
💾
Durable RCA History (SQLite)
Optional embedded SQLite store persists RCA reports and anomalies across restarts. Configurable retention, WAL journaling, no CGO — same single binary.
🔔
Alerts, Events & RCA Timeline
Review warnings, normal events, anomalies, and RCA reports as a searchable operational timeline. Slack notifications can post formatted incident cards for high-severity findings.
🔌
Service Topology Map
ArgoCD-style canvas links Ingress to Service to Workload to Pod with status colours, ports, and external IPs. Switch namespaces without leaving the page.
🤙
Copy Prompt & AI Health
Copy the exact CoPilot prompt from any analysis surface, and check the AI health chip so you know Ollama is ready before an incident depends on it.
☁️
Multi-Cluster Switching
Upload kubeconfigs, switch contexts, and route every dashboard query to the new cluster — no restart required.
🔒
Security-First Design
Optional auth, read-only defaults, mutation gates, CR-code approval for risky actions, and CORS policies. Safe for production from day one.
🚀
MCP Agent Protocol
Built-in MCP server for multi-cluster agent orchestration, remote AI coordination, and programmatic access.