Not just who's using CPU, but who's causing stalls. Linnix pairs the workload that is stalling (measured by PSI) with the workload inducing it (observed via eBPF) — and hands that evidence to you, your dashboards, or your AI agent.
CPU utilization tells you who is busy. Pressure Stall Information tells you who is waiting. Linnix pairs the victim's PSI with the offender inducing it — observed via eBPF on the same node.
THE OPERATING PRINCIPLE
PSI measures stall, not usage.
Drag the neighbour's load up. The victim's CPU barely moves — so a CPU dashboard stays green — while the share of time it spends stalled climbs.
40% CPU · 60% PSISLO risk
100% CPU · 5% PSIbusy, progressing
One node, two pods
illustrative model · not a measurement
node-7
media/image-resizer
neighbour
CPU
20%
PSI
2%
no attribution
payments/payment-api
serving · has an SLO
CPU
42%
PSI
3%
victim CPU42%
victim PSI3%
neighbour CPU20%
Both pods are progressing. Nothing to attribute.
PSI = share of time tasks were stalled waiting on a resource. The link is what Linnix adds: the victim's stall paired with the neighbour that contended with it.
02 / WATCH
Linnix in 90 seconds.
The short version: why utilization misses stalls, and how Linnix pairs a stalling workload with the one causing it.
Explainer video coming soonMeanwhile, the console above replays four real result states.
03 / REAL RESULT STATES
Four answers. All of them honest.
Linnix doesn't turn uncertainty into a confidence score. Every investigation ends in one of these states, and each says exactly what was measured, what was weighted, and what wasn't known.
Likely offender
Offenders ranked by their share of the stall that could be pinned on a neighbour. The victim's stall is measured by PSI; Linnix splits it across neighbours by a weighting of CPU share, fork rate and short-job churn, and shows you the inputs. The denominator is named, so 76% never reads as 76% of everything.
Victim: payments/payment-api lost 2.6s to stalls across 7 detection windows.
2.1s of that is attributed to neighbours; the percentages below split that figure.
Likely offender: media/image-resizer — 76% of attributed stall
Attributed stall: 1.6s across 6 windows
Dominant signal: CPU noisy neighbour
Evidence: peak CPU share 0.71, 186 forks, 0 short jobs
Unmeasured contenders
Rows that predate per-offender splitting have an unknown contribution — not zero. If one remains, Linnix refuses to crown anyone. An individual offender whose share can't be computed is reported as share unknown, never 0%.
Measured contributors, ranked:
media/image-resizer — 61% of attributed stall (1.2s across 5 windows, CPU noisy neighbour)
Unmeasured contenders: their rows predate per-offender stall splitting, so their contribution is unknown and may exceed any figure above.
batch/etl-runner — fork storm, blamed in 3 windows, peak CPU share 0.22
No single offender can be named while an unmeasured contender remains: it cannot be ranked against the rest.
No contention attributed
A real, useful result: the neighbours are ruled out, so you stop chasing them and look at the pod itself.
Investigation: shop/checkout over the last 15m
No contention attributed to any neighbour in this window.
The pod may still be slow — this only rules out other workloads on the node as the cause. Look at the pod's own limits, throttling and workload next.
Degraded mode
If the eBPF probes fail to attach, cognitod keeps serving PSI and /readyz reports NOT ready (503) instead of pretending to be healthy.
GET /readyz → 503 Service Unavailable
{
…
"kernel_instrumentation": "unavailable",
"ready": false,
"reason": "eBPF probes are not attached; running userspace-only, so no per-process stall attribution is being produced. Check the kernel version (5.12+ x86_64 / 5.18+ arm64), BTF availability, tracefs mount, and CAP_BPF/CAP_PERFMON.",
…
"transport": "userspace"
}
04 / MCP FOR AGENTS
Give your AI agent the machine.
“Reasoning is not the scarce thing any more… What it cannot do is look at the machine.”
Linnix supplies the facts; the model reasons. Five tools over stdio, each taking a detail argument so an agent pays for depth only once it has decided the host matters. If cognitod is unreachable, tools return a structured error rather than failing.
Tool
Endpoint
Answers
linnix_system_health
/status, /system
Is this host under pressure now; is the daemon healthy?
linnix_investigate_contention
/attribution
Which workloads contended with this pod while it stalled, on what evidence?
linnix_explain_process
/processes/{pid}, /graph/{pid}
What is this PID, what started it, what did it start?
payments/payment-api stalled across 7 detection window(s) in the last 20m. The largest contender was media/image-resizer with 76% of the stall attributed to neighbours, dominant signal: noisy_neighbor. This is contention attribution, not proven cause.
Investigation: payments/payment-api over the last 20m
Victim: payments/payment-api lost 2.6s to stalls across 7 detection windows.
2.1s of that is attributed to neighbours; the percentages below split that figure.
Likely offender: media/image-resizer — 76% of attributed stall
Attributed stall: 1.6s across 6 windows
Dominant signal: CPU noisy neighbour
Evidence: peak CPU share 0.71, 186 forks, 0 short jobs
Also contributing:
batch/etl-runner — 24% (500ms, fork storm)
This is contention attribution, not proven causality. To confirm, change one
thing — move the offender, or give it a CPU limit — and check whether the
victim's stall falls.
Evidence link: http://127.0.0.1:3000/attribution?pod=payment-api&namespace=payments&from=1790810844&to=1790812044&max_id=4096
ranked offenders, shares, windows, peak CPU share, dominant signal
The summary tier's own words. In full: this is contention attribution, not proven causality — confirm by changing one thing and watching the victim's stall.
blame_score is not a confidence. It is an unbounded weight — (CPU share + fork score + short-job score) × stall seconds — and each offender's attributed_stall_us is the victim's PSI-measured stall split in proportion to it. The inputs sit right beside it so you can check the arithmetic. Abridged to one row; the real raw tier returns every row and field cognitod sends.
05 / DETECTIONS
Signals that describe what the kernel is doing.
Each detection keeps the underlying measurements, so an alert is the start of an investigation — not the end of the evidence. Monitor-first: nothing is enforced unless you configure it.
D-01
Circuit breaker
CPU PSI > 40% and CPU usage > 90%, 15s grace Surfaces sustained distress. Enforcement is opt-in and requires explicit configuration or human approval.
D-02
Fork storm
configurable fork-rate thresholds Names rapid process creation — and, via the process tree, the parent doing it.
D-03
Memory leak
Sustained RSS growth over time, tied back to the process and workload.
D-04
Short-lived job churn
Rapid exec/exit churn that aggregate dashboards tend to smooth away.
D-05
Noisy neighbours
Pairs the offender contending for shared capacity with the victim losing time.
D-06
PSI saturation
CPU, memory and I/O stall pressure at cgroup and pod level.
06 / CONTROLLED SCENARIOS
Contention changes latency before it changes the dashboard.
Three controlled scenarios show why utilization alone misses the story — and what offender-to-victim attribution adds.
Synthetic demo results · being re-validated on real clusters
SCENARIO 01
18.0 → 97.4ms
Serving p99 became 5.4× worse.
Linnix named the 8 offender PIDs.
SCENARIO 02
~3×
Serving p99 worsened.
Throughput halved and the serving cgroup stalled ~50%, while the batch hogs showed zero pressure.
SCENARIO 03
6.7×
Serving p50 worsened.
A tight CPU limit made latency worse while reported CPU usage fell.
On the offender attributions above: this is contention attribution, not proven causality.
07 / OPEN + CLOUD
The detector stays open. The fleet view becomes the product.
Run the complete local evidence layer yourself. Add Linnix Cloud when the problem becomes multi-cluster context, history, and organizational workflow.
Open
AGPL-3.0
Everything needed to detect, inspect, and export contention evidence on your own infrastructure.
Agent + eBPF collector
Local incident evidence
Prometheus metrics
CLI + MCP tools
Grafana dashboard import
We don't cripple the open detector.
Cloud
PAID
The fleet-level operating layer for teams running contention investigations across clusters and time.
Multi-cluster offender → victim graph
Long-term history
Cross-node baselines
Slack, PagerDuty and Jira
SSO, RBAC and org runbooks
08 / OPERATING MODEL
Built for the machines you cannot afford to disturb.
Observe first, fail honestly, and keep incident evidence inside the boundary your platform team already controls.
Monitor-first
Detects and reports. Enforcement is opt-in: it requires explicit configuration or human approval.
<1% eBPF · ~4% full daemon
<1% CPU for the eBPF probes on the kernel side; ~4% for the full userspace daemon.
Scoped on bare metal
Runs with CAP_BPF + CAP_PERFMON (CAP_SYS_ADMIN on older kernels).
The Kubernetes DaemonSet currently runs privileged, for simplicity.
Local by default
Analysis runs on the host. No data leaves your infrastructure.
Linux 5.12+ (arm64 5.18+).
START WITH ONE NODE
Find the process behind your next “slow.”
Deploy Linnix, send its metrics to the stack you already use, and inspect the offender → victim pair when PSI spikes.