Repository navigation
Conversation
Signed-off-by: 1fanwang <1fannnw@gmail.com>
Signed-off-by: 1fanwang <1fannnw@gmail.com>
| @@ -522,6 +522,9 @@ func populateProcessFields(p *Process, proc procInfo) error { | |||
| } | |||
|
|
|||
| p.CPUTimeDelta = cpuTotalTime - p.CPUTotalTime | |||
There was a problem hiding this comment.
I think a comment would be useful here for future maintainers
There was a problem hiding this comment.
Have you seen this in the wild? I also think this should be used to invalidate the cache so the the proc info is re-read.
@sthaha, I have not seen a production occurrence. I reproduced PID reuse in a local Linux PID namespace. The fix rebuilds the process cache when cumulative CPU time drops, so Kepler reads the new process metadata.
Done in e8d6018.
sthaha
left a comment
There was a problem hiding this comment.
LGTM! Have you seen this in the wild?
I also think this should be used to invalidate the cache so the the proc info is re-read.
Signed-off-by: 1fanwang <1fannnw@gmail.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Equal CPU totals can still retain stale metadata after PID reuse.
Review effort: Balanced
Findings: 1
Open (1)
What changed in this PR
Adds PID-reuse handling to refresh stale process metadata and CPU accounting.
Changes:
- Rebuilds cached processes when cumulative CPU time decreases.
- Adds unit and refresh-level regression tests.
| File | Description |
|---|---|
internal/resource/informer.go |
Detects PID reuse and rebuilds process metadata. |
internal/resource/procfs_reader_test.go |
Tests metadata and CPU-delta reset behavior. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| return nil, err | ||
| } | ||
|
|
||
| if cpuTotalTime < cached.CPUTotalTime { |

Summary
When Linux reuses a PID, Kepler can keep the previous process's command and container assignment for the new process. If the new process starts with less cumulative CPU time, its metadata can also be skipped while the old cache entry is retained.
Kepler now replaces the process-cache entry when cumulative CPU time drops. The replacement reads current procfs data and starts its CPU delta from the new process's total.
Testing Done
I have no production sighting. I reproduced PID reuse with Kepler's procfs reader and real child processes in a Linux PID namespace. The new process reused the old PID with zero CPU time. Before cache invalidation, Kepler kept the old command name; after the change, it read the new command.
From the repository root, save the source below as internal/resource/pid_reuse_live_test.go and run:
Before cache invalidation:
After cache invalidation:
Reproducer source: pid_reuse_live_test.go