flock PR validation / validate (pull_request) Successful in 11s
Three defects enabled the 2026-08-16 Gitea blackhole (bug-wdgjpz3a00gd): 1. Orphaned allocation GC missing: ungraceful eviction (TaintManagerEviction) never calls CNI DEL, so the old node keeps advertising the pod's public /128 via BGP. Older allocation wins BGP path selection; live pod's node yields → blackhole. Fix: after the pod informer syncs at startup, sweep all committed allocations via orphanedCommitted(). Any allocation whose owner pod is absent from the node (or whose UID mismatches, indicating name reuse) is torn down, removed from the store, and released from IPAM. A 60 s periodic GC goroutine provides the same sweep while the agent runs. 2. renderBird outside-aggregate IP loop lacked pod liveness check: stale committed allocations caused BIRD to keep advertising the /128 even in steady state between GC ticks. Fix: before adding an outside-aggregate primary IP to the BIRD export, verify the pod is still in the node-scoped informer cache with a matching UID. Orphans are skipped silently; the GC cleans them on the next tick. 3. birdc startup race: the agent's first Render() fires before BIRD has bound /run/flock/bird.ctl, so the configure call silently fails with "Unable to connect" and the initial routes are never advertised. A container-only flock-agent restart (BIRD left running) avoids the race; a full pod restart re-hits it. Fix: reload() now retries up to 20 × 500 ms on socket-absent and "Unable to connect" conditions. Any other birdc failure (syntax error, etc.) is not retried. Fixes bug-wdgjpz3a00gd Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>