flock-agent: GC orphaned allocations; retry birdc on socket-not-ready
flock PR validation / validate (pull_request) Successful in 11s

Three defects enabled the 2026-08-16 Gitea blackhole (bug-wdgjpz3a00gd):

1. Orphaned allocation GC missing: ungraceful eviction (TaintManagerEviction)
   never calls CNI DEL, so the old node keeps advertising the pod's public
   /128 via BGP. Older allocation wins BGP path selection; live pod's node
   yields → blackhole.

   Fix: after the pod informer syncs at startup, sweep all committed
   allocations via orphanedCommitted(). Any allocation whose owner pod is
   absent from the node (or whose UID mismatches, indicating name reuse) is
   torn down, removed from the store, and released from IPAM. A 60 s
   periodic GC goroutine provides the same sweep while the agent runs.

2. renderBird outside-aggregate IP loop lacked pod liveness check: stale
   committed allocations caused BIRD to keep advertising the /128 even in
   steady state between GC ticks.

   Fix: before adding an outside-aggregate primary IP to the BIRD export,
   verify the pod is still in the node-scoped informer cache with a matching
   UID. Orphans are skipped silently; the GC cleans them on the next tick.

3. birdc startup race: the agent's first Render() fires before BIRD has
   bound /run/flock/bird.ctl, so the configure call silently fails with
   "Unable to connect" and the initial routes are never advertised. A
   container-only flock-agent restart (BIRD left running) avoids the race;
   a full pod restart re-hits it.

   Fix: reload() now retries up to 20 × 500 ms on socket-absent and
   "Unable to connect" conditions. Any other birdc failure (syntax error,
   etc.) is not retried.

Fixes bug-wdgjpz3a00gd

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
ops
2026-08-16 06:26:38 +00:00
co-authored by Claude Sonnet 4.6
parent 197bc6f3b8
commit c9950348ff
5 changed files with 247 additions and 8 deletions
+47
View File
@@ -82,6 +82,34 @@ func (s *Server) configureRuntime(ctx context.Context) error {
return fmt.Errorf("pod informer: %w", err)
}
// Startup orphan GC: the pod informer is now fully synced. Walk all
// committed allocations and release any whose owner pod is absent from
// this node. This catches ungraceful evictions where CNI DEL never ran
// (TaintManagerEviction path) and prevents stale public /128s from
// suppressing the live pod's BGP advertisement after rescheduling.
gcOrphans := func(label string) int {
orphans := orphanedCommitted(s.Store.Snapshot(), func(ns, name string) (string, bool) {
pod, ok := pods.Get(ns, name)
if !ok {
return "", false
}
return string(pod.UID), true
})
for _, a := range orphans {
s.Logger.Info(label,
"container_id", a.ContainerID,
"pod", a.Namespace+"/"+a.PodName,
"ip6", a.IP6,
"ip4", a.IP4,
)
_ = Teardown(a.ContainerID, net.ParseIP(a.IP6), net.ParseIP(a.IP4))
_ = s.Store.Delete(a.ContainerID)
ipam.Release(net.ParseIP(a.IP6), net.ParseIP(a.IP4))
}
return len(orphans)
}
gcOrphans("GC orphaned committed allocation (startup)")
// Keep NetworkUnavailable=False so the node.kubernetes.io/network-
// unavailable taint never gets re-applied. Calico's calico-node sets
// it on shutdown; without an owner replacing it, kubelet's controller
@@ -132,6 +160,25 @@ func (s *Server) configureRuntime(ctx context.Context) error {
}
}()
// Periodic orphan GC: defense-in-depth against allocations that escape
// the startup sweep (e.g. a pod evicted while the agent is running and
// the CNI DEL is never delivered). Keeps the store and IPAM in sync
// with the live pod set without requiring a full agent restart.
go func() {
t := time.NewTicker(60 * time.Second)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
if n := gcOrphans("GC orphaned committed allocation (periodic)"); n > 0 {
anycast.Trigger()
}
}
}
}()
// NetworkPolicy enforcement.
world := netpol.NewWorld(s.Logger)
if err := world.Start(ctx, s.restCfg); err != nil {