Files
hermes-agent/tests/hermes_cli
kshitijk4poor ed8f8498d3 fix(gateway): a degraded boot stays degraded and still counts as a completed restart
Gate review on the stack:

- `hermes gateway restart` under systemd only accepted `gateway_state == "running"`
  as proof of the replacement; a boot with a parked platform stamps `degraded` for
  its whole life, so every restart/update on such a host waited out the 60 s+
  budget and reported a false failure while the gateway was serving. The verifier
  now accepts `degraded` as restarted and prints one DEGRADED warning.
- Drain release and scale-to-zero wake re-stamped `running` unconditionally, wiping
  the parked-platform signal after the first `.drain_request.json` cycle. Every
  "we are serving" stamp now goes through `_serving_state()`; the mixed fatal +
  retryable boot path sets the flag too (it fell through as a plain run before).
  `_startup_parked_platforms` is a bool — the joined error text was only ever logged.
- `_wait_for_tcp_port_free`: an unresolvable or unreachable configured host raised
  a non-refused OSError on every probe and burned the full 10 s wait; only a
  connect timeout means "listener alive", anything else means nothing to wait for.
- Windows `restart()` replaces its blind `time.sleep(1.0)` "let Windows release the
  port" with the same configured-address wait.
- The api_server bind retry rebuilds the AppRunner per EADDRINUSE attempt instead of
  calling aiohttp's private `_unreg_site`; attempt count is a named constant.

Test: a `degraded` replacement is reported as restarted (red on the previous head).
E2E re-run: predecessor releases the port 0.5 s after the first bind → bound on the
third attempt (+0.61 s), connect() True.
2026-09-19 02:06:55 +05:30
..