get_model_info called get_model_context_length inline. The resolver chain runs several sequential provider probes, each with its own multi-second timeout, so an unreachable model.base_url held the response for tens of seconds. Bound the whole chain to a 5s budget on a throwaway worker thread; on timeout the response degrades to auto_context_length=0 while the abandoned probe finishes on its own. Fixes https://github.com/NousResearch/hermes-agent/issues/63214 (backend half)