Operations and the kill switch
Do this: stop all dialing now
Section titled “Do this: stop all dialing now”- Open Overview or Operations. Both show the Dialing card.
- Press Pause all dialing. Write a reason (at least 8 characters) and type
PAUSE ALLto confirm. The box says: no campaign launches a new call; calls already up finish. - The card now shows who paused it, when and why. To lift it, press Resume dialing and confirm.
For one company only, use Pause dialing for one company on Operations: pick the company, write a reason and type PAUSE. Each paused company is listed with its own Resume. Only platform_owner sees these buttons.
Behind the screen: PUT /ops/dialing/global, PUT /ops/dialing/tenants/{id}, GET /ops/dialing.
What the kill switch does
Section titled “What the kill switch does”- No campaign launches a new call while it is on. The dialer checks every second.
- Calls already up carry on. Nothing is hung up.
- Campaign states do not change. A running campaign stays running and reports the idle reason
platform_paused. When you lift the pause, dialing resumes by itself and nobody has to restart anything. - A person cannot dial around it. Manual dialing answers
platform_pausedand asking for the next person in a preview campaign answers 503. - It fails closed. If the dialer cannot read the switch, it launches nothing (
platform_pause_unreadable) until it can. - Pausing an already paused target changes nothing and records nothing new. Each change emits
vd.platform.dialing_pause_changed.
Watching the platform
Section titled “Watching the platform”Operations opens with four numbers (services answering, calls on the phone, dial rate, bot reply time p95), the kill switch card and six tabs. Live numbers refresh every 10 seconds while the page is visible. Every verdict is an icon plus a word, never colour alone.
| Tab | What it shows | Behind it |
|---|---|---|
| Services | Each service: health, version, answer time, uptime, server errors. | GET /ops/services |
| Queues | Webhook deliveries, service outboxes and the lag of every durable message consumer, each marked by how far behind it is. | GET /ops/queues |
| Live calls | Calls and AI sessions now, by company, and the dial rate overall and for the busiest companies. Denied attempts count. | GET /ops/live, GET /ops/dial-rate |
| Media nodes | Capacity and live calls per node, with Drain and Resume. | GET /telephony/nodes |
| Recordings | Companies with recording gaps in the last 48 hours. | GET /telephony/recordings/reconciliation |
| AI | Whether the AI upstream is configured and reachable, reply time p50 and p95 over 15 minutes, sessions live and failed. | GET /ai/health |
The models, base URLs and whether a key is present (yes or no, never the key) are on Settings, tab AI provider.
Maintenance mode
Section titled “Maintenance mode”In Settings, tab Maintenance, press Turn maintenance on, write the message and type MAINTENANCE. This turns on the global pause with the reason maintenance: <message> and shows the message to every company’s web app (PUT /settings/maintenance). Turn maintenance off lifts it. Turning it off lifts only a pause it set itself. A pause you set by hand stays.
Media nodes
Section titled “Media nodes”A media node is a server that carries the calls. On the Media nodes tab, Drain (type DRAIN to confirm) makes a node take no new calls while live ones end, and Resume puts it back (POST /telephony/nodes/{id}/drain, .../resume). Both need platform_owner. A node that is draining because its own process is shutting down is never resumed by this call.
Recording reconciliation
Section titled “Recording reconciliation”The Recordings tab (GET /telephony/recordings/reconciliation?hours=48, 1 to 720 hours) lists answered calls with a recording gap: no recording 15 minutes after the call ended, a failed recording, or one still waiting to upload 15 minutes after it was written. Only companies with a gap are listed, with up to five example call ids each.