ROLE · R/11 · OPERATIONS TRIBE

A digital ops manager so your team stops fighting fires.

Connect your stack (AWS / GCP / Vercel / Linear / PagerDuty / Statuspage / Bill.com / Ramp / Slack / Zendesk). Cyborg writes and maintains every SOP, monitors 24/7 vendor SLAs, runs procure-to-pay end-to-end (PR → PO → 3-way match → payment draft), tracks uptime across 47 systems, triages every ops ticket within 4 minutes, and runs the weekly ops review. Vendor renewals always founder-approved. Wires always founder-approved. Production deploys always engineer-approved.

FROM $599/month · SCALES TO $1,799 (head of ops) · ONBOARDING 7 days
— 02 / WHAT IT DOES

— CAPABILITIES

From SOP authoring to vendor renewal — on autopilot.

Routine ops autonomous. Vendor renewals + procurement + production access always human-gated. Your team gets calmer, your runway gets longer, your audit gets cleaner.

SOP authoring + maintenance

T1 · AUTONOMOUS

Writes new SOPs from observed workflow + interview transcripts. Versions every change. Quarterly review reminders to owner. RACI matrix per step. Linked to the Slack channels + tools they touch.

Vendor SLA monitoring

T1 · AUTONOMOUS

24/7 ping on every critical vendor (cloud, payments, comms, CRM, identity). SLA breach >5 min: PagerDuty page. Monthly SLA report with credits owed back to you under each contract.

Procure-to-pay (PR → PO)

T2 · GUIDED

Drafts purchase requests from Slack request, routes to budget owner, generates PO once approved, 3-way matches PO ↔ receipt ↔ invoice, schedules payment. Founder approves payment runs.

System uptime tracking

T1 · AUTONOMOUS

Composite uptime view across 47 systems (your APIs, vendor APIs, internal tools, status pages). Builds your status.yourco.com. RCA template auto-drafted on incident close.

Operational ticket triage

T1 · AUTONOMOUS

Every #ops-help request classified, prioritised, assigned in <4 min. P0 → page on-call. P1 → ticket + owner. P2/P3 → queue + ETA. Linked to relevant SOP.

Vendor renewal management

T3 · APPROVED

90 / 60 / 30 day renewal alerts. Drafts negotiation memo (usage data + market benchmark + alternatives). Founder + legal approve every renewal >$5k ARR. Zero auto-renewals.

Office / facilities ops

T2 · GUIDED

Office supplies reorder, cleaning vendor SLA, IT equipment dispatch (laptop unboxing → user → MDM enrolment), visitor management. Cyborg drafts, office manager approves.

Onboarding / offboarding ops

T2 · GUIDED

New hire: laptop ordered, 47 SaaS accounts provisioned, Slack channels added, SOPs assigned, day-1 buddy paired. Offboarding: 47 accounts revoked in <15 min, exit checklist closed.

Weekly ops review

T1 · AUTONOMOUS

Every Friday 16:00: 7-page deck. SLA summary, ticket queue, vendor spend, upcoming renewals, incident log, top 3 risks. Posted to #ops-weekly + emailed to leads.

Incident commander (P1+)

T2 · GUIDED

P1 incident → opens war-room channel, pages on-call, drafts customer Statuspage update (engineer approves before publish), keeps timeline, drafts post-mortem within 24h.

Compliance evidence collection

T3 · APPROVED

SOC 2 / ISO 27001 / HIPAA controls evidence auto-collected (access reviews, change logs, vendor reviews). Quarterly access review drafts; ops lead + security approve before submit.

Production deploys + access

T4 · NEVER ALONE

Cyborg drafts deploy notes + checks pre-flight (tests green, on-call available, change ticket open). It never clicks deploy. Prod database access requires engineer + founder dual sign.

TRUST LEVELS · T1 Cyborg posts, you can audit later T2 Cyborg drafts, human reviews T3 Human approves before action runs T4 Cyborg never signs alone
— 03 / STACK

— FLUENCY

Linear, PagerDuty, AWS — all native.

Read-only on production by default. Write-access only on operational tools (tickets, SOPs, status pages). Procurement is always human-approved at the payment step.

Issue tracking + on-call — READ-WRITE

  • Linear★★★★★
  • Jira / Atlassian★★★★★
  • PagerDuty★★★★★
  • Opsgenie★★★★★
  • incident.io★★★★★
  • Statuspage / BetterStack★★★★☆

Cloud + infra — READ ONLY

  • AWS (CloudWatch + Cost Explorer)★★★★★
  • GCP (Monitoring + Billing)★★★★★
  • Vercel / Netlify★★★★★
  • Kubernetes / EKS / GKE★★★★☆
  • Cloudflare★★★★★
  • Datadog / Grafana★★★★★

Procurement + finance ops — DRAFT

  • Bill.com★★★★★
  • Ramp★★★★★
  • Brex★★★★★
  • Coupa / SAP Ariba★★★★☆
  • Vendr / Tropic (SaaS buying)★★★★☆
  • Notion (vendor registry)★★★★★

Identity + comms + HR ops — READ-WRITE

  • Okta / Google Workspace★★★★★
  • Slack (50+ channels orchestrated)★★★★★
  • MS Teams / Entra ID★★★★★
  • Zendesk / Intercom (ops queue)★★★★★
  • Rippling / Gusto (HR ops)★★★★★
  • JumpCloud / 1Password (MDM + secrets)★★★★☆
— 04 / DAY IN LIFE

— THURSDAY · 06:00 → 18:00 IST · CALM WEEK, 1 P2 INCIDENT

A real Thursday. Series-B SaaS, 84 employees, 247 SOPs.

No fires today. 1 P2 (Stripe webhook latency) auto-triaged. 3 vendor renewals queued. Weekly ops deck shipped Friday-prep at 17:30.

  1. 06:00

    Overnight SLA scan

    47 systems polled at 03:00 UTC. 46 green. 1 yellow: Stripe webhook p95 latency 1,840ms vs SLO 1,200ms (last 6 hours). Auto-classified P2, ticket OPS-2026-7142 created, on-call notified for morning review — not paged.

  2. 07:30

    New-hire onboarding kicked off

    2 new engineers start Monday. Cyborg orders 2 MacBook Pros (Ramp), provisions 47 SaaS accounts (Okta + scim), drafts day-1 schedule, assigns onboarding buddy, sends welcome packet. Office manager approves the laptop spec — everything else autonomous.

  3. 09:15

    Vendor renewal queue refresh

    5 contracts due in next 30 days, $148.4k ARR. Top one: AWS Reserved Instance commit, $96k for 12 months. Cyborg pulls last-12-months usage, market benchmark (Vantage data), alternatives (GCP CUDs), drafts negotiation memo to founder + legal. Two more drafted, two awaiting usage data.

  4. 10:42

    Procure-to-pay run

    3 PRs in queue: $1,240 office furniture, $4,200 Datadog seat upgrade, $890 conference tickets. Cyborg routes each to budget owner. 2 approved within 38 min. 1 (Datadog) escalated — engineering lead asked for usage justification. Cyborg pulled the data, lead approved 14 min later.

  5. 11:48

    SOP update from incident

    Last week's P1 Stripe webhook RCA closed yesterday. Cyborg drafts SOP-2026-04-WHF-0042 v2.4: adds explicit retry policy + dead-letter queue check + new alarm threshold. Engineer reviews, approves with 1 inline edit. Published.

  6. 13:00

    Quarterly access review (in progress)

    SOC 2 control AC-2: quarterly review of all production access. Cyborg pulled Okta + AWS IAM + GitHub admin lists, cross-referenced active employees, flagged 4 stale accounts (3 ex-employees, 1 contractor whose project ended). Drafts revoke ticket + evidence package for security lead.

  7. 14:30

    Status page sync

    status.yourco.com auto-updated with the morning's Stripe webhook latency. Customer-facing message: "Investigating elevated webhook latency. No data loss. Updates in 30 min." Approved by on-call before publish (T2 gate on customer comms).

  8. 15:45

    Office supplies + facilities

    Coffee + snacks reorder triggered (par level hit). Office cleaning vendor invoice arrives, 3-way matched against contract + last visit log. Visitor pass for 16:30 partner meeting issued. Office manager copied on each.

  9. 16:30

    Friday ops deck draft

    Weekly ops review prep. 7-page deck: SLA summary (47/47 green this morning), ticket queue (12 open, 4 P2), vendor spend ($43.2k this week), 5 upcoming renewals, 1 P2 incident, top 3 risks (RI commit decision, hiring backfill, year-end audit prep).

  10. 17:00

    Stripe webhook P2 closed

    Engineering shipped the fix at 16:42. Cyborg verifies p95 back to 412ms (well under SLO). Closes OPS-2026-7142. Drafts mini-RCA. Adds to Friday deck.

  11. 18:00

    Tomorrow's queue ready

    Tomorrow: Friday ops review @ 16:00, AWS RI decision needed, 2 new-hire MDM enrolments (laptops arriving), 1 SOC 2 evidence push. Cyborg has it all queued. EOD post in #ops: "All systems green. 1 P2 closed. 5 renewals, 1 awaiting your call. Friday deck draft attached."

— 05 / SAMPLE DELIVERY

— REAL SOP DOC + OPERATIONS DASHBOARD · ANONYMISED

A real SOP. A real ops dashboard.

Two artefacts from a real customer's ops stack: the "Stripe Webhook Failure Recovery" SOP (v2.4, owned by SRE lead) and the live operations dashboard with vendor SLAs + ticket queue + uptime grid.

SOP-2026-04-WHF-0042 · ● PUBLISHED · v2.4 · LAST REVIEWED 2026-04-22

Stripe Webhook Failure — Detection & Recovery
Triggered when Stripe webhook delivery p95 latency exceeds 1,200 ms for 3 consecutive 5-min windows, OR any webhook returns non-2xx for >15 events in 10 min.

  • OWNERSRE Lead · @priya.r
  • REVIEWERSOn-call rotation, Eng Mgr
  • CADENCEQuarterly review
  • LINKED INCIDENTS4 (last: INC-2026-04-1207)
— RACI MATRIX
ActivitySRE on-callEng LeadFounderCyborg-Ops
Detect breachIIR · A
Triage + classifyIR · A
Investigate root causeR · ACS (data pull)
Customer comms (Statuspage)ACIR (drafts)
Implement fixRA
Verify recoveryRAIS (auto-verify)
Post-mortem + SOP updateCAIR (drafts)
R = Responsible · A = Accountable · C = Consulted · I = Informed · S = Support
  1. 01

    Confirm the breach is real (not a monitoring blip)

    Open the Datadog Stripe Webhook dashboard. Verify 3+ consecutive 5-min windows above 1,200 ms p95. Cross-check with Stripe Dashboard → Webhooks → Delivery attempts. If only 1 window red, mark false-positive in PagerDuty and exit.

    OWNER · SRE on-call · SLA 4 min
  2. 02

    Open war-room channel, page on-call backup if >P2

    Cyborg auto-opens #inc-YYYY-MM-DD-stripe-webhook, invites SRE on-call + Eng Lead. If failure rate >25% of webhooks (P1), Cyborg pages on-call backup via PagerDuty. Founder informed only if customer-facing.

    OWNER · Cyborg-Ops (autonomous)
  3. 03

    Pull diagnostic bundle

    Cyborg auto-attaches to war-room: last 100 webhook attempts (success / fail / latency), DB connection pool metrics, deploy log (last 24h), recent infra changes (Terraform), Stripe API status page. Reduces MTTR by ~12 min historically.

    OWNER · Cyborg-Ops (autonomous)
  4. 04

    Customer comms (if customer-facing failure)

    Cyborg drafts Statuspage update using template SP-T-007. SRE on-call reviews + approves before publish (T2 gate on all customer comms). First update within 12 min of breach. Subsequent updates every 30 min until "Resolved".

    OWNER · SRE on-call · Cyborg drafts
  5. 05

    Investigate root cause + ship fix

    SRE follows the runbook decision tree (linked: webhook-rca-tree.md). Common causes: deploy regression (rollback), DB pool exhaustion (scale up), Stripe API degradation (wait + replay). Fix shipped, Cyborg verifies recovery via SLO dashboard.

    OWNER · SRE on-call + Eng Lead
  6. 06

    Replay missed webhooks + reconcile orders

    If >0 webhooks failed: Cyborg drafts a replay job (Stripe webhook replay endpoint) + cross-checks order DB for any orders awaiting webhook. Engineer approves replay batch before run (T3, since this writes to prod).

    OWNER · Eng Lead approval · Cyborg drafts
  7. 07

    Post-mortem within 24h, SOP update if needed

    Cyborg drafts blameless post-mortem in post-mortems/ using template PM-T-003. Eng Lead reviews + approves. If SOP gap identified, Cyborg drafts SOP version bump (this doc went from v2.3 → v2.4 after INC-2026-04-1207).

    OWNER · Eng Lead · Cyborg drafts
VERSION HISTORY · v2.4 (2026-04-22, added DLQ check) ← v2.3 (2026-02-08) ← v2.2 (2025-11-14) ← v2.1 ← v2.0 ← v1.x NEXT REVIEW · 2026-07-22 · Open in Notion →
OPERATIONS LIVE DASHBOARD · refreshed 17:42 IST · OPS-DASH-W19 Status: ● ALL SYSTEMS GREEN · 1 P2 OPEN
UPTIME (30D)99.94%+0.02 vs LM
MTTR (P1, 30D)14 min↓ 6 min vs LM
VENDOR SLA CREDITS$2,840claimed YTD
RENEWALS DUE 30D5$148.4k ARR

— VENDOR SLA STATUS · LIVE

  • ● 99.99AWS (us-east-1)SLO 99.95 · healthy
  • ● 100.00CloudflareSLO 99.99 · healthy
  • ● 99.78Stripe webhooksSLO 99.90 · OPS-7142 closed 17:00
  • ● 99.97Postgres (RDS)SLO 99.95 · healthy
  • ● 100.00Okta SSOSLO 99.99 · healthy
  • ● 99.96Twilio (SMS)SLO 99.90 · healthy
  • ● 99.99SendGridSLO 99.95 · healthy
  • ● 100.00Vercel (web)SLO 99.99 · healthy

— TICKET QUEUE · OPEN 12 / DONE 38 (7D)

  • P2OPS-7148Datadog seat upgrade approval@nikhil · 2h
  • P2OPS-7144AWS RI commit decision (renewal)@founder · 6h
  • P2OPS-7140Stale Okta accounts (4) revoke@security · 1d
  • P3OPS-7138New-hire MacBook order × 2@office · 1d
  • P3OPS-7135Office coffee reorder@office · 2d
  • P3OPS-7129SOC 2 evidence — Q2 access review@security · 3d
  • P4OPS-7124Update vendor registry (Notion)@cyborg · 4d
  • P4OPS-7118Visitor pass — partner mtg Thu@office · 4d

— UPTIME GRID · LAST 30 DAYS · 16 SYSTEMS · GREEN = 100% / YELLOW = <SLO / RED = INCIDENT

AWS
Stripe
Cloudflare
RDS Postgres
Okta SSO
Composite uptime 99.94% · 2 yellow days (Stripe webhook latency) · 1 red day (Stripe API regional degradation, no customer impact, 4 min) Open full status →
— 06 / TOOLS IT TOUCHES

— INTEGRATIONS

Read prod, write the queue, never deploy alone.

Production access is read-only. Operational tools (tickets, SOPs, status pages) are read-write. Procurement is always human-approved at the payment step. Deploys never autonomous.

— 07 / PRICING

— FLAT MONTHLY · NO PER-TICKET FEE

3 tiers. No "per-vendor" fee.

Unlimited SOPs, unlimited vendors monitored. Cheaper than a junior ops analyst, more reliable than a 3-person ops team.

OPS-JR · OPS ANALYST

$599/month

SOP authoring, ticket triage, vendor SLA monitoring, 1 ticketing system. Up to 30 employees. Best for seed / pre-Series-A.

  • Up to 50 SOPs maintained
  • Up to 20 vendors monitored
  • 1 ticketing system + 1 IDP
  • Daily ticket triage
  • Weekly ops digest
  • Office + facilities ops
Start ops analyst
OPS-SR · HEAD OF OPS

$1,799/month

Multi-entity ops. Vendor portfolio >$2M ARR managed. Audit-ready evidence packages. Up to 500 employees.

  • Multi-entity ops (UK + US + IN)
  • Vendor portfolio review (quarterly)
  • Strategic procurement (RFP drafts)
  • Annual SOC 2 / ISO audit prep
  • BCP / DR plan maintenance
  • Vendor risk assessments
  • Dedicated #ops-lead channel
Start head of ops

14-DAY EVALUATION · 30-DAY MONEY-BACK · MONTH-TO-MONTH · UNLIMITED SOPS · See full pricing →

— 08 / ONBOARDING

— DAY 1 → DAY 7

Sign Monday, first SOP shipped Friday.

  1. DAY 0

    You sign

    Contract + DPA + region pick (EU / US). You add Cyborg as user (not admin) on Linear/Jira, read-only on AWS/GCP, scim-write on Okta, draft-only on Bill.com / Ramp.

  2. DAY 1

    Kickoff (90 min)

    Founder + ops lead + on-call lead + Cyborg. Stack walkthrough, SLO + SLA matrix, vendor list, escalation tiers, who-approves-what, current pain points (top 3).

  3. DAY 2-3

    SOP discovery

    Cyborg reads last 90 days of #ops + #help + ticket queue. Identifies the top 20 most-repeated questions. Drafts SOP for each. Owner-pairs each with a human reviewer.

  4. DAY 4

    Shadow mode

    For 24-48 hours, every Cyborg action drafted into #ops-drafts. Ops lead approves with one tap. Zero auto-actions. Tunes confidence thresholds.

  5. DAY 5

    First autonomous triage

    Ticket triage goes live. Cyborg classifies + assigns + applies SOP link in <4 min. Ops lead audits the next morning. If clean for 3 days, vendor SLA monitoring goes live next.

  6. DAY 6-7

    Full live mode

    SOPs maintained, SLAs monitored, ticket queue triaged, weekly review deck shipped. T3 procurement + T4 deploys still gated. First Friday ops review together with founder + ops lead.

— 09 / FAQ

— OPS LEAD + FOUNDER QUESTIONS

Sawaal jo har Head of Ops poochta hai.

Can it deploy to production or change AWS configs?

No. Production deploys are T4-protected: an engineer must click. Cloud writes (AWS / GCP IAM, security groups, RDS configs) are also T4 — Cyborg can read everything (CloudWatch, Cost Explorer, IAM lists for access reviews) but cannot write. The only autonomous writes are: tickets, SOPs, draft Statuspage updates (gated before publish), Slack messages, and SCIM provisioning on Okta (gated by joiner/leaver tickets).

Will it auto-renew vendors? I've been burned by surprise renewals.

Never. Every renewal >$5k ARR is T3 — founder + legal approve before the contract is signed. Cyborg does the heavy lifting: 90-day alert, usage data pull, market benchmark, alternatives memo, draft negotiation email. You decide. We've engineered it so the Cyborg literally cannot click "renew" on Vendr / direct vendor portals.

Will my ops lead lose their job?

Honestly, no — the role shifts. The 60% of an ops lead's day spent on tickets, SOP maintenance, vendor follow-ups, and chasing approvals goes away. What's left is the strategic 40%: org design, vendor portfolio strategy, BCP / DR planning, big procurement decisions. Most of our customers keep their ops lead + drop the 1-2 ops analyst hires they were planning.

What if it triages a ticket wrong (P3 instead of P1)?

Confidence-gated. A ticket with the words "down", "outage", "users affected", "production", combined with a known severity-keyword pattern, auto-classifies P0/P1 with >99% accuracy. Edge cases are queued for ops lead review (5-10 per week typically). Across our customer base: 0.4% mis-classification rate, 100% caught within 1 hour at the next ops standup. P0/P1 always pages on-call regardless — Cyborg can't delay an escalation.

Can it actually run a P1 incident as commander?

Yes — for the process, not the engineering decision. Cyborg opens the war-room channel, pages on-call, attaches diagnostic bundles, drafts customer comms (engineer approves before publish), keeps the timeline, drafts the post-mortem. The actual technical decision (rollback? scale up? wait?) is your engineer's call. Median MTTR for our customers dropped 38% — not because Cyborg is smarter than your engineers, but because they spend 100% of incident time fixing instead of process-managing.

SOC 2 evidence collection — can my auditor trust the trail?

Yes. Every evidence artefact is timestamped, source-linked, and signed by both Cyborg + the human approver. Quarterly access review for AC-2: Cyborg pulls Okta + AWS IAM + GitHub admin lists, cross-references HR roster, drafts the review packet. Security lead signs. Evidence package goes straight to your auditor's portal (Drata / Vanta / Secureframe). 3 of our customers passed SOC 2 Type II audits with zero exceptions on Cyborg-collected evidence.

What happens to the SOPs if I cancel?

You keep them. SOPs live in your Notion / Confluence — we never own the registry. On cancellation, Cyborg's user is deactivated, every SOP stays exactly as it is, you keep the version history + RACI + linked incidents. Optional 30-day grace period to swap to a human ops manager. Your processes, always.

— SEE OTHER ROLES

Operations is one of 12.

Pair with the Customer Support Cyborg for ticket overflow, the Accountant Cyborg for procure-to-pay GL sync, the HR Cyborg for joiner/leaver workflows, and the Legal Assistant Cyborg for vendor contract redlines.

Browse all 12 roles