Logo GH

Interaction of teams in operations

1) Why

The iGaming platform is dozens of domains (Payments, Games/Core, Risk/KYC, Data, Infra/SRE, Support, Compliance). Without formalized interoperability, MTTR, CFR and operational risks rise. The goal is to turn disparate functions into a single operating system: predictable contacts, transparent queues, common signals and consistent priority.

2) Principles

1. SLO-first: joint solutions are tied to SLO/error budgets.
2. A single source of truth: common dashboards, uniform statuses and artifacts.
3. Clear boundaries and interfaces: each pair of commands has a described contract (OLA/Runbook/API).
4. Small batches and reversibility: changes through phicheflags/canary, fast rollback.
5. No blame - yes data: parsing on facts, improvements - a mandatory part of the cycle.
6. Minimum required privileges and SoD: role separation for sensitive operations.
7. Automate the routine, standardize the rest.

3) Roles and RACI (end-to-end)

Head of Ops/SRE Lead is the owner of the operational framework, KPI/KRI. A

Service Owners (Payments/Games/KYC/Data) - domain goals, changes, risk. A/R

Platform/Infra - accessibility, performance, releases/canaries. R

Risk/Compliance/Security - SoD, RG/KYC/PII, audits. C/A

Support/CRM - front of complaints, communication to players. R/C

On-call IC/CL - incident management and external updates. R

Release Manager - calendar, CAB, status of changes. R

Data/Analytics - product and operational metrics, RCA support. R/C

4) Interaction Contracts (OLA/SLx)

OLA (Operational Level Agreement) - internal agreements between teams (not external SLAs). Include:
  • Areas of responsibility: whose area (for example, PSP routing - Payments; cache/DB - Infra).
  • Objectives/threshold metrics: Incident MTTA, escalation response time, post-monitoring window.
  • Queues and priorities: P1-P4, business criticality, freeze windows.
  • Interfaces: channels, bot commands, API/Runbook, owner directories.
  • Artifacts: what documents/logs/dashboards are required to accompany the event.
💡 The recommended OLA set: Payments↔Infra, Games↔Infra, Payments↔Risk/KYC, Risk↔Compliance, Ops↔Support, Ops↔Release.

5) Communication channels and protocols

Operational chat (shift): daily updates, mini-rituals, handover.
Var rooms for incidents: created by a bot; IC/CL roles are assigned by the command.
CAB/Change channel: discussion of changes, risks, release calendar.

Read-only SLO/Incident/Planned Activity summaries

Escalation: command templates '/page ', '/escalate', SLA reporting.

Unified message protocol: "fact → impact → ETA/ETR → next update window → owner."

6) Handovers between shifts and regions

Template 10-15 minutes:

1. SLO/SLI: where is the risk of budget burnout.

2. Open incidents/escalations and their ETAs.

3. Planned works/releases in the next 24-48 hours.

4. Providers (PSP/KYC/studios): active tickets, expectations.

5. On-call composition and contacts (IC/CL/domains).

6. "Watchlist" - areas of increased attention (queues/replication/cache).

Handover is recorded in a shift log, links - to var-rooms and dashboards.

7) Incident collaboration

Start: alert → bot creates a card '# inc-YYYY-MM-DD-XXX', assigns IC/CL and domain leads.
One vote rule: IC is the final decision; CL - communications.
Facts and hypotheses: we separate; "red" signals - priority.
Guardrails: phicheflags/PSP routing only change via runbook with SoD/dual-control.
Communications: drafts of public updates via CL, partners - targeted.
Closure: post-monitoring, post-mortem generation and improvement tasks with owners/deadlines.

8) Collaborate on changes

Release calendar: public, with freeze periods and on-call slots.
Quality gates: unit/contract/e2e, security, SLO gates staging.
Canary rolling: step by step 5%→25%→100% for GEO/tenants/banks.
Auto-rollback: policies by key SLI/KRI, WORM magazine.
Comm packages: draft updates agreed in advance with CL/Legal.
RACI changes: RM (A/R), SO (A/R), SRE (R), Sec/Compliance (C/A), CAB (A), IC/CL (R/C).

9) Unified telemetry and artifacts

Common metrics directory: SLI/SLO, business metrics, KRI (queues, PSP, replication).
Dashboard "Operation Map": summary by domains, regions, status of incidents/works.
Timelines: uniform format (time, author, action, result, links).
Post-mortems: template without charges, prevention measures, revision date.
Runbooks/Checklists: versioned; link from alerts and incident cards.

10) Prioritization and planning

Weekly Ops plan (30-45 min): coordination of top risks, releases, limits, improvements from post-mortems.
Kanban of operations: columns' Backlog → Ready → In Progress → Validate → Done ', WIP limits.
Priority criteria: impact on SLO/revenue/compliance, size/reversibility, dependence on providers.

11) Escalation matrix (squeeze)

EventTo whomSLA reactionsGot feedback
P1 payments (auth-success drop)IC + Payments + Infra≤ 5 minVar room, guardrails, canary pullback
P2 Settle DelaysGames/Core + Infra≤ 15 minIncrease in workers/quota, monitoring
PSP partner not availablePayments + Support≤ 15 minPartner Comm/Status, Temporary Routing
PII Leak/SuspicionSec/Compliance + IC/CLimmediatelyExport freeze, legal procedure
Release canary degradesRM + SRE + SO≤ 5 minAuto-rollback, comm inside, post-analysis

12) Policies and SoDs

SoD/4-eyes: conclusions/bonuses/PSP routing/PII export - only with double approval.
JIT rights: temporary escalation of privileges for runbook actions.
Data policies: PII ban in open channels/dashboards; geo-boundaries.
Audit - immutable activity logs (WORMs), policy revisions.

13) Interaction tools

Incident bot: '/incident new ', roles, update timers, comm drafts, '/runbook', '/flag ', '/config'.
Metrics API: common SLO view and KRI, exemplars (trace_id) for RCA.
Release-portal: manifests, gates, rolling/rollback status.
Owners directory/CMDB: domains, contacts, backup channels.

14) Collaboration Metrics (KPI/KRI)

MTTA/MTTR by domain and slot (day/night), proportion of incidents caught before complaints.
Handover Quality: transmission defects (checklist items not closed on time).
Change Collaboration:% of releases with ready-made comm packages and no rollbacks.
Guardrail Discipline: frequency of SoD/policy violations (target 0).
Comms Cadence: adherence to public update intervals when P1/P2.
Post-mortem SLA: proportion of post-mortems ≤ D + 5, completion of actions.
Fair-share Load: Distribution of nights/peaks by people/teams.
Customer Signal Lead: Lag between objective degradation and first complaints.

15) Implementation Roadmap (6-10 weeks)

Ned. 1-2: domain/owner inventory; OLA templates; launch of the replacement channel and handover checklist; base escalation matrix.
Ned. 3-4: incident-bot (MVP), shared status channel, single SLO/SLI/KRI card; runbooks directory.
Ned. 5-6: CAB/release calendar, comm packages and freeze windows; SoD/4-eyes for sensitive operations.
Ned. 7-8: canary rolling and auto rollback as standard; post-mortem template, Exec/Ops-dashboards collaboration.
Ned. 9-10: P1 exercises, cross-regional handovers, WORM audits, KPI/KRI reports, OLA adjustments.

16) Templates (fragments)

16. 1 OLA (Payments ↔ Infra/SRE)

yaml ola:
scope: "Payments-Auth & Routing"
contacts:
payments_so: "@pay-so"
infra_oncall: "@sre-oncall"
objectives:
mtta_p1: "≤5m"
rollback_ttr: "≤10m canary"
interfaces:
runbooks: ["psp-failover", "reroute", "auth-throttle"]
dashboards: ["auth_success", "psp_latency", "queue_lag"]
escalation:
p1: ["IC","Payments Lead","SRE L2"]
p2: ["Payments OnCall","SRE OnCall"]
artifacts:
status_templates: ["public","partners"]
postmortem_due: "D+5"

16. 2 Handover checklist (10 items)

1. SLO domain statuses

2. Open Incidents (ETA/Owners)

3. Planned activities/releases + observation windows

4. Providers (PSP/KYC/Studios) - Risks/Expectations

5. Queues/replication/cache - lag/anomalies

6. Limit/Phicheflag Changes

7. Complaints/tickets and load thresholds

8. Comma plans and status drafts

9. On-call composition and reserve

10. "Watchlist" per slot

17) Antipatterns

"Will anyone deal?" without RACI and owner.
Incidents without IC/CL and update timers.
Hidden changes (manual clicks), no Git/Audit.
Non-common telemetry: different numbers in different teams.
Releases without comm packs and canaries.

SoD violations "for the sake of speed."

Handovers orally, without records and checklists.
Post-mortems without actions and deadlines.

Total

The interaction of teams in operations is a contractual collaboration: OLA/SLx, clear channels and roles, handover discipline, general telemetry, coordinated releases and incident processes. Such a framework reduces MTTR and CFR, aligns priorities, protects SLO, revenue and compliance - and makes daily operations predictable and sustainable.

Contact

Get in Touch

Reach out with any questions or support needs.We are always ready to help!

Start Integration

Email is required. Telegram or WhatsApp — optional.

Your Name optional
Email optional
Subject optional
Message optional
Telegram optional
@
If you include Telegram — we will reply there as well, in addition to Email.
WhatsApp optional
Format: +country code and number (e.g., +380XXXXXXXXX).

By clicking this button, you agree to data processing.