Перейти к содержимому

Production support process

Это содержимое пока не доступно на вашем языке.

Defines how Chronacta production incidents are triaged and escalated.

Tier Scope Examples
L1 First response, gather facts, run read-only checks Acknowledge page, capture version/cluster status, open incident channel
L2 Diagnose from dashboards, execute runbooks Interpret Grafana panels, chronacta verify, rolling restart, follower lag
L3 Storage/HA engineering, code or data surgery Corruption, replication divergence, migration failures, patch releases

Escalate L1 → L2 within 15 minutes for SEV-1/SEV-2; L2 → L3 when runbooks do not resolve within SLA.

Level Definition Response target
SEV-1 Data loss risk, prolonged write outage, split-brain Immediate page; continuous work until mitigated
SEV-2 Degraded HA (no quorum / high lag), backup pipeline down Respond within 30 minutes
SEV-3 Single feature impaired (projection, subscription, Admin UI) Business hours
SEV-4 Docs, cosmetic, non-prod Backlog

Use Grafana overview before shell access:

Symptom Panel / metric Next step
Writes failing Engine ready (chronacta_engine_ready) Check /readyz, disk free space, troubleshooting.md
HA unstable Cluster leader, Replication lag chronacta cluster status; confirm single leader
Slow appends Append latency Compare with baseline; check disk sync, load
Projections stale Projection health, Lag by component chronacta projection list; resume/rebuild
Auth storm Auth and rate limits Review chronacta_auth_failures_total; check IdP / tokens
Disk filling On-disk storage Run scavenge dry-run; expand volume

Prometheus alerts: prometheus/chronacta-alerts.yaml. SLO table: monitoring.md.

  1. Acknowledge the page and open an incident channel.
  2. Capture: version (chronacta-server --version), cluster status, recent deploy, symptoms.
  3. Open Grafana Chronacta Overview dashboard; note firing alerts.
  4. Follow troubleshooting.md; prefer dry-run before destructive ops.
  5. Collect support bundle: chronacta support collect (see support-bundle.md).
  6. Take a backup before scavenge, restore, or membership surgery when data risk exists.
  1. On-call engineer (L2) → platform owner (L3) if SEV-1 not mitigated in 30 minutes.
  2. Involve storage/HA specialists for corruption or replication divergence.
  3. Customer-facing status updates for SEV-1/SEV-2 every 30 minutes until mitigated.
  • Timeline, root cause, blast radius, follow-ups.
  • Update runbooks if a gap blocked recovery.
  • File ops tickets for systemic fixes.
  • Prefer rolling upgrade (ha-rolling-upgrade.md).
  • Storage migrations: keep CHRONACTA_AUTO_MIGRATE=false unless the target release documents an automatic migration path.
  • Run pre-release checks before promote: make proto test test-race vet build.

Scenario: replication lag alert fires during business hours.

  1. L1 acknowledges and posts chronacta-server --version + link to Grafana.
  2. L2 confirms ChronactaReplicationLagHigh, checks leader panel and chronacta cluster status.
  3. L2 decides: transient load vs follower down; follows HA runbook.
  4. Document outcome; verify alert clears within SLO window.