Skip to content

Troubleshooting runbooks

Operational playbooks for Chronacta single-node and HA deployments.

Symptom First checks
gRPC unavailable Process up? Port open? chronacta health -server HOST:PORT
Append rejected Expected version? Auth/RBAC? Disk space / min free bytes? Quotas?
Read lag / empty Wrong stream ID? Scavenge removed indexes? Wrong from version?
Admin UI login fails Auth enabled? Token store path? Development mode only for labs
Backup upload fails Remote credentials, bucket, network; list-remote
Follower lag Leader healthy? Network partition? cluster status lag fields
  1. chronacta storage-status — note data_bytes, readiness reason.
  2. Dry-run scavenge: chronacta scavenge -max-events-per-stream N -dry-run.
  3. Review plan; execute with -confirm only after backup.
  4. Create backup: chronacta backup create … (+ auto-upload if configured).
  1. Confirm process and host health on each node.
  2. chronacta cluster status from a reachable node.
  3. If quorum lost, restore connectivity before membership changes.
  4. Promote only with documented procedure (ha-rolling-upgrade.md).
  1. Stop writes on contested nodes (take offline or fence).
  2. Identify the node with the highest committed global position / epoch from durable storage + cluster metadata.
  3. Bring that node up as sole writer; rejoin followers from snapshot or catch-up.
  4. Never force two leaders with overlapping epochs — discard or reseed the divergent replica after backup.
  1. chronacta subscription list / Admin UI Subscriptions.
  2. Inspect lag, in-flight, dead letters.
  3. Pause → replay from a safe checkpoint → resume.
  4. Ack stuck in-flight only when safe for the consumer.

Default (v1.1+): slow live subscribers enter catch-up mode; the stream stays open and SubscribeResponse.control may report CATCHING_UP. Tune client processing or use pkg/client/resilience SubscribeLive for auto-resubscribe after server restarts.

Legacy disconnect: set CHRONACTA_SUBSCRIPTION_CATCHUP_MODE=false to restore pre-v1.1 behavior (ErrSlowSubscriber disconnect).

  1. Check chronacta-connector logs and job handler HTTP status.
  2. Verify subscription checkpoint advances (subscription get -id JOB_SUB).
  3. Handlers must be idempotent on event_id (at-least-once delivery).
  4. See connector.md.
  1. Admin UI Projections → Detect stalled / status panel.
  2. Check errors and dead letters.
  3. Resume or rebuild from a known-good checkpoint after fixing the program.

Prometheus: chronacta_* including tenant labels chronacta_tenant_events_total{tenant_id=…}. See monitoring.md.