---
name: oneuptime-observability-aware-coding
description: "Incident-aware, observable coding practices modeled on the OneUptime observability platform. Use when adding logging, metrics or alerts."
---

# OneUptime Observability-Aware Coding

- Instrument with OpenTelemetry conventions: propagate trace context across
  service and queue boundaries, and attach meaningful span attributes rather
  than logging free-form strings.
- Emit structured, machine-parseable logs with severity, and include the
  identifiers an on-call responder needs to pivot (request/trace ID, tenant
  or project ID, resource ID). Never log secrets or credentials.
- Write error paths for the 2 a.m. responder: error messages should state
  what failed, on which resource, and what was being attempted - not just
  rethrow a bare exception.
- Treat health and readiness endpoints, probes, and monitor checks as public
  contracts; changing their paths, status codes, or semantics silently can
  break external uptime monitors and alerting.
- Make failure handling explicit: timeouts on all outbound calls, retries
  only for idempotent operations, and degraded-mode behavior that keeps
  status/alerting paths alive when a dependency is down.
- Alert-worthy conditions belong in metrics or events, not buried in logs;
  ensure new failure modes are detectable by a monitor before shipping them.
