Cronitor API

Custom Metrics

Job and heartbeat monitors can carry named metrics on each telemetry event. Built-in metrics (count, error_count, length, latency, duration) retain their existing behavior and units. Any other valid name is stored as a custom metric and can be discovered, graphed, and asserted on.

This page covers the full loop: send, discover, graph, assert, and troubleshoot. The same assertion strings work in the Monitors API, YAML, the dashboard text input, and MCP.

Send metrics

Attach metrics to run, tick, complete, and fail events as repeated metric query parameters. The canonical form is name:value. name=value is also accepted.

export CRONITOR_PING_URL="https://cronitor.link/p/YOUR_API_KEY/YOUR_MONITOR_KEY"

curl --fail --get "$CRONITOR_PING_URL" \
  --data-urlencode "state=complete" \
  --data-urlencode "metric=queue_depth:142" \
  --data-urlencode "metric=quality_score:0.87"

Replace YOUR_API_KEY with an API key that has monitor:telemetry permission and YOUR_MONITOR_KEY with your monitor key. See Telemetry API for the URL components and job and heartbeat examples.

Zero, negative, and fractional custom metric values are valid. Non-finite values (NaN, Inf) are dropped.

Each event is its own sample. Sending the same name and value on both run and complete records two observations.

Names and limits

RuleDetail
CharactersLetters, digits, underscore, and hyphen only (a-z, 0-9, _, -)
LengthAt most 63 characters after trim
NormalizationNames are trimmed and lowercased. Queue_Depth is stored as queue_depth
Per eventAt most 10 custom metrics. The first 10 valid custom names are kept
DuplicatesThe last valid value for a repeated name wins
Built-inscount, error_count, length, latency, and duration do not count against the 10-name limit

There is no per-monitor name registry and no account-wide name cap. An assertion can be saved before its metric first arrives; it stays skipped until data exists.

Dropped metrics

Malformed names, over-length names, non-finite values, and metrics beyond the first 10 valid custom names are dropped silently. The event and its valid metrics are still stored. The HTTP response means the ping was received, not that every metric was stored. Cronitor does not return an error body or increment a user-visible counter for dropped metrics.

Repeat metric= once per value. Do not combine several metrics into one comma-separated value.

Discover names

Known custom names for a monitor and environment come from the events in the selected range, most recently seen first, capped at 100 names.

curl "https://cronitor.io/api/metrics?monitor=queue-worker&env=production&time=7d&metricNames=true" \
  -u API_KEY:
{
  "metric_names": {
    "queue-worker": {
      "production": ["quality_score", "queue_depth"]
    }
  }
}

Monitor detail responses include metric_names for the current environment over the last 7 days. MCP get_monitor exposes the same list. MCP get_metrics with metric_names: true uses the selected range.

Do not pass a built-in name as metric. Built-in series use the field parameter on the Metrics API. Without time, start, or end, that endpoint defaults to the last hour. Monitor detail and MCP get_monitor discovery stay on the last 7 days.

Graph metrics

On a job or heartbeat monitor, the chart selector includes a Metrics option. Pick one custom name. The line shows the average value per time bucket, with the sample count in the tooltip. Heartbeats default to this chart when custom names are discovered in the selected range, unless you have explicitly selected another chart. Bucket width follows the existing range conventions: hourly unless the selected range is longer than 14 days, then daily.

Missing buckets are gaps, never zeros. The dashboard and the API use the same query. Request withNulls=true when a charting client needs explicit null points for empty buckets.

curl "https://cronitor.io/api/metrics?monitor=queue-worker&env=production&time=24h&metric=queue_depth" \
  -u API_KEY:
{
  "monitors": {
    "queue-worker": {
      "production": [
        {"stamp": 1640995200, "avg": 142.0, "count": 12},
        {"stamp": 1640998800, "avg": 0.0, "count": 4}
      ]
    }
  }
}

avg: 0 with a nonzero count is a real zero. A missing stamp, or a null avg/count when withNulls=true, means no samples in that bucket.

MCP get_metrics accepts exactly one of fields, metric, or metric_names. REST GET /api/metrics accepts field, metric, or metricNames; metric cannot be combined with field or metricNames.

Expanded event details on jobs and heartbeats show custom metrics on that event, including explicit zeros. Heartbeat Latest Activity summarizes metrics from the latest eligible event; expand the row to see each event’s metrics. Built-in zero values are omitted from heartbeat activity because they cannot be distinguished from missing values.

Assert on metrics

Metric assertions are strings. REST, YAML, MCP, and the dashboard store and validate the same expression. In the dashboard, add a Metric Assertion and type the rest of the expression after the fixed metric. prefix.

metric.<name> <operator> <value>
metric.<name>.<aggregate> <operator> <value> over <period>

Most assertions state the healthy condition. Cronitor alerts when telemetry violates it. "metric.queue_depth < 100" means a healthy event is below 100, so 100 or more triggers an alert.

FormExampleEvaluates
Latestmetric.queue_depth < 100The latest eligible event's value
Latestmetric.quality_score >= 0.8The latest eligible event's value
Windowmetric.queue_depth.mean < 100 over 24 hoursMean of samples in the trailing window
Windowmetric.queue_depth.p99 < 250 over 1 hourApproximate 99th percentile over the window
Windowmetric.quality_score.min > 0 over 1 hourMinimum over the window

Aggregates are sum, mean, min, max, p50, p90, p95, and p99. A window is required for aggregates and forbidden on latest-value assertions. Windows are at most 90 days.

Operators are <, <=, =, >, >=, and !=. Values are finite numbers. Window assertions on duration and latency accept time values such as 30s or 500ms; a bare number is in the metric's own unit (seconds for duration, milliseconds for latency).

These assertions apply to job and heartbeat monitors. They can be saved before the metric exists.

curl --user API_KEY: \
--header "Content-Type: application/json" \
--request PUT \
--data '{
    "monitors": [
        {
            "type": "heartbeat",
            "key": "queue-worker",
            "schedule": "every 1 minute",
            "assertions": [
                "metric.queue_depth < 100",
                "metric.quality_score >= 0.8",
                "metric.queue_depth.mean < 100 over 24 hours"
            ]
        }
    ]
}' \
https://cronitor.io/api/monitors

The assertions array replaces the monitor's existing assertions. Include the full desired set on every update.

heartbeats:
  queue-worker:
    schedule: "every 1 minute"
    assertions:
      - "metric.queue_depth < 100"
      - "metric.quality_score >= 0.8"
      - "metric.queue_depth.mean < 100 over 24 hours"
    notify:
      - default

Built-in latest assertions (metric.duration, metric.count, metric.error_count, metric.error_rate) keep their existing types and units. The window form works with stored metric names — for example "metric.duration.p95 < 30s over 1 day" or "metric.count.sum > 0 over 1 hour". error_rate is a derived latest assertion, not a stored name, so it has no window form.

Evaluation

  • Latest custom assertions look at the latest eligible run, complete, tick, or fail event only. If that event omits the metric, the rule is skipped. Cronitor does not fall back to an older event.
  • Window assertion results can be cached for up to three minutes. A new monitor event triggers an earlier refresh for metric window assertions. Values aging out of the window can change the result without a new event.
  • Windows longer than 7 days refresh on a timer only. New events do not trigger an earlier refresh. The interval is the window divided by 720, with a three-minute minimum: about 1 hour for 30 days and 3 hours for 90 days.
  • Missing data is skipped, never treated as zero. An empty window cannot recover an open issue.
  • An explicit 0 on a custom name is a real sample and is included in latest checks, window aggregates, and charts. Built-in column metrics (count, error_count, length, latency, duration) still treat zero as absent.
  • failure_tolerance does not apply to metric assertions. They alert on the first violating value.
  • Percentiles (p50, p90, p95, p99) are approximate.
  • A recovery on any rule resolves the incident. If another metric assertion is still failing, it re-alerts on its next evaluation.

Alerts include the full assertion string and a readable sentence, for example: metric.queue_depth averaged 142 over the last 24 hours; expected less than 100.

Troubleshoot

I sent a metric and it does not appear. Check the name against the character and length rules. Invalid, over-length, non-finite, and 11th-and-later custom names are dropped silently. A 200 on the ping URL does not confirm that a given metric was stored. Repeat the name exactly as you intend to query it; it is stored lowercased.

The assertion never fires. Missing data is skipped. If the latest event does not carry the name, a latest assertion does nothing even when older events did. An empty window is skipped and cannot recover an issue. Confirm the metric is present on recent events in expanded activity, then confirm the assertion string round-trips on a monitor GET.

The chart shows a gap, not a zero. Gaps mean no samples in that bucket. Send an explicit 0 when zero is a real measurement. withNulls=true returns those gaps as nulls; it does not fill zeros.

Discovery is missing a name. Names are listed for one monitor and environment, most recent first, limited to 100. Names that have not appeared in the selected range (7 days by default on monitor detail) are omitted. An assertion can still be configured for a name that has not arrived yet.

A window assertion was rejected. Aggregates require over <period>, latest assertions cannot include over, and the period cannot exceed 90 days.

failure_tolerance did not absorb a bad sample. That setting applies to state=fail events and failed checks, not to metric assertions.

Two thresholds on the same metric share history. Latest or window assertions that share the same name, aggregate, and window are grouped together even if their thresholds differ. A threshold edit inherits the previous failure state. Use a different aggregate or window when you need independent history.

Run and complete both sent the value and the sum looks doubled. Each event is a sample. Send the metric on the event that represents the measurement, not on both, unless you intend two observations.

Previous
Telemetry