Probe diagnostics v10.6

Probe diagnostics gives you a fleet-wide view of probe health: which probes are failing, degraded, or stale — including probes left stale by an unreachable agent — and how long each one takes to run. It has two pages: a probe-health list across every target, and a per-probe detail page with charts and a per-target breakdown.

To open it, select Management > Probes > Probe Diagnostics.

Prerequisites

  • You need the pem_manage_probe privilege to see the tab at all. Without it, PEM shows an access-denied message and the underlying data endpoints also refuse the request.
  • Probe telemetry must be enabled. It's controlled by the probe_telemetry_enabled configuration parameter, which ships on by default. If an administrator turns it off, Probe Diagnostics shows no data until it's turned back on.
  • An agent must support probe telemetry to appear in the diagnostics data. Unsupported agents are listed by name in the dashboard itself, so you don't need to check versions manually.

Probe health

The list page shows one row per probe, aggregated across every target it runs on:

ColumnContents
ProbeThe probe's display name.
Is System probe?Whether it's a built-in probe or a custom one.
Target typeWhat kind of object the probe runs on: Agent, Tool, Server, Database, Schema, or Extension.
ScopeThe object type the probe applies to.
TargetsThe number of objects this probe currently runs on.
Target StatusA count of this probe's targets in each health state. See Target status.
Success %The success rate over the selected time window.
RunsThe number of runs over the selected time window.
Execution TimeAverage, 95th percentile, and maximum run time over the window.

Filters

FilterNotes
ForThe time window: 1 hour, 4 hours, 12 hours, 24 hours (default), 3 days, or 7 days. This is the only filter applied on the server; every other filter runs in your browser over the rows already loaded.
Prior toThe end of the time window. Defaults to now.
Servers / AgentsNarrow the list to specific servers or agents.
Show System Probes?Include or exclude built-in probes.
Search probes…Filter by probe name.

A fixed set of target-type buttons (Agent, Tool, Server, Database, Schema, Extension) is always shown on the toolbar, even when no probe of that type falls in the current window. Use Columns to choose which columns are visible, and Refresh to reload the list. Column definitions opens a short legend explaining Target Status, Targets, Success % / Runs, and Execution Time.

Target status

Each target a probe runs on is in exactly one of five health states, checked in this order — the first match wins:

StatusMeaning
DisabledThe probe is turned off for this target.
StaleEnabled but hasn't reported on schedule — its last run (as of the selected window) is older than about twice its execution frequency, plus a short grace. A non-reporting agent also surfaces here.
FailingThe last run failed, or the success rate over the window is under 50%.
DegradedThe success rate over the window is between 50% and 85%.
HealthyThe success rate over the window is above 85%, and the last run succeeded.

Status reflects the selected Prior to window, not necessarily right now — a historical window shows each target's state as of that point in time, so viewing the past won't show every target as stale.

Select Target status legend on either page to see this same reference without leaving the dashboard.

Probe detail

Select a probe from the list to open its detail page. It has eight KPI tiles, all scoped to the selected time window and target: Target status, Success %, Runs, Failures, average execution time, 95th percentile execution time, maximum execution time, and Metric writes.

Charts

ChartShows
Executions & failuresSuccessful vs. failed probe executions per time bucket.
Time breakdownAverage time per run, split into query execution, sync/commit, and connection-pool wait, stacked to the total.
Execution timePer-run execution time: average, 95th percentile, and maximum.

The window is split into roughly 100 time buckets for these charts, so the bucket size adjusts automatically to how wide a window you select.

Per-target breakdown

Use Select target to scope the whole page to one target, or leave it on All targets. The breakdown table lists every target the probe runs on:

ColumnNotes
StatusThe target's current health state.
TargetThe target's breadcrumb (for example, a server, or a server and database).
ProfileThe profile managing this target's probe configuration, if any. Hidden by default.
Success %Success rate over the window.
RunsNumber of runs over the window.
Execution Time (Average / 95th percentile / Maximum)Shown by default.
SyncAverage sync/commit time. Hidden by default.
Pool waitAverage time spent waiting for a database connection. Hidden by default. Only applies to probes that run SQL queries against a target server.
Metric Writes (Insert / Update / Delete)Rows written to the metrics tables over the window. Shown by default.
Last runWhen this target last ran.
Last errorThe error from the target's most recent run, tagged EXEC or SYNC depending on which phase it came from. Clears once a later run succeeds, so a target that has recovered shows no error. Shown by default.

Use Columns to show or hide the optional columns.

Targets that are dropped or unregistered fall off this list — that's expected, not data loss. The list always reflects what PEM currently monitors; history for a target that no longer exists stops contributing to the probe's aggregate numbers on the list page too.

Configuring a probe for one target

Select Configure probe for this target to open the inline configuration form for a specific target, without leaving the diagnostics page.

FieldNotes
Probe enabled?Default follows the probe's own setting, or Enabled/Disabled overrides it for this target.
Execution frequency (HH:MM:SS)Untick Use default to set a custom interval. Must be between 00:00:10 and 48:00:00.
History retention (days)Untick Use default to set a custom retention period. Must be a whole number of days between 1 and 365. The out-of-the-box default is 30 days.

Two things affect whether you can use this form:

  • Without the pem_config_probe privilege, the form is read-only.
  • If a profile manages the target's configuration, the read-only form notes that profile changes affect every assigned target, with a link to the profile. If you're missing pem_config_probe, that note isn't shown — no-privilege takes precedence over profile-managed.

The Manage Profile link itself is hidden unless you also hold pem_manage_profile.

Privileges and access control

PrivilegeControls
pem_manage_probeThe entire Probe Diagnostics tab. Without it, the tab shows an access-denied message, and the data endpoints return the same result if called directly.
pem_config_probeWhether the inline per-target configuration form is editable. Without it, the form is read-only regardless of profile management.
pem_manage_profileThe Manage Profile deep-link on a profile-managed target.

Holding pem_manage_probe doesn't automatically show every target in your fleet. Each target is additionally checked against your team's visibility on the underlying agent and server — a target whose agent or server your team can't see doesn't appear in Probe Diagnostics, even though you can otherwise use the feature.

Data source and retention

Probe Diagnostics reads from pemhistory.probe_telemetry, populated by agents that support probe telemetry. History retention isn't a single fleet-wide setting — it's configured per probe, the same History retention (days) setting described above, defaulting to 30 days and adjustable from 1 to 365.

Troubleshooting

SymptomCause and what to do
Probe Diagnostics shows no data at allProbe telemetry is disabled. An administrator needs to turn on probe_telemetry_enabled.
An agent's probes never show upThe agent's version doesn't support probe telemetry. Check the unsupported-agents notice on the dashboard.
I can't see the tabYou don't have the pem_manage_probe privilege.
The configuration form is read-only and I don't see whyEither you're missing pem_config_probe, or the target is profile-managed — the form tells you which.
Saving the configuration form fails with a frequency errorExecution frequency must be between 00:00:10 and 48:00:00 (HH:MM:SS).
Saving the configuration form fails with a retention errorHistory retention must be a whole number of days between 1 and 365.
A target I removed still shows old data on the probe's aggregate numbersIt shouldn't — targets no longer monitored stop contributing once PEM's view of current targets updates. If it persists, refresh the page.