WarehousePG Enterprise Manager (WEM) collects metrics through two separate channels. An OTel Collector pipeline gathers WHPG cluster and host hardware metrics and stores them in ClickHouse, where the WEM interface queries them directly. Two separate endpoints can also expose these metrics directly, in Prometheus format. WEM exposes its own internal metrics, serving WEM's own state only, at /prom/metrics. The OTel gateway collector separately exposes WHPG cluster and host metrics at <host>:<port>/metrics, for an external Prometheus instance to scrape directly, when you set OTEL_PROMETHEUS_SCRAPE_ENDPOINT in wem.conf. See Configuration reference.

Note

Earlier WEM versions collected cluster metrics with a standalone SQL exporter that pushed warehousepg_observability_* metrics to Prometheus via remote write. That pipeline was removed and replaced with the OTel Collector/ClickHouse pipeline described below. If you have external tooling or dashboards built against the old warehousepg_observability_* metric names, they need to be rebuilt against the metric names in this section.

WHPG cluster and host metrics

The OTel Collector running on each cluster node gathers these metrics and writes them to ClickHouse. Metrics carry a whpg_host attribute identifying the node (and, for per-device or per-state breakdowns, additional attributes as noted). The WEM interface queries these from ClickHouse. Set OTEL_PROMETHEUS_SCRAPE_ENDPOINT to also expose them on a Prometheus scrape endpoint for an external Prometheus instance.

Cluster connectivity

MetricTypeDescription
whpg.cluster.connectedGaugeIndicates whether the database connection is valid (1=connected, 0=not connected).

Coordinator and segment status

MetricTypeDescription
whpg.coordinator.active_up_countGaugeNumber of active coordinator nodes currently up.
whpg.coordinator.active_down_countGaugeNumber of active coordinator nodes currently down.
whpg.coordinator.standby_up_countGaugeNumber of standby coordinator nodes currently up.
whpg.coordinator.standby_down_countGaugeNumber of standby coordinator nodes currently down.
whpg.segment.total_primary_countGaugeTotal number of primary segments in the cluster.
whpg.segment.total_mirror_countGaugeTotal number of mirror segments in the cluster.
whpg.segment.primary_down_countGaugeNumber of primary segments currently down.
whpg.segment.mirror_down_countGaugeNumber of mirror segments currently down.
whpg.segment.statusGaugePer-segment up/down status (1=up, 0=down).

Monitoring query and connection states

MetricTypeDescription
whpg.queries.totalGaugeTotal number of connections.
whpg.queries.activeGaugeNumber of connections with active queries.
whpg.queries.idleGaugeNumber of idle connections.
whpg.queries.idle_in_transactionGaugeNumber of connections idle inside an open transaction.
whpg.queries.blockedGaugeNumber of queries blocked waiting for locks.
whpg.queries.long_running_120sGaugeNumber of queries that have been running for more than 120 seconds.
whpg.queries.in_waitGaugeNumber of queries currently in a wait state.
whpg.query.max_duration_secondsGaugeDuration of the longest-running query currently active, in seconds.
whpg.transactions.total_executedGaugeTotal number of queries executed.

Tracking database sizes and skew

MetricTypeDescription
whpg.db.sizeGaugeSize of each database in bytes.
whpg.data_skew.cvGaugeCoefficient of variation of data distribution across segments, used to detect data skew.

Collecting host hardware metrics

These metrics reflect the physical resources of each cluster host, collected by the OTel Collector's host metrics receiver using OpenTelemetry semantic-convention names.

MetricTypeDescription
system.cpu.timeCounterCumulative CPU time, broken out by the state attribute (user, system, idle, iowait, nice, softirq, steal, interrupt, wait).
system.memory.usageGaugeMemory usage in bytes, broken out by the state attribute (used, free, cached, buffered, slab_reclaimable, slab_unreclaimable).
system.memory.limitGaugeTotal physical memory on the host, in bytes.
system.cpu.load_average.1mGauge1-minute load average.
system.cpu.load_average.5mGauge5-minute load average.
system.cpu.load_average.15mGauge15-minute load average.
system.disk.ioCounterCumulative bytes read from or written to disk, broken out by the direction (read, write) and device attributes.
system.network.ioCounterCumulative bytes received or transmitted over the network, broken out by the direction (receive, transmit) and device attributes.
system.filesystem.usageGaugeFilesystem usage in bytes, broken out by the state (used, free, reserved), mountpoint, type, and device attributes.

Querying the ClickHouse tables directly

WEM's interface queries the metrics and logs above from ClickHouse through its own API, so you don't need direct ClickHouse access for day-to-day use. If you want to query them directly instead, for example from a BI tool or a custom Grafana panel, this section documents the underlying tables.

These tables are created and populated by the OTel Collector's own ClickHouse exporter, using that exporter's default schema, not a schema WEM defines itself — column names follow OTel Collector conventions rather than WEM's own naming. They live in the database set by CLICKHOUSE_DB (default acp_observability).

Metrics tables

TableDescription
otel_metrics_gaugePoint-in-time metrics, for example connection counts, load averages, and segment status.
otel_metrics_sumCumulative counter metrics, for example CPU time and disk and network I/O.

WEM's own queries read these columns from both tables:

ColumnDescription
MetricNameThe metric name, matching the names listed under WHPG cluster and host metrics above (for example whpg.queries.active, system.cpu.time).
ValueThe metric's numeric value.
TimeUnixTimestamp the sample was collected.
AttributesMap of string labels for the sample. WEM's own queries rely on the whpg_host key (the reporting node), plus metric-specific keys such as state, direction, device, mountpoint, and type — see the metric descriptions above for which attributes apply to which metric.
Note

The OTel Collector's ClickHouse exporter also creates otel_metrics_histogram, otel_metrics_exponential_histogram, and otel_metrics_summary tables by default, and WEM doesn't currently query any of them. Confirm with engineering whether these tables are actually populated in a WEM deployment.

Logs table

TableDescription
otel_logsWHPG log entries collected from each node. This replaced WEM's earlier Loki-based log pipeline.

WEM's own queries read these columns:

ColumnDescription
TimestampWhen the log line was collected.
ServiceNameThe OTel service name of the collecting node.
BodyThe raw log line.
TraceIdTrace ID associated with the log entry, if any.
ResourceAttributesMap of resource-level attributes, such as the reporting host.
LogAttributesMap of per-log attributes parsed from the WHPG CSV log. WEM's Logs panel filters and summary counts rely on the severity, log_user, and database keys specifically.
Note

By default, ClickHouse retains OTel-collected metrics and logs indefinitely. To change the retention policy, edit the internal tables in CLICKHOUSE_DB, otel_logs for logs, and each otel_metrics_* table (such as otel_metrics_gauge) for the metric types you collect:

ALTER TABLE acp_observability.<table_name> MODIFY TTL <timestamp_column> + INTERVAL <retention_days> DAY;

Substitute <table_name> and <retention_days> with your values, and <timestamp_column> with that table's own timestamp column, Timestamp for otel_logs, TimeUnix for otel_metrics_* tables.

WEM internal metrics

These metrics describe the health and behavior of WEM itself. Prometheus scrapes them directly from the /prom/metrics endpoint, which is unauthenticated. All WEM internal metrics use the prefix wem_.

Monitoring canary checks

Each canary metric carries a check_name label identifying the canary check it belongs to.

MetricTypeDescription
wem_canary_duration_msGaugeExecution time of the most recent canary check run, in milliseconds.
wem_canary_statusGaugeResult of the most recent canary check run (0=success, 1=warning, 2=critical).
wem_canary_last_run_timestampGaugeUnix timestamp of the most recent canary check execution.
wem_canary_row_countGaugeRow count returned by the most recent canary check query.
wem_canary_checks_totalCounterTotal number of canary check executions since WEM started.
wem_canary_failures_totalCounterTotal number of canary check executions that returned a warning or critical result.

Tracking WEM system state

MetricTypeDescription
wem_upGaugeIndicates whether the WEM process is running (1=running).
wem_scheduler_lock_heldGaugeIndicates whether this WEM instance currently holds the shared scheduler lock used by canary checks and resource-usage collection (1=held). WEM's alert evaluator, log scan, and Safeguard config-drift sync run under their own separate locks, which aren't currently reflected in this metric.
wem_build_infoGaugeStatic build metadata for this WEM instance, with version, commit, and go_version labels. Always 1.

Monitoring connection pools

Each pool metric carries a pool_name label (for example whpg-main, wem-state, or query-editor) identifying the connection pool it describes. Because the Query Editor opens a pool per session, pool_name values for it are dynamic and unbounded in cardinality.

MetricTypeDescription
wem_pool_total_connsGaugeTotal number of connections in the pool (idle and acquired).
wem_pool_idle_connsGaugeNumber of idle connections available in the pool.
wem_pool_acquired_connsGaugeNumber of connections currently in use.
wem_pool_max_connsGaugeMaximum number of connections configured for the pool.
wem_pool_utilization_percentGaugePool utilization expressed as a percentage of the maximum connection limit.

Could this page be better? Report a problem or suggest an addition!