Polyaxon v3 is coming →

Prometheus exporters: tutorial and best practices

Learn how Prometheus exporters expose third-party metrics, how to build a simple exporter, and what practices keep metrics useful.

November 12, 2024by Polyaxon

Prometheus works best when targets expose metrics directly. Many systems do not. Exporters fill the gap by translating application, database, or infrastructure signals into Prometheus-friendly metrics.

The exporter itself should stay boring. Collect useful metrics, label them clearly, expose them reliably, and avoid turning the monitoring path into another fragile service.

What is a Prometheus exporter?

A Prometheus exporter collects measurements from a system that does not expose Prometheus-compatible metrics and translates them into a format Prometheus can scrape. It acts as an adapter between an application, database, or infrastructure service and the Prometheus server.

The exporter exposes metric names, optional labels, and numeric values through an HTTP endpoint, usually /metrics. Prometheus periodically reads that endpoint and stores the samples as time series for queries, dashboards, and alert evaluation. The exporter exposes the measurements; Prometheus stores their history.

How do Prometheus exporters work?

Prometheus implements an HTTP pull model to gather metrics from targets. When an application already exposes compatible metrics, Prometheus can scrape it directly. Otherwise, an exporter provides the translation layer. This is periodic metrics collection, rather than a stream of individual events or log records.

A Prometheus exporter's working mechanism typically involves the following:

  • Running an HTTP server with a metrics endpoint that Prometheus periodically queries.
  • Reading measurements from the application's API, status endpoint, or other supported source.
  • Translating those measurements into metric families with names, types, labels, and values.
  • Returning the metrics in a supported exposition format, typically using a client library.

Prometheus exporter implementation types

In an ecosystem of stateful and stateless applications, there are two common approaches to exposing metrics. Choosing between them depends on whether you can instrument the application itself.

Application built-in exporters

Applications can natively expose metrics such as request counts, errors, and duration through a Prometheus client library. This is usually called direct instrumentation; a separate exporter process is unnecessary. Common use cases include:

  • Developing an application from scratch: developers integrate instrumentation during application design, when request handling and other useful measurement points are easy to identify.
  • Instrumenting an existing application: teams modify the code to expose the metrics they need.

Third-party/standalone exporters

Standalone exporters expose metrics on behalf of another system. Applications may already provide a statistics API or status page, but its output needs translation before Prometheus can ingest it. An exporter can perform that translation without changing the application. Some exporters also derive metrics from logs, although logs and metrics serve different diagnostic purposes.

How to set up a Prometheus exporter for monitoring and alerts

The Prometheus project maintains a Python client library for building instrumentation and custom collectors. In this section, we'll build a basic exporter using Python, package it in a Docker image, deploy it on Kubernetes, and configure an existing Prometheus server to scrape it.

The example exposes an illustrative queue depth of 10. That fixed value makes it easy to follow a metric from Python through Kubernetes to Prometheus. It does not measure a real queue, cluster memory, or HTTP traffic. A real exporter would read its values from the system being monitored.

Prerequisites

You need:

  • A development Kubernetes cluster and kubectl configured for it, with permission to create a namespace, Deployment, and Service.
  • Docker installed, a Docker Hub repository you can push to, and a CLI session authenticated with docker login. Cluster nodes must be able to pull the image; a private repository needs an appropriate image pull Secret.
  • curl for inspecting the HTTP endpoint.
  • For the scraping and alerting steps, an existing Prometheus server running inside the cluster, with permission to change its scrape configuration and network access to the exporter.

The tutorial creates the namespace exporter-demo. Check that your current Kubernetes context points to the development cluster before applying the manifests. It does not install Prometheus or Alertmanager; the Kubernetes monitoring guide explains how exporters fit into the broader monitoring setup.

Creating the exporter using a Python script

Create the working directory and its code subdirectory, then enter the project directory:

mkdir -p custom-exporter/code
cd custom-exporter

Save the following complete script as code/collector.py. It has three parts: imports, a custom collector, and the HTTP server that makes its output available for scraping.

import time

from prometheus_client import REGISTRY, start_http_server
from prometheus_client.core import GaugeMetricFamily


class CustomCollector:
    def collect(self):
        queue_depth = GaugeMetricFamily(
            "demo_queue_depth",
            "Illustrative number of items waiting in a queue",
            labels=["queue"],
        )
        queue_depth.add_metric(["example"], 10)
        yield queue_depth


if __name__ == "__main__":
    REGISTRY.register(CustomCollector())
    start_http_server(8000, addr="0.0.0.0")
    while True:
        time.sleep(60)

collect() belongs inside CustomCollector: the registry calls that method to obtain the metric families. GaugeMetricFamily describes a value that may increase or decrease, while add_metric() adds a sample with the label queue="example". A cumulative request count would use a counter instead; this example only defines the queue-depth gauge.

Registration makes the collector available to the client library. start_http_server() starts the exporter's HTTP listener on port 8000; it does not start a Prometheus server. The sleep loop keeps the process alive, while HTTP requests trigger collection. The default registry also exposes Python process metrics. The Python custom collector documentation describes the collector interface and registration behavior.

Create code/pip-requirements.txt with the client dependency:

prometheus-client==0.26.0

Building the Docker image

From the custom-exporter directory, create a file named Dockerfile:

FROM python:3.13.15-slim-bookworm

WORKDIR /code
COPY code/pip-requirements.txt .
RUN pip install --no-cache-dir -r pip-requirements.txt
COPY --chmod=644 code/collector.py .

USER 65532:65532
EXPOSE 8000
CMD ["python", "collector.py"]

The image installs the dependency before copying the script so code changes can reuse the dependency layer. It runs the exporter as a non-root user. EXPOSE documents the listening port; the Kubernetes Service below provides access to it. The example pins a Python image and client release; keep those dependencies maintained when adopting the exporter.

Replace YOUR_DOCKERHUB_USERNAME in the commands and manifest with your Docker Hub username. Build an image for your cluster nodes' CPU architecture; an ARM laptop and an x86 cluster may need a cross-platform build rather than Docker's default local architecture.

docker build -t YOUR_DOCKERHUB_USERNAME/custom-exporter:0.1.0 .

Confirm that Docker lists the image with the expected repository and tag:

docker image ls YOUR_DOCKERHUB_USERNAME/custom-exporter:0.1.0

Push the image to Docker Hub by running the command:

docker push YOUR_DOCKERHUB_USERNAME/custom-exporter:0.1.0

Deploying the exporter in a Kubernetes cluster

Once the image is pushed, create a Deployment to run the exporter and a Service to give it a stable endpoint. These are ordinary Kubernetes resources; no custom operator is required.

Create the namespace and a folder for the manifests. Continue running commands from custom-exporter:

kubectl create namespace exporter-demo
mkdir templates

Save both resources below in templates/custom-exporter-deployment.yaml, including the --- separator, and replace the image's username:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: custom-exporter
  namespace: exporter-demo
spec:
  replicas: 1
  selector:
    matchLabels:
      app: custom-exporter
  template:
    metadata:
      labels:
        app: custom-exporter
    spec:
      automountServiceAccountToken: false
      securityContext:
        runAsNonRoot: true
        runAsUser: 65532
        runAsGroup: 65532
        seccompProfile:
          type: RuntimeDefault
      containers:
        - name: exporter
          image: YOUR_DOCKERHUB_USERNAME/custom-exporter:0.1.0
          ports:
            - name: metrics
              containerPort: 8000
          securityContext:
            allowPrivilegeEscalation: false
            capabilities:
              drop: ["ALL"]
          resources:
            requests:
              cpu: 50m
              memory: 64Mi
            limits:
              cpu: 250m
              memory: 128Mi
          readinessProbe:
            tcpSocket:
              port: metrics
            periodSeconds: 5
---
apiVersion: v1
kind: Service
metadata:
  name: custom-exporter
  namespace: exporter-demo
spec:
  selector:
    app: custom-exporter
  type: ClusterIP
  ports:
    - name: metrics
      protocol: TCP
      port: 8000
      targetPort: metrics

The Deployment and Service use the same app label, and the Service's targetPort resolves to the container port named metrics. ClusterIP keeps the endpoint internal. The resource values are small starting values for this demonstration; measure collection cost before using them for a real exporter. The TCP readiness probe checks that the listener accepts connections, while the next steps verify the metric output and a Prometheus scrape.

Apply the manifest and wait for the Deployment:

kubectl apply -f templates/custom-exporter-deployment.yaml
kubectl --namespace exporter-demo rollout status deployment/custom-exporter --timeout=180s

Inspect the Pod and Service:

kubectl --namespace exporter-demo get pods -l app=custom-exporter
kubectl --namespace exporter-demo get service custom-exporter

Expect one ready exporter Pod and a Service exposing port 8000. To inspect the exporter locally, start a port-forward and leave it running:

kubectl --namespace exporter-demo port-forward \
  --address=127.0.0.1 service/custom-exporter 8080:8000

In another terminal, request its metrics:

curl --fail http://127.0.0.1:8080/metrics

Among the process metrics, expect the sample defined in Python:

# HELP demo_queue_depth Illustrative number of items waiting in a queue
# TYPE demo_queue_depth gauge
demo_queue_depth{queue="example"} 10.0

The path is local port 8080 to Service port 8000 to the exporter's listener on 8000. This endpoint is the exporter, not the Prometheus UI.

Configure Prometheus to scrape the exporter

A running exporter is not automatically a Prometheus target. For a Prometheus server running inside the cluster, add this job to its existing scrape_configs list:

scrape_configs:
  - job_name: custom-exporter
    scrape_interval: 15s
    scrape_timeout: 5s
    metrics_path: /metrics
    static_configs:
      - targets:
          - custom-exporter.exporter-demo.svc:8000

Keep the server's other jobs. Apply and reload the configuration through the mechanism used by your Prometheus installation. If a Helm chart or Prometheus Operator manages that configuration, use its supported configuration or monitor resources instead of editing a generated file. A Service or scrape annotation alone does not configure every Prometheus installation. See the Prometheus configuration reference for scrape jobs and discovery options.

The .svc address uses cluster DNS and must be reachable from Prometheus; it is not a workstation address. This static Service target is sufficient for the single-replica exercise. For multiple exporter instances, discover and scrape each instance separately so load balancing does not mix different instances' samples into one target.

In your existing Prometheus UI, check the job's target status and query:

up{job="custom-exporter"}

After a successful scrape, the value should be 1. Then query the sample:

demo_queue_depth{job="custom-exporter", queue="example"}

Expect 10. Together, these checks show that Prometheus can reach the exporter and ingest its metric. They do not demonstrate the health of a real queue, since this tutorial uses an illustrative value.

Troubleshoot collection and access

Check each step between the container and Prometheus:

SymptomWhat to inspect
ImagePullBackOffThe image username and tag, registry access, and any required image pull Secret.
Container exits or restartsThe Python traceback in the logs, installed dependencies, image architecture, and resource limits.
Port-forward cannot select a ready PodThe Service selector, Pod labels, readiness, and named container port.
No custom-exporter target in PrometheusWhether the scrape job was loaded or the installation's discovery configuration selected the exporter.
up is 0The target's scrape error, cluster DNS, network policies, port, and HTTP response.
up is 1 but the sample is absentCollector registration, the image version, the metric name, and any metric relabeling rules that drop samples.

These commands expose the container logs and Deployment events:

kubectl --namespace exporter-demo logs deployment/custom-exporter
kubectl --namespace exporter-demo describe deployment custom-exporter
kubectl --namespace exporter-demo describe pods -l app=custom-exporter

For a real exporter, a successful HTTP scrape and a successful upstream collection are separate checks. Give upstream requests bounded timeouts and report failures explicitly. The exporter-writing guide describes failure reporting through a failed scrape or an exporter-specific health metric; a failed measurement should not silently become a plausible zero.

Best practices when using Prometheus exporters

Some best practices to adopt when using Prometheus exporters include:

Use an existing Prometheus exporter

Before building a custom exporter, check for an existing exporter that covers the system and metrics you need. Maintaining a custom collector adds work: upstream APIs change, labels need stable meanings, and collection failures need diagnosis. Evaluate the project's maintenance, supported application versions, required privileges, and collection cost alongside its metric coverage.

The official exporters and integrations list is a useful starting point. Inclusion does not mean every exporter is maintained by the Prometheus project; check the individual repository before adopting it.

Use labels and annotations to help understand metrics

Each exporter exposes a particular set of measurements. Clear names, units, and HELP descriptions make those measurements understandable. Labels add dimensions for grouping and querying, but metric labels and Kubernetes metadata have different roles:

  • Instrumentation labels describe a measurement inside the application, such as the example's queue="example". Keep the possible values bounded.
  • Target labels describe the scrape target. Prometheus adds labels such as job and instance; discovery and relabeling can add selected infrastructure metadata.
  • Kubernetes labels and annotations describe Kubernetes objects. Labels support object selection, while annotations can carry additional metadata or tool-specific configuration. Neither becomes a Prometheus metric label automatically; discovery and relabeling determine what is retained.

Avoid user IDs, request IDs, arbitrary paths, or unique run IDs as general-purpose metric labels. Every unique combination creates another time series. The metric naming guide explains naming and cardinality; our Kubernetes monitoring article applies those choices to ML workloads.

Configure actionable alerts

Apart from capturing the right metrics, monitoring teams need alerts for conditions that require action. If a threshold is too sensitive, the team receives unnecessary alarms. If it is too lenient, the team may miss a sustained failure. Choose thresholds and durations using the service's workload and the impact on its users.

For this exporter, up{job="custom-exporter"} == 0 with a for: 5m duration can identify sustained scrape failures. That duration is illustrative. A target removed from configuration needs a separate missing-target check, because it no longer produces up = 0. For a real queue, also consider waiting time and whether work is progressing, rather than alerting on queue depth alone.

Prometheus evaluates configured alert rules; Alertmanager handles notification routing, grouping, and silences. Rules must be loaded into Prometheus, and notification delivery requires a configured Alertmanager integration. Include an owner and a useful diagnostic next step with each alert. The Prometheus alerting guidance and our Alertmanager guide cover those responsibilities.

Adopt an appropriate scaling mechanism

As deployments grow, more exporters and time series increase collection, memory, storage, and query costs. Track active series, ingestion rate, scrape duration, query latency, and retention needs. Remove unused metrics and unbounded labels before increasing capacity.

Choose a scaling mechanism for the actual bottleneck. Dividing scrape targets across Prometheus instances distributes collection work. Federation can collect selected series from other Prometheus servers. Remote storage can provide longer retention and a broader query layer, but remote write does not remove the local cost of scraping or make exporter discovery automatic. See the Prometheus storage documentation and our managed Prometheus guide for the operating tradeoffs.

Administer reliable metric access privileges

Exporters can expose sensitive information about applications, services, and hosts. Keep metrics endpoints private and restrict access to the monitoring components that need them. The demo's ClusterIP Service avoids publishing a node port, but it does not itself prevent other Pods from connecting. Use appropriate network policies and workload boundaries.

Kubernetes RBAC controls access to Kubernetes resources; it does not automatically authorize HTTP requests to an exporter or Prometheus. Configure authentication and TLS explicitly where required. Prometheus supports HTTPS and authentication, but those protections are not enabled merely by starting the server. Apply the corresponding protection to each exporter or its proxy, and consult the Prometheus security model when designing access.

Clean up the example

Stop the local port-forward with Ctrl+C. Remove the custom-exporter scrape job and any demo alert rules from your Prometheus configuration, then apply that configuration through your installation's normal process. From the custom-exporter directory, remove the tutorial resources:

kubectl delete -f templates/custom-exporter-deployment.yaml

If exporter-demo was created only for this tutorial and contains no other workloads, delete it:

kubectl delete namespace exporter-demo

Final thoughts

Exporters are glue. Good glue is boring: stable endpoints, clear labels, limited cardinality, and alerts that point to action instead of panic.

Polyaxon teams can use Prometheus-style metrics alongside run metadata to connect infrastructure health with ML workload behavior. The exporter gives the metric; the platform context tells you why it matters. Start with platform observability for the infrastructure view and run monitoring for run statuses, events, and resource consumption. Correlate the affected workload and time window without copying every unique run identifier into metric labels.