Skip to content

Observability

# Observability & Monitoring

This guide covers metrics collection, monitoring, and visualization for eoAPI deployments. All monitoring components are optional and disabled by default.

Overview

eoAPI observability is implemented through conditional dependencies in the main eoapi chart:

Core Monitoring

Essential metrics collection infrastructure including Prometheus server, metrics-server, kube-state-metrics, node-exporter, and prometheus-adapter.

Integrated Observability

Grafana dashboards and visualization tools are available as conditional dependencies within the main chart, eliminating the need for separate deployments.

Configuration

Prerequisites: Kubernetes cluster with Helm 3 installed.

Quick Deployment

# Deploy with monitoring and observability enabled
helm install eoapi eoapi/eoapi \
  --set monitoring.prometheus.enabled=true \
  --set observability.grafana.enabled=true

# Access Grafana (see "Accessing Grafana" section below for credentials)
kubectl port-forward -n eoapi svc/eoapi-obs-grafana 3000:80

Using Configuration Files

For production deployments, use configuration files instead of command-line flags:

# Deploy with integrated monitoring and observability
helm install eoapi eoapi/eoapi -f values-full-observability.yaml

For a complete example: See production profile

Architecture & Components

Component Responsibilities:

  • Prometheus Server: Central metrics storage and querying engine
  • metrics-server: Provides resource metrics for kubectl top and HPA
  • kube-state-metrics: Exposes Kubernetes object state as metrics
  • prometheus-node-exporter: Collects hardware and OS metrics from nodes
  • prometheus-adapter: Enables custom metrics for Horizontal Pod Autoscaler
  • Grafana: Dashboards and visualization of collected metrics

Data Flow: Exporters expose metrics → Prometheus scrapes and stores → Grafana/kubectl query via PromQL → Dashboards visualize data

Detailed Configuration

Basic Monitoring Setup

# values.yaml - Enable core monitoring in main eoapi chart
monitoring:
  metricsServer:
    enabled: true
  prometheus:
    enabled: true

prometheus:
  server:
    persistentVolume:
      enabled: true
      size: 50Gi
    retention: "30d"
  kube-state-metrics:
    enabled: true
  prometheus-node-exporter:
    enabled: true

Observability Chart Configuration

observability.grafana.enabled is only the on/off switch (it's the condition on the grafana dependency in Chart.yaml); actual subchart configuration - persistence, service type, resources, datasources - must go under the top-level grafana: key instead, matching the dependency name:

observability:
  grafana:
    enabled: true

# Basic Grafana setup
grafana:
  service:
    type: LoadBalancer

# Production Grafana configuration
grafana:
  persistence:
    enabled: true
    size: 10Gi
  resources:
    limits:
      cpu: 200m
      memory: 400Mi
    requests:
      cpu: 50m
      memory: 200Mi

PostgreSQL Monitoring

Enable PostgreSQL metrics collection:

postgrescluster:
  monitoring: true  # Enables postgres_exporter sidecar

Available Metrics

Core Infrastructure Metrics

  • Container resources: CPU, memory, network usage
  • Kubernetes state: Pods, services, deployments status
  • Node metrics: Hardware utilization, filesystem usage
  • PostgreSQL: Database connections, query performance (when enabled)

Custom Application Metrics

When prometheus-adapter and nginx ingress are both enabled, these custom metrics become available: - nginx_ingress_controller_requests_rate_stac_eoapi - nginx_ingress_controller_requests_rate_raster_eoapi - nginx_ingress_controller_requests_rate_vector_eoapi - nginx_ingress_controller_requests_rate_multidim_eoapi

Requirements: - nginx ingress controller with prometheus metrics enabled - Ingress must use specific hostnames (not wildcard patterns) - prometheus-adapter must be configured to expose these metrics

Application Metrics

The raster (titiler-pgstac), vector (tipg), and stac (stac-fastapi-pgstac) services can natively expose a Prometheus metrics endpoint with request-count and latency histograms broken down by a low-cardinality operation label (e.g. tiles, search, list_items). This is opt-in per service and requires monitoring.prometheus.enabled=true for Prometheus to discover and scrape the endpoint:

monitoring:
  prometheus:
    enabled: true

raster:
  metrics:
    enabled: true  # exposes titiler_pgstac_http_requests_total, titiler_pgstac_http_request_duration_seconds at /metrics

vector:
  metrics:
    enabled: true  # exposes tipg_http_requests_total, tipg_http_request_duration_seconds at /metrics

stac:
  metrics:
    enabled: true  # exposes generic http_requests_total, http_request_duration_seconds at /_mgmt/metrics

stac's metrics require stac-fastapi-pgstac>=7.0.0 (the chart's pinned tag); unlike raster/vector, its metric names are unprefixed, matching stac-fastapi-api's shared implementation.

stac-auth-proxy (when stac-auth-proxy.enabled=true) exposes its own /_mgmt/metrics endpoint by default — no toggle needed — with generic http_requests_total / http_request_duration_seconds metrics. It's scraped via a dedicated prometheus.extraScrapeConfigs entry rather than pod annotations, since the vendored stac-auth-proxy subchart doesn't yet expose a prometheus.io/scrape annotation hook.

Prometheus Operator (ServiceMonitor)

Clusters running prometheus-operator (rather than, or alongside, this chart's bundled monitoring.prometheus) don't discover targets via prometheus.io/scrape annotations — their Prometheus custom resource only picks up ServiceMonitor objects matching its serviceMonitorSelector. Each service above has a matching metrics.serviceMonitor toggle (stac-auth-proxy.serviceMonitor for the proxy) that creates one, gated on the ServiceMonitor CRD actually being installed:

raster:
  metrics:
    enabled: true
    serviceMonitor:
      enabled: true
      interval: 30s
      # Match your Prometheus CR's serviceMonitorSelector, e.g. for kube-prometheus-stack:
      additionalLabels:
        release: prometheus

Without a matching label, most prometheus-operator installs will silently ignore the ServiceMonitor — check your Prometheus CR's spec.serviceMonitorSelector for the expected label. This is independent of monitoring.prometheus.enabled: you can use either mechanism, or both.

Pre-built Dashboards

The eoapi-observability chart provides ready-to-use dashboards:

eoAPI Services Dashboards

Per-service dashboards (raster, vector, stac, stac-auth-proxy) backed by the application metrics above, when enabled: - Request rates per service - Response times and error rates - Traffic patterns by operation

Infrastructure Dashboard

  • CPU usage rate by pod
  • CPU throttling metrics
  • Memory usage and limits
  • Pod count tracking

Container Resources Dashboard

  • Resource consumption by container
  • Resource quotas and limits
  • Performance bottlenecks

PostgreSQL Dashboard (when enabled)

  • Database connections
  • Query performance
  • Storage utilization

Production Configuration

monitoring:
  prometheus:
    enabled: true

prometheus:
  server:
    # Persistent storage
    persistentVolume:
      enabled: true
      size: 100Gi
      storageClass: "gp3"
    # Retention policy
    retention: "30d"
    # Resource allocation
    resources:
      limits:
        cpu: "2000m"
        memory: "4096Mi"
      requests:
        cpu: "1000m"
        memory: "2048Mi"
    # Security - internal access only
    service:
      type: ClusterIP

Resource Requirements

Core Monitoring Components

Minimum resource requirements (actual usage varies by cluster size and metrics volume):

Component CPU Memory Purpose
prometheus-server 500m 1Gi Metrics storage
metrics-server 100m 200Mi Resource metrics
kube-state-metrics 50m 150Mi K8s state
prometheus-node-exporter 50m 50Mi Node metrics
prometheus-adapter 100m 128Mi Custom metrics API
Total ~800m ~1.5Gi

Observability Components

Component CPU Memory Purpose
grafana 100m 200Mi Visualization

Operations

Accessing Grafana

# Get Grafana admin username (usually 'admin')
kubectl get secret -n eoapi -l app.kubernetes.io/name=grafana \
  -o jsonpath="{.items[0].data.admin-user}" | base64 --decode

# Get Grafana admin password
kubectl get secret -n eoapi -l app.kubernetes.io/name=grafana \
  -o jsonpath="{.items[0].data.admin-password}" | base64 --decode

# Port-forward to access Grafana UI
kubectl port-forward -n eoapi svc/eoapi-obs-grafana 3000:80
# Access at http://localhost:3000

Verification Commands

# Check Prometheus is running
kubectl get pods -n eoapi -l app.kubernetes.io/name=prometheus

# Verify metrics-server
kubectl get apiservice v1beta1.metrics.k8s.io

# List available custom metrics
kubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1" | jq '.resources[].name'

# Test metrics collection
kubectl port-forward svc/eoapi-prometheus-server 9090:80 -n eoapi
# Visit http://localhost:9090/targets

Monitoring Health

# Check Prometheus targets
curl -X GET 'http://localhost:9090/api/v1/query?query=up'

# Verify Grafana datasource connectivity
kubectl exec -it deployment/eoapi-obs-grafana -n eoapi -- \
  wget -O- http://eoapi-prometheus-server/api/v1/label/__name__/values

Advanced Features

Alerting Setup

Enable alertmanager for alert management:

monitoring:
  prometheus:
    enabled: true

prometheus:
  alertmanager:
    enabled: true
    config:
      global:
        # Configure with your SMTP server details
        smtp_smarthost: 'your-smtp-server:587'
        smtp_from: 'alertmanager@yourdomain.com'
      route:
        receiver: 'default-receiver'
      receivers:
      - name: 'default-receiver'
        webhook_configs:
        - url: 'http://your-webhook-endpoint:5001/'

Note: Replace example values with your actual SMTP server and webhook endpoints.

Batch Job Metrics

Enable pushgateway for batch job metrics:

monitoring:
  prometheus:
    enabled: true

prometheus:
  prometheus-pushgateway:
    enabled: true  # For batch job metrics collection

Custom Dashboards

Add custom dashboards by creating ConfigMaps with the appropriate label:

apiVersion: v1
kind: ConfigMap
metadata:
  name: custom-dashboard
  namespace: eoapi
  labels:
    eoapi_dashboard: "1"
data:
  custom.json: |
    {
      "dashboard": {
        "id": null,
        "title": "Custom eoAPI Dashboard",
        "tags": ["eoapi"],
        "panels": []
      }
    }

The ConfigMap must be in the same namespace as the Grafana deployment and include the eoapi_dashboard: "1" label.

Troubleshooting

Common Issues

Missing Metrics 1. Check Prometheus service discovery:

kubectl port-forward svc/eoapi-prometheus-server 9090:80 -n eoapi
# Visit http://localhost:9090/service-discovery

  1. Verify target endpoints:
    kubectl get endpoints -n eoapi
    

Grafana Connection Issues 1. Check datasource connectivity in Grafana UI → Configuration → Data Sources 2. Verify Prometheus URL accessibility from Grafana pod

Resource Issues - Monitor current usage: kubectl top pods -n eoapi - Check for OOMKilled containers: kubectl describe pods -n eoapi | grep -A 5 "Last State" - Verify resource limits are appropriate for your workload size - Consider increasing Prometheus retention settings if storage is full

Security Considerations

  • Network Security: Use ClusterIP services for Prometheus in production
  • Access Control: Configure network policies to restrict metrics access
  • Authentication: Enable authentication for Grafana (LDAP, OAuth, etc.)
  • Data Privacy: Consider metrics data sensitivity and retention policies