OpenTelemetry Collector Configuration Tutorial

by Liam Foster
OpenTelemetry Collector Configuration Tutorial

OpenTelemetry Collector is a vendor-neutral agent that receives, processes, and exports telemetry data. It sits between your applications and observability backends, decoupling instrumentation from storage. This tutorial walks the configuration model, then the mechanics, then failure modes.

The Three-Stage Pipeline Model

Every Collector config routes data through three stages:

  1. Receivers — accept telemetry (metrics, traces, logs) from applications or infrastructure
  2. Processors — transform, filter, enrich, or sample the data
  3. Exporters — send processed data to backends (Prometheus, Jaeger, Datadog, etc.)

Data flows: Receiver → Processor → Exporter. A pipeline connects them by name. You can have multiple receivers feeding one processor, or one receiver fanning out to multiple exporters. This is not a linear chain; it's a directed graph.

Gotcha: A receiver or exporter that's defined but not referenced in any pipeline is ignored. It won't error; it just sits idle.

Minimal Working Config

Here's a production-ready skeleton:

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  batch:
    send_batch_size: 1024
    timeout: 10s

exporters:
  otlp:
    endpoint: backend.example.com:4317
    headers:
      Authorization: Bearer token123

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [batch]
      exporters: [otlp]

This receives OpenTelemetry Protocol (OTLP) data on ports 4317 (gRPC) and 4318 (HTTP), batches it to reduce export overhead, then ships it to a backend. The service.pipelines section is mandatory — it declares which components are active.

Gotcha: The OTLP exporter endpoint must match your backend's ingestion port. Datadog, New Relic, and self-hosted Jaeger all use different ports or protocols.

Receivers: Where Data Enters

Common receivers:

  • otlp — OpenTelemetry Protocol (gRPC and HTTP). Standard choice for instrumented apps.
  • prometheus — scrape Prometheus targets. Use this to ingest existing metrics.
  • jaeger — accept Jaeger spans. Useful for migrating from Jaeger agents.
  • syslog — parse syslog messages. Handles structured and unstructured logs.
  • hostmetrics — collect system CPU, memory, disk from the Collector host.

Example: scrape Prometheus targets and ingest OTLP traces:

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
  prometheus:
    config:
      scrape_configs:
        - job_name: 'app'
          static_configs:
            - targets: ['localhost:8080']

Each receiver can be configured independently. The prometheus receiver runs its own scrape loop; the otlp receiver listens passively.

Gotcha: Receiver names are unique. If you define two otlp receivers, the second overwrites the first. Use different receiver types or export multiple receivers in separate pipelines if you need multiple ports.

Processors: The Transformation Layer

Processors run in order and can drop, modify, or enrich data.

Essential processors:

  • batch — group spans/metrics/logs before export. Reduces network calls. Always use this in production.
  • memory_limiter — drop data if Collector memory exceeds a threshold. Prevents OOM crashes.
  • attributes — add or remove attributes (key-value pairs) from all spans.
  • resource_detection — auto-detect cloud provider, container, or host metadata.
  • sampling — keep only a fraction of traces (probabilistic or tail-based).

Example with memory limit and sampling:

processors:
  memory_limiter:
    check_interval: 1s
    limit_mib: 512
    spike_limit_mib: 128
  sampling:
    sampling_percentage: 10
  batch:
    send_batch_size: 1024
    timeout: 10s

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, sampling, batch]
      exporters: [otlp]

Processor order matters. The memory_limiter runs first and can drop data if the spike exceeds 128 MiB. Then sampling keeps 10% of remaining spans. Finally, batch groups them.

Gotcha: memory_limiter is aggressive. If your Collector is under sustained load and hits the spike limit, it will discard data. Monitor Collector metrics (otelcol_processor_memory_limiter_*) to tune the thresholds.

Exporters: Where Data Leaves

Exporters push data to backends. Each exporter can have its own endpoint, authentication, and retry logic.

Common exporters:

  • otlp — OpenTelemetry Protocol (gRPC or HTTP).
  • prometheus — expose metrics on an HTTP endpoint (for scraping).
  • jaeger — send traces to Jaeger (gRPC or HTTP).
  • logging — print to stdout (debugging only).
  • datadog — send to Datadog API.

Example with retry and timeout:

exporters:
  otlp:
    endpoint: backend.example.com:4317
    timeout: 10s
    retry_on_failure:
      enabled: true
      initial_interval: 100ms
      max_interval: 10s
      max_elapsed_time: 5m

The exporter will retry with exponential backoff up to 5 minutes before giving up. This is sensible for transient network hiccups.

Gotcha: If the backend is unavailable, the Collector buffers data in memory (or disk, if persistent storage is enabled). Without bounds, this can cause OOM. Pair exporters with memory_limiter and set a max_elapsed_time on retries.

Pipelines: Connecting the Pieces

Pipelines are declared in service.pipelines. Each pipeline has a signal type (traces, metrics, logs) and lists receivers, processors, and exporters.

service:
  pipelines:
    traces:
      receivers: [otlp, jaeger]
      processors: [memory_limiter, sampling, batch]
      exporters: [otlp, logging]
    metrics:
      receivers: [prometheus]
      processors: [batch]
      exporters: [prometheus]
    logs:
      receivers: [syslog]
      processors: [batch]
      exporters: [otlp]

This config:

  • Receives traces from OTLP and Jaeger, processes them, and exports to both the OTLP backend and stdout.
  • Scrapes Prometheus metrics, batches them, and re-exports as a Prometheus endpoint.
  • Receives syslog, batches it, and exports as OTLP logs.

Gotcha: Each signal type (traces, metrics, logs) has its own pipeline. A processor in the traces pipeline does not affect metrics. If you need the same logic for all signals, define the processor separately and reference it in each pipeline.

When This Breaks

Receiver port already in use: The Collector fails to start. Check netstat -tlnp | grep 4317 and kill the conflicting process, or change the receiver endpoint.

Exporter timeout: Data queues up in memory. Monitor otelcol_exporter_queue_size and increase timeout or max_elapsed_time, or scale the backend.

Processor drops data silently: The memory_limiter or sampling processor is too aggressive. Lower limit_mib thresholds or increase sampling_percentage.

Config syntax error: The Collector exits on startup. Validate YAML with yamllint and check the Collector logs for parsing errors.

Undefined receiver/exporter in pipeline: The Collector ignores the reference and may drop data. Verify all names in service.pipelines are defined in receivers, processors, and exporters.

Production Checklist

  • Enable memory_limiter processor with conservative thresholds.
  • Set batch processor send_batch_size and timeout based on backend capacity.
  • Configure exporter retry and timeout.
  • Monitor Collector metrics: queue size, memory, dropped spans.
  • Run Collector as a sidecar or DaemonSet, not on the same host as the backend. If you're self-hosting on a cloud VM, the comparison on tinjauhost.biz.id covers flexible VPS options worth considering.
  • Use TLS/mTLS for exporter endpoints.
  • Test failover: kill the backend and verify the Collector doesn't crash.

Takeaway

OpenTelemetry Collector configuration is a three-stage pipeline: receivers ingest, processors transform, exporters ship. Start with the minimal config (OTLP receiver, batch processor, OTLP exporter), add memory_limiter and sampling, then tune based on metrics. If you prefer isolating the Collector in a reproducible local setup, devcontainer workflows make it straightforward to version-control your environment. The model is simple; the gotchas are resource management and misconfigured pipelines.