OpenTelemetry Collector is a vendor-neutral agent that receives, processes, and exports telemetry data. It sits between your applications and observability backends, decoupling instrumentation from storage. This tutorial walks the configuration model, then the mechanics, then failure modes.
The Three-Stage Pipeline Model
Every Collector config routes data through three stages:
- Receivers — accept telemetry (metrics, traces, logs) from applications or infrastructure
- Processors — transform, filter, enrich, or sample the data
- Exporters — send processed data to backends (Prometheus, Jaeger, Datadog, etc.)
Data flows: Receiver → Processor → Exporter. A pipeline connects them by name. You can have multiple receivers feeding one processor, or one receiver fanning out to multiple exporters. This is not a linear chain; it's a directed graph.
Gotcha: A receiver or exporter that's defined but not referenced in any pipeline is ignored. It won't error; it just sits idle.
Minimal Working Config
Here's a production-ready skeleton:
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
send_batch_size: 1024
timeout: 10s
exporters:
otlp:
endpoint: backend.example.com:4317
headers:
Authorization: Bearer token123
service:
pipelines:
traces:
receivers: [otlp]
processors: [batch]
exporters: [otlp]
This receives OpenTelemetry Protocol (OTLP) data on ports 4317 (gRPC) and 4318 (HTTP), batches it to reduce export overhead, then ships it to a backend. The service.pipelines section is mandatory — it declares which components are active.
Gotcha: The OTLP exporter endpoint must match your backend's ingestion port. Datadog, New Relic, and self-hosted Jaeger all use different ports or protocols.
Receivers: Where Data Enters
Common receivers:
- otlp — OpenTelemetry Protocol (gRPC and HTTP). Standard choice for instrumented apps.
- prometheus — scrape Prometheus targets. Use this to ingest existing metrics.
- jaeger — accept Jaeger spans. Useful for migrating from Jaeger agents.
- syslog — parse syslog messages. Handles structured and unstructured logs.
- hostmetrics — collect system CPU, memory, disk from the Collector host.
Example: scrape Prometheus targets and ingest OTLP traces:
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
prometheus:
config:
scrape_configs:
- job_name: 'app'
static_configs:
- targets: ['localhost:8080']
Each receiver can be configured independently. The prometheus receiver runs its own scrape loop; the otlp receiver listens passively.
Gotcha: Receiver names are unique. If you define two otlp receivers, the second overwrites the first. Use different receiver types or export multiple receivers in separate pipelines if you need multiple ports.
Processors: The Transformation Layer
Processors run in order and can drop, modify, or enrich data.
Essential processors:
- batch — group spans/metrics/logs before export. Reduces network calls. Always use this in production.
- memory_limiter — drop data if Collector memory exceeds a threshold. Prevents OOM crashes.
- attributes — add or remove attributes (key-value pairs) from all spans.
- resource_detection — auto-detect cloud provider, container, or host metadata.
- sampling — keep only a fraction of traces (probabilistic or tail-based).
Example with memory limit and sampling:
processors:
memory_limiter:
check_interval: 1s
limit_mib: 512
spike_limit_mib: 128
sampling:
sampling_percentage: 10
batch:
send_batch_size: 1024
timeout: 10s
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, sampling, batch]
exporters: [otlp]
Processor order matters. The memory_limiter runs first and can drop data if the spike exceeds 128 MiB. Then sampling keeps 10% of remaining spans. Finally, batch groups them.
Gotcha: memory_limiter is aggressive. If your Collector is under sustained load and hits the spike limit, it will discard data. Monitor Collector metrics (otelcol_processor_memory_limiter_*) to tune the thresholds.
Exporters: Where Data Leaves
Exporters push data to backends. Each exporter can have its own endpoint, authentication, and retry logic.
Common exporters:
- otlp — OpenTelemetry Protocol (gRPC or HTTP).
- prometheus — expose metrics on an HTTP endpoint (for scraping).
- jaeger — send traces to Jaeger (gRPC or HTTP).
- logging — print to stdout (debugging only).
- datadog — send to Datadog API.
Example with retry and timeout:
exporters:
otlp:
endpoint: backend.example.com:4317
timeout: 10s
retry_on_failure:
enabled: true
initial_interval: 100ms
max_interval: 10s
max_elapsed_time: 5m
The exporter will retry with exponential backoff up to 5 minutes before giving up. This is sensible for transient network hiccups.
Gotcha: If the backend is unavailable, the Collector buffers data in memory (or disk, if persistent storage is enabled). Without bounds, this can cause OOM. Pair exporters with memory_limiter and set a max_elapsed_time on retries.
Pipelines: Connecting the Pieces
Pipelines are declared in service.pipelines. Each pipeline has a signal type (traces, metrics, logs) and lists receivers, processors, and exporters.
service:
pipelines:
traces:
receivers: [otlp, jaeger]
processors: [memory_limiter, sampling, batch]
exporters: [otlp, logging]
metrics:
receivers: [prometheus]
processors: [batch]
exporters: [prometheus]
logs:
receivers: [syslog]
processors: [batch]
exporters: [otlp]
This config:
- Receives traces from OTLP and Jaeger, processes them, and exports to both the OTLP backend and stdout.
- Scrapes Prometheus metrics, batches them, and re-exports as a Prometheus endpoint.
- Receives syslog, batches it, and exports as OTLP logs.
Gotcha: Each signal type (traces, metrics, logs) has its own pipeline. A processor in the traces pipeline does not affect metrics. If you need the same logic for all signals, define the processor separately and reference it in each pipeline.
When This Breaks
Receiver port already in use: The Collector fails to start. Check netstat -tlnp | grep 4317 and kill the conflicting process, or change the receiver endpoint.
Exporter timeout: Data queues up in memory. Monitor otelcol_exporter_queue_size and increase timeout or max_elapsed_time, or scale the backend.
Processor drops data silently: The memory_limiter or sampling processor is too aggressive. Lower limit_mib thresholds or increase sampling_percentage.
Config syntax error: The Collector exits on startup. Validate YAML with yamllint and check the Collector logs for parsing errors.
Undefined receiver/exporter in pipeline: The Collector ignores the reference and may drop data. Verify all names in service.pipelines are defined in receivers, processors, and exporters.
Production Checklist
- Enable
memory_limiterprocessor with conservative thresholds. - Set
batchprocessorsend_batch_sizeandtimeoutbased on backend capacity. - Configure exporter retry and timeout.
- Monitor Collector metrics: queue size, memory, dropped spans.
- Run Collector as a sidecar or DaemonSet, not on the same host as the backend. If you're self-hosting on a cloud VM, the comparison on tinjauhost.biz.id covers flexible VPS options worth considering.
- Use TLS/mTLS for exporter endpoints.
- Test failover: kill the backend and verify the Collector doesn't crash.
Takeaway
OpenTelemetry Collector configuration is a three-stage pipeline: receivers ingest, processors transform, exporters ship. Start with the minimal config (OTLP receiver, batch processor, OTLP exporter), add memory_limiter and sampling, then tune based on metrics. If you prefer isolating the Collector in a reproducible local setup, devcontainer workflows make it straightforward to version-control your environment. The model is simple; the gotchas are resource management and misconfigured pipelines.