Distributed Tracing in Go with OpenTelemetry: Spans, Context, and Sampling That Works

Your microservice takes eleven seconds to answer a request that used to take two hundred milliseconds. The on-call engineer opens the logs and finds nothing wrong: every service logged “success.” The problem is invisible because each service only knows its own story, and the story nobody is telling is the one that matters — which of the eight services in the request path added nine seconds, and why.

This is the gap distributed tracing fills. A trace is the complete journey of one request through your system: a tree of timed spans, stitched together across process boundaries by two small HTTP headers. In this post we build tracing into a Go service with OpenTelemetry — provider setup, automatic HTTP instrumentation, manual spans, context propagation, and the sampling strategy that keeps it affordable in production.

The Two Headers That Bind a Trace Together

Every span in a trace carries a trace ID, its own span ID, and the span ID of its parent. When a request crosses a network boundary, the caller serializes that identity into headers and the callee parses it back out. The W3C Trace Context standard defines the wire format: a traceparent header like 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01, encoding version, trace ID, span ID, and sampled flag. A second header, tracestate, carries vendor-specific context.

That is the entire magic. There is no central coordinator assigning IDs; the first service to see a request generates a trace ID and everyone downstream inherits it. This is why you can add tracing to services incrementally — a newly instrumented service joins existing traces automatically, and un-instrumented hops just pass the headers through untouched.

Wiring Up the Tracer Provider

The OpenTelemetry Go SDK splits into a stable API (the go.opentelemetry.io/otel packages you code against) and an SDK (the implementation you configure at startup). The wiring happens once, in your main function: build an exporter, build a resource describing your service, wrap both in a TracerProvider, and register it globally.

package main

import (
	"context"
	"log"
	"os"
	"os/signal"
	"syscall"
	"time"

	"go.opentelemetry.io/otel"
	"go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp"
	"go.opentelemetry.io/otel/propagation"
	"go.opentelemetry.io/otel/sdk/resource"
	sdktrace "go.opentelemetry.io/otel/sdk/trace"
	semconv "go.opentelemetry.io/otel/semconv/v1.41.0"
)

func initTracer(ctx context.Context) (func(context.Context) error, error) {
	exporter, err := otlptracehttp.New(ctx,
		otlptracehttp.WithEndpoint("localhost:4318"),
		otlptracehttp.WithInsecure(),
	)
	if err != nil {
		return nil, err
	}

	res, err := resource.New(ctx,
		resource.WithAttributes(
			semconv.ServiceName("checkout-api"),
			semconv.ServiceVersion("1.6.2"),
		),
	)
	if err != nil {
		return nil, err
	}

	tp := sdktrace.NewTracerProvider(
		sdktrace.WithBatcher(exporter),
		sdktrace.WithResource(res),
		sdktrace.WithSampler(sdktrace.ParentBased(sdktrace.TraceIDRatioBased(0.1))),
	)

	otel.SetTracerProvider(tp)
	otel.SetTextMapPropagator(
		propagation.NewCompositeTextMapPropagator(
			propagation.TraceContext{},
			propagation.Baggage{},
		),
	)

	return tp.Shutdown, nil
}

func main() {
	ctx := context.Background()

	shutdown, err := initTracer(ctx)
	if err != nil {
		log.Fatal(err)
	}

	stop := make(chan os.Signal, 1)
	signal.Notify(stop, os.Interrupt, syscall.SIGTERM)
	<-stop

	shutdownCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
	defer cancel()
	if err := shutdown(shutdownCtx); err != nil {
		log.Printf("tracer shutdown: %v", err)
	}
}

Three details here matter more than they look. The resource is the identity card attached to every span your service emits — the service.name attribute is what your tracing backend groups traces by, and setting it wrong (or leaving the default unknown_service) makes multi-service views useless. The propagator list decides which wire formats you speak; TraceContext{} is the W3C standard and should almost always come first, with Baggage{} added if you pass business context like tenant IDs alongside traces. And shutdown must run on a clean exit: the SDK batches spans in memory and exports asynchronously, so killing the process without flushing silently drops the last several seconds of spans — precisely the ones around the crash you are debugging.

The exporter above speaks OTLP over HTTP to a collector on localhost:4318, the default port. In production, a lightweight OpenTelemetry Collector usually runs as a sidecar or DaemonSet; the SDK exports to it cheaply over localhost, and the collector handles batching, retry, sensitive-data filtering, and forwarding to whatever backend you use — Jaeger, a commercial SaaS, or both at once. This one indirection means changing tracing vendors becomes a collector config change, not a redeploy of every service.

Free Spans: Auto-Instrumenting HTTP Servers and Clients

Before writing any manual instrumentation, let the otelhttp package do the boring part. On the server side, one wrapper around your mux creates a server span per request, extracts the incoming traceparent, names the span after the operation, and records the HTTP status.

mux := http.NewServeMux()
mux.HandleFunc("POST /orders", createOrder)

// One line: every request now gets a server span,
// parented to whatever traceparent header arrived.
wrapped := otelhttp.NewHandler(mux, "checkout-api")
http.ListenAndServe(":8080", wrapped)

On the client side, wrap the transport. This is the half people forget, and it is the half that makes traces continuous: a wrapped http.RoundTripper injects the traceparent header into every outbound request, so the downstream service's server span automatically becomes a child of yours.

client := &http.Client{
	Transport: otelhttp.NewTransport(http.DefaultTransport),
	Timeout:   5 * time.Second,
}

// Pass the inbound request's context — that is what links the spans.
req, _ := http.NewRequestWithContext(r.Context(), http.MethodPost,
	"http://inventory.internal/reserve", body)
resp, err := client.Do(req)

Two failure modes account for most "why is there a gap in my trace" confusion. First, forgetting NewRequestWithContext: if you build the request without the inbound context, the client transport has no span context to inject, and the downstream call starts a brand-new trace. Second, constructing a new http.Client deep in some helper without the wrapped transport — the request goes out bare, and your beautiful waterfall turns into two disconnected islands.

Manual Spans: Instrumenting What Actually Matters

HTTP spans tell you which service is slow. Manual spans tell you why. Any meaningful unit of work — a database transaction, a call to a payment provider, a cache-priming loop — deserves a child span with attributes that make the waterfall self-explanatory.

var tracer = otel.Tracer("checkout-api/orders")

func placeOrder(ctx context.Context, order Order) error {
	ctx, span := tracer.Start(ctx, "orders.place",
		trace.WithSpanKind(trace.SpanKindInternal),
		trace.WithAttributes(
			attribute.String("order.id", order.ID),
			attribute.Int("order.item_count", len(order.Items)),
		),
	)
	defer span.End()

	if err := chargeCard(ctx, order); err != nil {
		span.RecordError(err)
		span.SetStatus(codes.Error, "payment failed")
		return fmt.Errorf("charge card: %w", err)
	}

	span.SetStatus(codes.Ok, "")
	return nil
}

The context dance is the part to internalize. tracer.Start returns a new context that carries the new span; every downstream call in this function must receive that returned ctx, not the original, or the child spans you expect will become siblings instead. Note also the status convention: an error log line is not enough. A span only shows as failed in the backend if you call SetStatus(codes.Error, ...), and the exception only appears if you RecordError it. Spans that carry neither look successful no matter what happened inside them — the tracing equivalent of logging "success."

Attribute choice deserves as much care as span choice. High-cardinality values like user IDs or request IDs as span attributes are fine, because spans are sampled individually — the cardinality discipline that matters for metrics does not apply here. But attribute strategy is a one-way door in the other direction: backends only let you filter and group by what was recorded, so put the things you will slice traces by — order channel, payment provider, retry count — into attributes now, not after the first bad incident.

Sampling: Paying for What You Keep

At any real traffic level, recording every trace is expensive and mostly pointless — the overwhelming majority of requests behave exactly as designed. The SDK's sampling decides which traces survive. The default sampler keeps everything, which means an unconfigured service at 5,000 requests per second emits 5,000 traces per second forever.

The setup shown earlier uses the configuration most teams should start with:

sdktrace.WithSampler(
	sdktrace.ParentBased(sdktrace.TraceIDRatioBased(0.1)),
)

TraceIDRatioBased(0.1) keeps ten percent of the traces that reach this service as the head of the chain. Wrapping it in ParentBased is the critical subtlety: sampling decisions happen once, at the first service, and the decision travels in the traceparent header. Without ParentBased, each downstream service independently drops ten percent of its spans, and you get Frankenstein traces with random segments missing. With it, a service honors the upstream decision and keeps its spans for the whole trace or none of it.

Head sampling has a known weakness: the interesting traces — the failures, the p99 outliers — do not identify themselves until after they finish. The answer is tail-based sampling, configured at the Collector: keep traces in memory, then retain all error traces and slow traces at 100%, plus a small baseline of the healthy ones. That gives you complete visibility into everything that matters at a fraction of full-retention cost, at the price of the Collector holding traces in memory and the operational maturity to run it well.

Reading a Waterfall Like an Engineer

A trace view is a tree drawn as a timeline, and it answers questions in a specific order. Start with span width: the widest span is where the wall-clock time went. Then check whether its children fill that width — if a database call spans 800ms but its children account for 60ms, the time went into connection acquisition, pool queuing, or network transfer, none of which produce child spans on their own. Then look for serial chains: three 200ms service calls back-to-back that could run concurrently is a 400ms saving hiding in plain sight. Finally, look for the same downstream service appearing many times under one parent — the classic N+1 pattern, where one span per iteration means one network round trip per iteration.

Traces also make error paths legible in a way logs cannot. When a request fails, the failed span carries the error status and recorded exception, and the waterfall shows exactly which downstream dependency returned what. Correlate the trace ID from the response header back into your structured logs (put the trace ID in every log line — Go's slog handlers make this a one-line middleware change) and a user complaint becomes a reproducible, navigable artifact instead of a guessing game.

Where to Go From Here

The full path from zero is smaller than it looks: run a Collector and a Jaeger instance in your dev environment, wire the provider block in your main, wrap your mux and your outbound client, and add manual spans to the three slowest operations you already know about. The official Go getting-started guide walks the same setup end to end, and the SDK docs cover the sampler and processor options in detail.

Once the first service is live, resist the urge to instrument everything at once. Add tracing to the one service on your slowest user journey, watch a hundred traces, and let the waterfalls tell you where to go next. Distributed debugging stops being a superpower the moment every request hands its story forward in two headers.

Leave a Reply

Your email address will not be published. Required fields are marked *