A practical guide to distributed tracing with Ion. Learn how to create spans, propagate context, and debug production issues.
A Span represents a single unit of work in your application. Think of it as a "timer with context."
[Request]────────────────────────────────────────────────────────────>
[ProcessOrder]────────────────────────>
[ValidatePayment]───────>
[WriteDB]───>
Each box above is a span. Together, they form a Trace — the complete picture of a request's journey through your system.
Ion provides a clean Span interface that wraps OpenTelemetry:
type Span interface {
// End marks the span as complete. MUST be called.
End()
// SetStatus sets the span status (StatusOK, StatusError, or StatusUnset).
SetStatus(code ion.StatusCode, description string)
// RecordError records an error as an event on the span timeline.
RecordError(err error)
// SetAttributes adds key-value metadata for filtering/searching.
SetAttributes(attrs ...ion.Attr)
// AddEvent adds a timestamped event to the span.
AddEvent(name string, attrs ...ion.Attr)
}Every span follows this pattern: Start → Work → End
import "go.opentelemetry.io/otel/attribute"
func ProcessOrder(ctx context.Context, orderID string) error {
// 1. Get a named tracer (use Child() for full observability)
tracer := app.Tracer("order.processor")
// 2. Start a span (returns enriched context + span)
ctx, span := tracer.Start(ctx, "ProcessOrder")
// 3. ALWAYS defer End() immediately after Start()
defer span.End()
// 4. Add attributes for filtering in Jaeger/Tempo
span.SetAttributes(attribute.String("order.id", orderID))
// 5. Do work...
if err := validateOrder(ctx, orderID); err != nil {
// 6. Record errors properly
span.RecordError(err)
span.SetStatus(ion.StatusError, "validation failed")
return err
}
// 7. Mark success (optional, default is Unset/Ok)
span.SetStatus(ion.StatusOK, "order processed")
return nil
}Note
Ion provides ion.StatusOK, ion.StatusError, and ion.StatusUnset so you never need to import go.opentelemetry.io/otel/codes. For span attributes, use attribute.String(), attribute.Int64(), etc. from the standard OTel attribute package.
Marks the span as complete and records its duration.
ctx, span := tracer.Start(ctx, "MyOperation")
defer span.End() // ← CRITICAL: Always defer immediatelyCaution
Forgetting to call End() causes memory leaks and broken traces. Always use defer.
Sets the span's final status. This affects error rate metrics and trace visualization.
| Code | When to Use | Visual |
|---|---|---|
ion.StatusUnset |
Default, operation outcome unknown | Gray |
ion.StatusOK |
Explicit success | Green |
ion.StatusError |
Operation failed | Red |
// On success
span.SetStatus(ion.StatusOK, "completed")
// On failure
span.SetStatus(ion.StatusError, "database timeout")Important
By default, spans finish as Unset (success). You MUST call SetStatus(ion.StatusError, ...) on failures, or your error rate dashboard will show 0%.
Records an error as an event on the span timeline. This captures:
- Error message
- Error type
- Stack trace (if available)
if err != nil {
span.RecordError(err) // Adds "exception" event to timeline
span.SetStatus(ion.StatusError, "operation failed") // Marks span as failed
return err
}Tip
Always use both RecordError() AND SetStatus() together. RecordError adds detail; SetStatus triggers alerts.
Adds searchable metadata to the span. Use for low-cardinality data that helps you filter traces.
span.SetAttributes(
attribute.String("user.id", userID),
attribute.Int("retry.count", retryCount),
attribute.Bool("cache.hit", cacheHit),
)Good Attributes (Low Cardinality):
http.status_code(50 values)region(10 values)customer.tier(3 values)
Bad Attributes (High Cardinality):
error.message(infinite)request.body(huge)user.email(PII risk)
Adds a timestamped marker to the span timeline. Use for significant moments during execution.
// Mark when cache was missed
span.AddEvent("cache_miss", attribute.String("key", cacheKey))
// Mark retry attempts
span.AddEvent("retry_attempt", attribute.Int("attempt", 3))
// Mark important state changes
span.AddEvent("payment_authorized")Events appear as points on the span timeline in Jaeger/Tempo, helping you understand the sequence of operations.
Understanding when to create a new span versus adding an event is key to clean traces.
| Feature | tracer.Start(...) (New Span) |
span.AddEvent(...) |
|---|---|---|
| Concept | Duration (Start & End) | Point in Time (Snapshot) |
| Visual | A bar with length | A dot on the timeline |
| Best For | DB Queries, API Calls, Major Functions | Retries, Cache Misses, Validation Failures |
| Overhead | Higher (Context switching, allocs) | Very Low (Log entry) |
Rule of Thumb:
- "How long did this take?" → New Span
- "When did this happen?" → Add Event
RecordError(err) is just a specialized Event.
- It adds an event named
"exception"with the error message and stack trace. - Where to put it? Record the error on the active span where the failure occurred.
- Bubble Up: If a child span fails (e.g., DB call), record the error there. If that error causes the parent operation to fail, record it on the parent span as well.
When starting a span, you can customize its behavior:
Pre-populate attributes at span creation:
ctx, span := tracer.Start(ctx, "ProcessPayment",
ion.WithAttributes(
attribute.String("payment.method", "card"),
attribute.Float64("amount", 99.99),
),
)Set the span's role in the trace:
import "go.opentelemetry.io/otel/trace"
// For incoming requests (HTTP server, gRPC server)
tracer.Start(ctx, "HandleRequest", ion.WithSpanKind(trace.SpanKindServer))
// For outgoing requests (HTTP client, DB calls)
tracer.Start(ctx, "CallExternalAPI", ion.WithSpanKind(trace.SpanKindClient))
// For async message producers
tracer.Start(ctx, "PublishEvent", ion.WithSpanKind(trace.SpanKindProducer))
// For async message consumers
tracer.Start(ctx, "ConsumeEvent", ion.WithSpanKind(trace.SpanKindConsumer))
// Default: internal operations
tracer.Start(ctx, "Calculate", ion.WithSpanKind(trace.SpanKindInternal))Connect spans that are causally related but not in a parent-child relationship:
// Capture link from parent context
link := ion.LinkFromContext(parentCtx)
// Create new span with link to parent
ctx, span := tracer.Start(context.Background(), "BackgroundJob",
ion.WithLinks(link),
)func HandleRequest(ctx context.Context, req *Request) (*Response, error) {
tracer := app.Tracer("api.handler")
ctx, span := tracer.Start(ctx, "HandleRequest")
defer span.End()
// Add request metadata
span.SetAttributes(
attribute.String("http.method", req.Method),
attribute.String("http.path", req.Path),
)
// Process
result, err := process(ctx, req)
if err != nil {
span.RecordError(err)
span.SetStatus(ion.StatusError, err.Error())
return nil, err
}
span.SetStatus(ion.StatusOK, "success")
return result, nil
}func (r *Repository) GetUser(ctx context.Context, id string) (*User, error) {
tracer := app.Tracer("repository.user")
ctx, span := tracer.Start(ctx, "GetUser",
ion.WithSpanKind(trace.SpanKindClient),
)
defer span.End()
span.SetAttributes(attribute.String("db.operation", "SELECT"))
user, err := r.db.QueryRow(ctx, "SELECT * FROM users WHERE id = $1", id)
if err != nil {
span.RecordError(err)
span.SetStatus(ion.StatusError, "query failed")
return nil, err
}
return user, nil
}When spawning goroutines, use links to preserve trace causality:
func (s *Service) ProcessAsync(ctx context.Context, job *Job) {
// Capture causality before spawning
link := ion.LinkFromContext(ctx)
go func() {
// Create FRESH context (won't be canceled when parent returns)
newCtx := context.Background()
tracer := app.Tracer("worker.background")
// Link back to original request
ctx, span := tracer.Start(newCtx, "ProcessJob",
ion.WithLinks(link),
)
defer span.End()
// Process...
if err := s.doWork(ctx, job); err != nil {
span.RecordError(err)
span.SetStatus(ion.StatusError, "job failed")
}
}()
}For loops, don't trace every item — trace the batch:
func ProcessBatch(ctx context.Context, items []*Item) error {
tracer := app.Tracer("batch.processor")
ctx, span := tracer.Start(ctx, "ProcessBatch")
defer span.End()
span.SetAttributes(attribute.Int("batch.size", len(items)))
var errors int
for i, item := range items {
if err := processItem(ctx, item); err != nil {
errors++
// Add event instead of child span
span.AddEvent("item_failed",
attribute.Int("index", i),
attribute.String("error", err.Error()),
)
}
}
span.SetAttributes(attribute.Int("batch.errors", errors))
if errors > 0 {
span.SetStatus(ion.StatusError, fmt.Sprintf("%d items failed", errors))
return fmt.Errorf("%d items failed", errors)
}
return nil
}Warning
This is the #1 source of broken dashboards.
When an error occurs, you must do three things:
if err != nil {
// 1. Record the error details (adds event to timeline)
span.RecordError(err)
// 2. Mark the span as failed (turns red, affects error rate)
span.SetStatus(ion.StatusError, "operation failed")
// 3. Log for human debugging (correlated via trace_id)
app.Error(ctx, "operation failed", err)
return err
}| Action | What It Does | Where It Shows |
|---|---|---|
RecordError(err) |
Adds exception event with stack trace | Span timeline |
SetStatus(ion.StatusError, ...) |
Marks span as failed | Red in UI, Error Rate metrics |
app.Error(ctx, ...) |
Logs with trace correlation | Loki, searching by trace_id |
Once configured correctly, your traces will appear in Jaeger/Tempo like this:
ion-service 12.5ms
├── HandleRequest [http.method=POST, http.path=/orders] 12.0ms
│ ├── ValidateOrder 2.1ms
│ ├── ProcessPayment [payment.method=card] 5.2ms
│ │ └── CallPaymentGateway [SpanKindClient] 4.8ms
│ └── SaveOrder 3.1ms
│ └── PostgreSQL.INSERT [SpanKindClient] 2.9ms
Each span shows:
- Duration (right side)
- Attributes (in brackets)
- Span kind (for client/server relationships)
- Events (as points on the timeline)
- Errors (red highlighting)
Enable tracing in your Ion config:
cfg := ion.Default().
WithService("my-service").
WithTracing("localhost:4317") // OTEL Collector address
// Or with full control:
cfg := ion.Config{
ServiceName: "my-service",
Tracing: ion.TracingConfig{
Enabled: true,
Endpoint: "localhost:4317",
Protocol: "grpc",
Sampler: "always", // or "ratio:0.1" for 10%
},
}
app, _, err := ion.New(cfg)- Always
defer span.End()immediately aftertracer.Start() - Always call both
RecordError()andSetStatus(ion.StatusError, ...)on failures - Use low-cardinality attributes — avoid putting error messages or dynamic IDs as attribute values
- Pass context everywhere — broken context = broken traces
- Use Links for goroutines — prevents ghost spans and preserves causality
- Trace boundaries, not everything — DB calls, HTTP requests, major logic blocks
- Name spans descriptively —
ProcessOrdernotDoStuff - Call
app.Shutdown(ctx)on exit — flushes all buffered traces - Use
Child()for components —svc := app.Child("http"), thensvc.Tracer("http.handler")
// Get tracer (prefer Child() for component-scoped observability)
http := app.Child("http")
tracer := http.Tracer("http.handler")
// Start span
ctx, span := tracer.Start(ctx, "OperationName")
defer span.End()
// Add attributes
span.SetAttributes(attribute.String("key", "value"))
// Add event
span.AddEvent("event_name", attribute.Int("count", 5))
// Record error
span.RecordError(err)
span.SetStatus(ion.StatusError, "failed")
// Mark success
span.SetStatus(ion.StatusOK, "success")
// Link spans (for goroutines)
link := ion.LinkFromContext(parentCtx)
ctx, span := tracer.Start(context.Background(), "AsyncOp", ion.WithLinks(link))