- Resources
- What is Alloy? When does it make sense to use it?
- Alloy configuration language 101
- Learning environment setup
- Foundation 1: Traces
- Foundation 2: Logs
- Foundation 3: Metrics
- How to use the Alloy UI to debug pipelines
- Mission I: Rogue Dimension - cardinality filtering
- Mission II: Operation Cold Storage - log splitting & S3 archival
- Mission III: Selective Surveillance - head sampling
- Mission IV: Leave No (Error) Trace - tail sampling
To instruct Alloy on how we want that done, we must write these instructions in a language (Alloy syntax) that Alloy understands.
The usage section gives you an example of how this particular component can be configured.
The arguments and blocks sections list what you can do with the data. Pay close attention to the name, type, description, default, and required columns so Alloy knows what you want it to do!
Focusing on these 3 things will point us in the right direction as we configure our pipeline.
Important
Prerequisites
Make sure you have the following installed before continuing:
- Docker Desktop or Docker Engine
- Docker Compose (included with Docker Desktop)
Step 1: Clone the repo for the learning environment:
git clone https://github.com/grafana/grafanacon2026-alloy-in-action.gitStep 2: Start the environment from within the project's root directory:
make startYou should see the following message in the terminal:
✅ Mission Control is online
Health check: curl http://localhost:8080/health
Tip
To stop the environment at any time, run make stop from the project root.
- Open Grafana (localhost:3000). You should see the Grafana page:
- Open the Alloy UI (localhost:12347). You should see the Alloy UI:
If both pages load, you're good to go!
Warning
Setup not working?
- Docker not running?
If you're using Docker Desktop, start it (or restart it if it is already running) and try again. If you're using Docker Engine, make sure the Docker daemon is running (
sudo systemctl start docker). - Port conflicts? Make sure ports 3000, 12347, 4317, and 4318 are free.
Open the project using a text editor of your choice.
- Expand the
alloy/folder and open theconfig.alloyfile. - We will use this file to build pipelines for the training exercises and missions.
Note
Throughout this workshop, you'll edit alloy/config.alloy and then reload Alloy to apply your changes. Each section builds on the previous one, so keep your earlier work in the file as you add new pipelines.
- Use
otelcol.receiver.otlpto receive traces from the application - Use
otelcol.processor.batchto batch traces for efficient export - Use
otelcol.exporter.otlphttpto send traces to Tempo
Field agents are conducting operations around the world, and mission-control is tracking every request that flows through the system.
These operations generate traces, detailed records of each request's journey through the application.
Right now, those traces are being generated but going nowhere, since Alloy isn't configured to receive them.
Your task: Build a trace pipeline so you can see what's happening inside mission-control.
Once it's configured, you'll be able to view traces in the Mission Control Grafana dashboard and drill into individual operations.
Pipeline:
otelcol.receiver.otlp → otelcol.processor.batch → otelcol.exporter.otlphttp
- Receive OTLP traces on
0.0.0.0:4317(gRPC) and0.0.0.0:4318(HTTP) - Batch traces before sending (improves efficiency)
- Export to Tempo at
http://tempo:4318
Copy this into your config.alloy file and fill in the TODOs:
Tip
How the starter code works: Each exercise gives you a code block with TODO placeholders. Replace each TODO with the correct value. The comments next to each TODO give you a hint about what goes there.
/*
Foundation 1: Traces Pipeline
Pipeline: otelcol.receiver.otlp -> otelcol.processor.batch -> otelcol.exporter.otlphttp
*/
// Step 1: Receive OTLP traces from mission-control
//The receiver listens for incoming traces. Mission-control sends traces using OTLP, so you need both gRPC and HTTP endpoints which are 4317 and 4318 by default.
otelcol.receiver.otlp "default" {
grpc {
endpoint = "0.0.0.0:4317"
}
http {
endpoint = "0.0.0.0:4318"
}
// OpenTelemetry components use `output` to send data and `.input` to receive it:
output {
traces = [TODO] // Forward to the batch processor's input hint: component_type.label.input
}
}
// Step 2: Batch traces before exporting
otelcol.processor.batch "default" {
send_batch_size = 100
send_batch_max_size = 200
timeout = "250ms"
output {
traces = [TODO] // Forward to the Tempo exporter's input
}
}
// Step 3: Export traces to Tempo
otelcol.exporter.otlphttp "docker_tempo" {
client {
endpoint = "TODO" // Send to http://tempo:4318
tls {
insecure = true
insecure_skip_verify = true
}
}
}Full solution
/*
Foundation 1: Traces Pipeline
Pipeline: otelcol.receiver.otlp -> otelcol.processor.batch -> otelcol.exporter.otlphttp
*/
// Receive OTLP traces from mission-control
otelcol.receiver.otlp "default" {
grpc {
endpoint = "0.0.0.0:4317"
}
http {
endpoint = "0.0.0.0:4318"
}
output {
traces = [otelcol.processor.batch.default.input]
}
}
// Batch traces before exporting
otelcol.processor.batch "default" {
output {
traces = [otelcol.exporter.otlphttp.docker_tempo.input]
}
}
// Export traces to Tempo
otelcol.exporter.otlphttp "docker_tempo" {
client {
endpoint = "http://tempo:4318"
tls {
insecure = true
insecure_skip_verify = true
}
}
}Whenever you make changes to the config file, you need to reload Alloy:
make alloy-reloadIf the config is valid, you'll see:
config reloaded
Caution
If you see an error instead of config reloaded, double-check your config for typos, missing commas, or smart quotes. See the Troubleshooting section for common issues.
- Go to the Mission Control Overview dashboard (localhost:3000). The Traces panel at the bottom should now show data.
- When you click on one of the traces, you can see the full breakdown of what happened and how long it took.
- Use
loki.source.fileto discover and read log files - Use
loki.processto parse JSON and extract labels - Use
loki.writeto send logs to Loki
Mission-control generates operational logs: JSON-formatted records of everything happening in the system. Agent check-ins, status updates, and internal events.
These logs are written to files inside the container but no one's reading them, just like the previous section.
The environment is set up to mimic a scenario where Alloy runs on the host and has access to its filesystem (a DaemonSet in Kubernetes or simply an Alloy collector running on each VM).
Your task: Build a log pipeline to read mission-control's log files, parse the JSON, and send them to Loki.
Once it's configured, you'll see logs flowing into Grafana where you can search and filter by component.
Pipeline:
loki.source.file → loki.process → loki.write
- Discover and read log files at
/var/log/alloy/*.log - Parse JSON and extract the
componentfield and add it as a label - Write to Loki at
http://loki:3100/loki/api/v1/push
Add this to your config.alloy file (below your traces pipeline) and fill in the TODOs:
/*
Foundation 2: Logs Pipeline
Pipeline: loki.source.file -> loki.process -> loki.write
*/
// Step 1: Discover and read files
loki.source.file "mission_control" {
targets = [{"__path__" = "/var/log/alloy/*.log"}]
file_match {
enabled = true
}
forward_to = [TODO] // Forward to loki.process receiver
}
/*
Step 2: Extract "component" from JSON and promote to label
Original log: {"level":"INFO","component":"agents","msg":"Request received"}
Labels are indexed - makes filtering fast in Grafana
*/
loki.process "mission_control_logs" {
//`stage.json` parses the log body
stage.json {
expressions = {
component = "component", // Extract the "component" field from the JSON log body
}
}
// Add the extracted field as a Loki label
//`stage.labels` adds extracted fields to Loki labels
stage.labels {
values = {
component = "", // Empty string means "use the extracted value with the same name."
}
}
forward_to = [TODO] // Forward to the loki.write receiver
}
// Step 3: Send logs to Loki
loki.write "docker_loki" {
endpoint {
url = "http://loki:3100/loki/api/v1/push"
}
}Full solution
/*
Foundation 2: Logs Pipeline
Pipeline: loki.source.file -> loki.process -> loki.write
*/
// Discover and read log files
loki.source.file "mission_control" {
targets = [{"__path__" = "/var/log/alloy/*.log"}]
file_match {
enabled = true
}
forward_to = [loki.process.mission_control_logs.receiver]
}
// Parse JSON logs and extract labels
loki.process "mission_control_logs" {
stage.json {
expressions = {
component = "component",
}
}
stage.labels {
values = {
component = "",
}
}
forward_to = [loki.write.docker_loki.receiver]
}
// Send logs to Loki
loki.write "docker_loki" {
endpoint {
url = "http://loki:3100/loki/api/v1/push"
}
}Tip
You can also use local.file_match to perform file discovery. This used to be the only way to do it. However, using the file_match block inside loki.source.file has less overhead and results in a simpler pipeline.
Important
Remember to reload Alloy after every config change: make alloy-reload
- Check the Mission Control Overview dashboard (localhost:3000). The logs panel should now show data.
-
Next, open Explore from the Grafana sidebar.
Explore is Grafana's query playground. It lets you run ad-hoc queries without building a dashboard.
-
Select Loki as the data source and try this query:
{component="agents"} -
You should see logs filtered to just field agent activity.
- Check
Common labelsat the top. It showscomponent=agents. That's the label we extracted from the JSON.
Tip
Why labels matter: Without that label, you'd have to search the entire log body to filter by component. With it, you can filter instantly. Labels are indexed in Loki, making queries fast and efficient.
- Use
prometheus.scrapeto scrape metrics from the application - Use
prometheus.relabelto standardize label names for consistency - Use
prometheus.remote_writeto export metrics to Mimir
The mission-control application exposes a set of metrics in Prometheus format, accessible at the standard /metrics endpoint. Among those metrics are agent check-ins, used to track
which field operatives are active and where they're located.
Your task:
Build a metrics pipeline to scrape mission-control, standardize the labels, and send them to Mimir.
Once it's configured, you'll see a world map showing agent locations, grouped by country.
Two metrics track your agents: active_agents tells you who's currently online, and agent_comms_total counts their check-ins. Together, they paint a picture of field activity.
But these metrics label the same data differently. active_agents uses agent_id and country_code, while agent_comms_total uses id and region. This makes it harder to correlate metrics later. You'll use prometheus.relabel to standardize them.
This kind of inconsistency is common as organizations grow and teams or departments develop different observability practices. Luckily, it's all solvable at the collector level!
Pipeline:
prometheus.scrape → prometheus.relabel → prometheus.remote_write
- Scrape metrics from
mission-control:8080 - Rename
id→agent_idandregion→country_code - Send to Mimir at
http://mimir:9009/api/v1/push
Add this to your config.alloy file (below your logs pipeline) and fill in the TODOs:
/*
Foundation 3: Metrics Pipeline
Pipeline: prometheus.scrape -> prometheus.relabel -> prometheus.remote_write
*/
// Step 1: Scrape metrics from mission-control
prometheus.scrape "mission_control" {
scrape_interval = "5s"
scrape_timeout = "4s"
targets = [{"__address__" = "mission-control:8080"}]
forward_to = [TODO] // Forward to prometheus.relabel.standardize_agent_labels receiver
}
// Step 2: Standardize label names
/*
Standardize labels so all agent metrics use consistent naming
Before:
active_agents{agent_id="ALPHA-007", country_code="US"}
agent_comms_total{id="ALPHA-007", region="US"}
After:
active_agents{agent_id="ALPHA-007", country_code="US"}
agent_comms_total{agent_id="ALPHA-007", country_code="US"}
Rename: id -> agent_id, region -> country_code
*/
prometheus.relabel "standardize_agent_labels" {
rule {
action = "replace"
source_labels = ["id"]
regex = "(.+)"
target_label = "agent_id"
}
rule {
action = "labeldrop"
regex = "^id$"
}
// We'll do the same with the "region" label to rename it to "country_code" and drop the old label
rule {
action = "replace"
source_labels = ["TODO"]
regex = "(.+)"
target_label = "TODO"
}
rule {
action = "labeldrop"
regex = "^TODO$"
}
forward_to = [TODO] // Forward to prometheus.remote_write receiver
}
// Step 3: Send metrics to Mimir
prometheus.remote_write "docker_mimir" {
endpoint {
url = "http://mimir:9009/api/v1/push"
}
}
Full solution
/*
Foundation 3: Metrics Pipeline
Pipeline: prometheus.scrape -> prometheus.relabel -> prometheus.remote_write
*/
// Scrape metrics from mission-control
prometheus.scrape "mission_control" {
scrape_interval = "5s"
scrape_timeout = "4s"
targets = [{"__address__" = "mission-control:8080"}]
forward_to = [prometheus.relabel.standardize_agent_labels.receiver]
}
// Step 2: Standardize label names
/*
Standardize labels so all agent metrics use consistent naming
Before:
active_agents{agent_id="ALPHA-007", country_code="US"}
agent_comms_total{id="ALPHA-007", region="US"}
After:
active_agents{agent_id="ALPHA-007", country_code="US"}
agent_comms_total{agent_id="ALPHA-007", country_code="US"}
Rename: id → agent_id, region → country_code
*/
prometheus.relabel "standardize_agent_labels" {
rule {
action = "replace"
source_labels = ["id"]
regex = "(.+)"
target_label = "agent_id"
}
rule {
action = "labeldrop"
regex = "^id$"
}
rule {
action = "replace"
source_labels = ["region"]
regex = "(.+)"
target_label = "country_code"
}
rule {
action = "labeldrop"
regex = "^region$"
}
forward_to = [prometheus.remote_write.docker_mimir.receiver] // Forward to prometheus.remote_write receiver
}
// Step 3: Send metrics to Mimir
prometheus.remote_write "docker_mimir" {
endpoint {
url = "http://mimir:9009/api/v1/push"
}
}Important
Remember to reload Alloy after every config change: make alloy-reload
- Check the Mission Control Overview dashboard (localhost:3000) and view the Active Agents panel.
-
To verify your relabel rules are working, go to Explore, select Mimir as a data source and run this query:
agent_comms_total{agent_id=~".+"} -
You should see that
agent_comms_totalmetric now hasagent_idandcountry_codelabels (purple box). That confirms that our relabeling rules worked.
Alloy's UI is a useful tool that helps you visualize how Alloy is configured and what it is doing so you are able to debug efficiently.
Navigate to localhost:12347 to see the list of components (orange box) that alloy is currently configured with.
Click on the blue ‘view’ button on the right side (red arrow).
This page shows us the health of the component, the arguments it's using, and its current exports (green box).
This page also gives us quick access to the component’s documentation (orange arrow) and a Live Debugging view (yellow arrow).
When we click on the Live Debugging view, we will be able to see a real-time stream of telemetry flowing through a component.
Navigate to the Graph tab to access the graph of components and how they are connected.

The number (pink box) shown on the dotted lines shows the rate of transfer between components. The window at the top (pink box) configures the interval over which Alloy should calculate the per-second rate, so a window of ‘5’ means that Alloy should look over the last 5 seconds to compute the rate.
The color of the dotted line signifies what type of data are being transferred between components. See the color key (green box) for clarification.
Tip
Quick debugging checklist with the Alloy UI:
- Ensure that no component is reported as unhealthy.
- Ensure that the arguments and exports for misbehaving components appear correct.
- Use live debugging to verify the data is what you expect.
Note
Nice work completing the foundations!
You now have a working pipeline for traces, logs, and metrics.
If you got stuck on any foundation, you can copy a checkpoint file to catch up:
cp alloy/checkpoints/foundation3.alloy alloy/config.alloy
make alloy-reloadTraining's over. Each mission throws a new crisis at the pipelines you've built. Activate with make missionN, reset all with make reset.
make mission1An adversary discovered that our server records the full request path as a metric label. They're now flooding us with requests to thousands of random URLs, paths like /api/a3f8c2e1 that don't map to any real endpoint. Every unique path creates a new time series in Mimir, and cardinality is climbing fast.
Your orders:
You'll be expanding on what you did in the Metrics foundation. There is a pre-made regular expression provided for you to use at http://mission-control:8080/api/metrics/allowed-paths. Use prometheus.relabel with a keep action on the path label to filter out the noise. The Alloy standard library functions may be helpful here!
Open Explore in Grafana (localhost:3000/explore), select Mimir as the data source, and run:
http_requests_total
If the query times out, try narrowing the time range to Last 5 minutes. That's the cardinality explosion in action.
Scroll through the series list and notice the random paths like /api/a3f8c2e1 with status="404". After you apply the fix, you should only see legitimate paths.
Add these components to your metrics section (below standardize_agent_labels) and fill in the TODOs. Relabeling must happen before remote_write so only allowed path labels reach Mimir.
/*
Mission I: Rogue Dimension
Pipeline: remote.http -> prometheus.relabel -> prometheus.remote_write
*/
// Step 1: Fetch the allowlist of legitimate paths
remote.http "allowed_paths_regex" {
url = "http://mission-control:8080/api/metrics/allowed-paths"
}
// Step 2: Filter - only keep metrics where "path" is a known legitimate route
prometheus.relabel "mission1" {
rule {
action = "keep"
source_labels = ["path"]
regex = TODO
}
forward_to = [prometheus.remote_write.docker_mimir.receiver]
}
Hint 1: you don't need to write regex
The API at http://mission-control:8080/api/metrics/allowed-paths returns a ready-made regex for you.
Try looking at the endpoint to see what the response looks like: localhost:8080/api/metrics/allowed-paths
Use localhost when curling from your terminal. Inside the Alloy config, use mission-control, which is the Docker-internal hostname.
Hint 2: accessing the response body
remote.http exposes the fetched response body as .content. You can pass it to standard library functions or other components!
Hint 3: parsing the response
The response body is JSON, not a plain string. Use encoding.from_json() from the Alloy standard library to parse it into an object you can access.
Hint 4: putting it together
In Alloy, you can chain expressions together. The general pattern for this particular exercise looks like:
encoding.from_json(some_component.content).some_field
Parse the content, then access the field you need with dot notation.
Hint 5: pipeline order
Scraped metrics should flow through your standardize relabel component, then through this new keep filter, then remote_write. If the filter is bypassed, cardinality in Mimir will not improve.
Full solution
// Fetch the allowlist (unchanged)
remote.http "allowed_paths_regex" {
url = "http://mission-control:8080/api/metrics/allowed-paths"
}
// Filter: only keep series whose `path` label matches the allowlist
prometheus.relabel "mission1" {
rule {
action = "keep"
source_labels = ["path"]
regex = encoding.from_json(remote.http.allowed_paths_regex.content).regex
}
forward_to = [prometheus.remote_write.docker_mimir.receiver]
}Make sure to rewire the existing standardize_agent_labels component so metrics flow through the new filter before reaching Mimir:
prometheus.relabel "standardize_agent_labels" {
// ...existing rules unchanged...
forward_to = [prometheus.relabel.mission1.receiver]
}How it works
remote.http.allowed_paths_regex.contentexposes the raw response body from the fetched URL, a JSON string like{"regex":"^(/api/agents|/metrics|...)$"}.encoding.from_jsonparses that string into an object so you can reach into it. Chaining.regexpulls the pre-built pattern out of the object.- The
keepaction tellsprometheus.relabelto drop every series whosepathlabel does not match the regex, so the attacker's random/api/a3f8c2e1-style paths never make it toremote_write. - The final pipeline is
scrape -> standardize_agent_labels -> mission1 -> remote_write. Make surestandardize_agent_labels.forward_topoints atprometheus.relabel.mission1.receiverso the filter sits in the path.
Note
After reloading, give it ~20 seconds for scraped metrics to flow through the pipeline and land in Mimir before verifying.
make mission1-verifyThen you can confirm with the Alloy livedebugging view:
- Navigate to the Alloy UI (localhost:12347)
- Click on
Viewforprometheus.remote_write.docker_mimir - Click
livedebuggingat the top, under the component identifier - Wait for one scrape to come in and click
Stopon the top right - In the search bar, enter
http_requests_totalto filter for only the relevant samples
You should see only legitimate paths (like /api/agents, /metrics, etc.). No more random paths!
[!TIP] Once you're done, you can run
make mission1-stopto lower the cardinality again
make mission2After the last incident, the higher-ups want us to collect the DEBUG logs we were previously dropping. It turns out those include request logs that could have helped us track down the attacker. But pumping everything into Loki would blow the budget. The directive: archive all logs to a new S3 bucket named audit-logs, but only send INFO/WARN/ERROR logs to Loki for fast queries.
Your orders: The skills you picked up in Foundation II will come in handy here. Split your log pipeline into two parallel paths:
- All logs -> S3 via
otelcol.receiver.loki->otelcol.processor.batch->otelcol.exporter.awss3 - Non-DEBUG only -> Loki via a second
loki.processand using astage.dropto drop anyDEBUGlogs
Extend your logs section (below loki.process.mission_control_logs). Update its forward_to to fan out to two destinations, then add the components below and fill in the TODOs.
/*
Mission II: Operation Cold Storage
Pipeline:
Path 1 (all logs): otelcol.receiver.loki -> otelcol.processor.batch -> otelcol.exporter.awss3
Path 2 (non-DEBUG): loki.process -> loki.write
*/
// TODO: update loki.process.mission_control_logs forward_to with two receivers
// Path 1: Bridge Loki logs to OTLP format for S3 export
otelcol.receiver.loki "all_logs" {
output {
logs = [TODO]
}
}
otelcol.processor.batch "s3_logs" {
timeout = "10s"
send_batch_size = 100
send_batch_max_size = 200
output {
logs = [TODO]
}
}
otelcol.exporter.awss3 "audit_logs" {
s3_uploader {
s3_bucket = "TODO"
s3_prefix = "logs"
endpoint = "http://localstack:4566"
disable_ssl = true
s3_force_path_style = true
}
marshaler {
type = "body"
}
sending_queue {
batch {
flush_timeout = "500ms"
min_size = 100
sizer = "items"
}
}
}
// Path 2: Filter logs before sending to Loki
loki.process "filter_debug" {
stage.json {
expressions = {
// Which JSON field contains the log level? Try using livedebugging to see what an example log line looks like!
level = "TODO",
}
}
// TODO: add a stage to drop the logs you don't want forwarded to Loki
forward_to = [loki.write.docker_loki.receiver]
}Note
Why OTLP components for Path 1?
The S3 exporter (otelcol.exporter.awss3) is an OpenTelemetry component. There's no native Loki component for writing to S3. To bridge the gap, otelcol.receiver.loki accepts Loki log entries and converts them to OTLP format, so they can flow through the otelcol pipeline to S3.
Hint 1: wiring the fan-out
forward_to takes an array, so you can list multiple receivers to send logs to both paths simultaneously. otelcol.receiver.loki exposes a .receiver export that accepts Loki log entries.
Hint 2: filtering logs
Check the loki.process docs for stages that can filter log lines based on an extracted field value.
Hint 3: stage details
The stage can take a source (the extracted field to compare) and value (the string to match against) to directly compare the value. Only
source and value are necessary for this mission, but it is possible to perform more complex matching.
Full solution
Update loki.process.mission_control_logs so its forward_to fans out to both paths:
loki.process "mission_control_logs" {
// ...existing stage.json and stage.labels unchanged...
forward_to = [
otelcol.receiver.loki.all_logs.receiver, // Path 1: all logs -> S3
loki.process.filter_debug.receiver, // Path 2: non-DEBUG -> Loki
]
}Path 1 (archive everything to S3):
otelcol.receiver.loki "all_logs" {
output {
logs = [otelcol.processor.batch.s3_logs.input]
}
}
otelcol.processor.batch "s3_logs" {
timeout = "10s"
send_batch_size = 100
send_batch_max_size = 200
output {
logs = [otelcol.exporter.awss3.audit_logs.input]
}
}
otelcol.exporter.awss3 "audit_logs" {
s3_uploader {
s3_bucket = "audit-logs"
s3_prefix = "logs"
endpoint = "http://localstack:4566"
disable_ssl = true
s3_force_path_style = true
}
marshaler {
type = "body"
}
sending_queue {
batch {
flush_timeout = "500ms"
min_size = 100
sizer = "items"
}
}
}Path 2 (drop DEBUG before Loki):
loki.process "filter_debug" {
stage.json {
expressions = {
level = "level",
}
}
stage.drop {
source = "level"
value = "DEBUG"
}
forward_to = [loki.write.docker_loki.receiver]
}How it works
forward_toaccepts a list, so listing both receivers inloki.process.mission_control_logsfans every log line out to both downstream paths simultaneously.- Path 1 (S3 archive): The S3 exporter is an OpenTelemetry component, so Loki-format entries have to be converted first.
otelcol.receiver.lokiaccepts Loki entries and emits them as OTLP logs. From there the pipeline is a standard OTel chain:batchholds entries up to 10s or 100 items to keep the S3 object count sane, andawss3writes them to theaudit-logsbucket under alogs/year=.../month=...prefix. Theendpointoverride points the exporter at the local localstack container instead of real AWS. - Path 2 (filtered Loki):
stage.jsonpulls thelevelfield out of the JSON body and stores it in the stage's extracted map under the keylevel.stage.dropthen references that same key:source = "level"tells it which extracted field to look at, andvalue = "DEBUG"is the string to compare against. Conceptually it is evaluatingextracted["level"] == "DEBUG", and any line that matches is dropped. Everything else (INFO, WARN, ERROR) is forwarded toloki.write.docker_loki, the same Loki sink you wired up in the Logs foundation. - The result is a tiered storage pattern: all logs are archived cheaply in S3 for compliance and post-incident forensics, and only the higher-signal levels are indexed in Loki for interactive querying.
make mission2-verifyThen confirm both paths manually:
Path 2 (Loki, no DEBUG): Open Explore in Grafana (localhost:3000/explore), select Loki, and set the time range to Last 5 minutes. Run:
{filename=~".+"}
You should still see INFO logs but no DEBUG logs in the time since you reloaded your Alloy config.
Path 1 (S3, all logs): Run make s3-list in your terminal. Look for:
- File keys like
logs/year=2026/month=04/day=13/hour=21/minute=54/logs_...txtshow log files are being written to S3 organized by timestamp - DEBUG entries in the content (
"level":"DEBUG") mixed with INFO and WARN confirms S3 is receiving all logs, not just the filtered ones that go to Loki
make mission3We need to keep our network traffic to a minimum to maintain a low profile. Right now we're sending every single trace to Tempo, and that kind of volume is going to attract attention.
Your orders:
Go back to the pipeline you built in Foundation I and add head sampling to cut the volume down. Insert an otelcol.processor.probabilistic_sampler component between the OTLP receiver and batch processor. Set sampling_percentage = 5.0 to keep only 5% of traces. While you're at it, think about what we're trading away here. What intelligence might slip through the cracks?
Open the Alloy UI (localhost:12347), click the Graph tab, and note the trace rate on the edges leading to Tempo. You'll compare this to the rate after you add sampling.
Add this otelcol.processor.probabilistic_sampler component to your config.alloy and fill in the TODOs. Traces should enter the sampler before the batch processor.
/*
Mission III: Selective Surveillance
Pipeline: otelcol.receiver.otlp -> otelcol.processor.probabilistic_sampler -> otelcol.processor.batch -> otelcol.exporter.otlphttp
*/
otelcol.processor.probabilistic_sampler "mission3" {
sampling_percentage = TODO
output {
traces = [TODO]
}
}Hint: wiring
OTLP components use .input to receive data, just like you wired the batch processor in the foundations.
Full solution
Update the existing otelcol.receiver.otlp "default" so traces flow into the sampler before the batch processor:
otelcol.receiver.otlp "default" {
// ...grpc and http blocks unchanged...
output {
traces = [otelcol.processor.probabilistic_sampler.mission3.input]
}
}Add the sampler itself, forwarding survivors to the existing batch processor:
otelcol.processor.probabilistic_sampler "mission3" {
sampling_percentage = 5.0
output {
traces = [otelcol.processor.batch.default.input]
}
}How it works
- The sampler makes a keep-or-drop decision the instant a trace enters the collector. By default, that decision is derived from a hash of the trace ID, so it is deterministic: any collector that sees spans from the same trace will arrive at the same verdict, and traces stay intact rather than getting partially exported.
sampling_percentage = 5.0keeps roughly 1 in every 20 traces. The other 95% are dropped immediately and never reach the batch processor, which is why volume to Tempo falls off sharply after reload.- Trade-off: head sampling is cheap and stateless, but it is blind to what happens inside the trace. A trace that contains a 500 error has the same 5% survival odds as a healthy trace, which is what the "needle-in-a-haystack" warning is pointing at. Mission IV addresses this by swapping to tail sampling, where the decision can be based on the trace's contents.
- Make sure
otelcol.receiver.otlp.default.output.tracespoints atotelcol.processor.probabilistic_sampler.mission3.input. If it still points straight at the batch processor, the sampler is bypassed entirely.
make mission3-verifyThen confirm in the Alloy UI: go to the Alloy UI (localhost:12347), click the Graph tab, and check the trace rate on the edges leading to Tempo. You should see a significant drop compared to before.
make mission4A field agent has sent us an encrypted dead drop and is transmitting the decryption key in fragments through span attributes on error traces. One piece every 15 seconds, five pieces total. With Head Sampling it will take too long to assemble the key.
Your orders:
Swap out the head sampler for otelcol.processor.tail_sampling. Configure two policies:
status_codepolicy: keep all error traces (every fragment counts)probabilisticpolicy: sample 5% of everything else
Once you've recovered all the fragments, piece together the token and crack open the dead drop.
make access-token # Check token fragment recovery (wait ~75s)
make deaddrop KEY="your-token" # Unlock the dead drop with the assembled tokenImportant
Before adding tail sampling, remove the head sampler from Mission III. You can either delete it, comment it out, or rewire the pipeline to skip it.
Add tail sampling to your trace pipeline and fill in the TODOs. Point the OTLP receiver at this processor, then continue to batch and send to Tempo as before.
/*
Mission IV: Leave No (Error) Trace
Pipeline: otelcol.receiver.otlp -> otelcol.processor.tail_sampling -> otelcol.processor.batch -> otelcol.exporter.otlphttp
*/
otelcol.processor.tail_sampling "mission4" {
decision_wait = "10s"
policy {
name = "keep_all_errors"
type = "TODO"
// TODO: add the inner block for this policy type
}
policy {
name = "sample_normal_traffic"
type = "TODO"
// TODO: add the inner block for this policy type
}
output {
traces = [otelcol.processor.batch.default.input]
}
}Hint 1: policy types
Check the otelcol.processor.tail_sampling docs for the available policy types. You need one that matches on span status, and one that samples by percentage.
Hint 2: inner block structure
Each policy type has a corresponding inner block with the same name as the type, containing that type's arguments. For example:
policy {
name = "example"
type = "latency"
latency {
threshold_ms = 5000
}
}Hint 3: error status codes
The status_code block lists the valid values for status_codes.
Full solution
Update the existing otelcol.receiver.otlp "default" so traces flow into the tail sampler instead of the head sampler from Mission III:
otelcol.receiver.otlp "default" {
// ...grpc and http blocks unchanged...
output {
traces = [otelcol.processor.tail_sampling.mission4.input]
}
}Add the tail sampler with both policies, forwarding kept traces to the existing batch processor:
otelcol.processor.tail_sampling "mission4" {
decision_wait = "10s"
policy {
name = "keep_all_errors"
type = "status_code"
status_code {
status_codes = ["ERROR"]
}
}
policy {
name = "sample_normal_traffic"
type = "probabilistic"
probabilistic {
sampling_percentage = 5.0
}
}
output {
traces = [otelcol.processor.batch.default.input]
}
}How it works
- Tail sampling buffers every span of a trace in memory as it arrives. The decision about whether to keep or drop the trace is deferred until either the root span is received or
decision_wait(10s here) elapses, whichever comes first. That delay is what lets the sampler see the full picture (status codes, attributes, latency) before committing. - Policies are evaluated with OR logic: if any policy votes to sample, the whole trace is kept. That's why the two policies in this config compose naturally:
type = "status_code"withstatus_codes = ["ERROR"]retains every trace that contains at least one error span. Each dead-drop fragment the field agent is leaking rides on an error span, so this policy guarantees every fragment survives.type = "probabilistic"withsampling_percentage = 5.0keeps 5% of healthy traffic, preserving a baseline for general observability.
- Swapping head sampling for tail sampling is exactly what unlocks the mission: under head sampling, error traces had the same 5% survival odds as anything else, so most fragments were dropped before anyone could read them. Under tail sampling, the
status_codepolicy makes error traces deterministic keepers. - Make sure
otelcol.receiver.otlp.default.output.tracespoints atotelcol.processor.tail_sampling.mission4.input, and remove or rewire around theprobabilistic_samplerfrom Mission III so traces do not flow through both.
[!IMPORTANT] Tail sampling is more involved to operate than head sampling. The workshop runs a single Alloy instance so this is handled for you, but in a real deployment there are two operational constraints to plan for:
- Memory: every span is buffered until
decision_waitexpires or the root span arrives, so the collector needs enough memory to holddecision_waitworth of in-flight traces.- Trace locality: all spans belonging to the same trace must reach the same Alloy instance, otherwise the sampler only sees part of the trace and cannot make an informed decision. Two common ways to guarantee this:
- A load balancer in front of Alloy that does consistent hashing on the trace ID, so every span for a given trace lands on the same instance.
- The
otelcol.exporter.loadbalancingcomponent, which routes spans to downstream Alloy instances by trace ID.For a fuller walkthrough of both sampling strategies and how to lay them out in a real deployment, see Sampling with the OpenTelemetry Collector.
make mission4-verify| Command | Description |
|---|---|
| Environment | |
make start |
Start all services |
make stop |
Stop all services and clean volumes |
make restart |
Restart everything |
make clean |
Full cleanup including logs |
| Config | |
make alloy-reload |
Reload Alloy config after changes |
| Monitoring | |
make logs |
Tail mission-control logs |
make alloy-logs |
Tail Alloy logs |
make metrics |
View metrics endpoint |
make status |
Check mission status |
| Missions | |
make mission1 / mission2 / mission3 / mission4 |
Activate a mission |
make mission1-verify / mission2-verify / mission3-verify / mission4-verify |
Verify a mission solution |
make reset |
Reset all missions |
| Mission IV Endgame | |
make access-token |
Check token fragment recovery (wait ~75s) |
make deaddrop KEY="..." |
Unlock the dead drop with your assembled token |
Caution
illegal character U+201C
This means you have “smart quotes” (curly quotes) in your config, likely from copying text via a browser or rich-text editor. Replace all “ and ” with plain ASCII double quotes (”).
- Alloy not receiving traces?
Run
make alloy-logsto check for connection errors. - Port conflicts? Check ports 8080, 3000, 3100, 3200, 4317, 4318, 9009, 12347.
scrape_timeout greater than scrape_interval? Addscrape_timeout = "4s"to yourprometheus.scrapeblock (or any value less thanscrape_interval).- Mission 2:
NoSuchBucket? The S3 init container may have failed on startup. Recreate the bucket manually:docker compose exec localstack curl -X PUT http://localhost:4566/audit-logs
