Core Loki for Log Analysis and Observability

Part of our "Kubernetes & Cloud" courses

2 days

Core Loki
Outline Last updated:

Course Overview

Learn to query your logs with Grafana Loki in this hands-on course for developers and support engineers who troubleshoot production systems. You write LogQL log and metric queries in Grafana and in logcli, Loki's command-line interface, and learn how Loki stores and reads your logs, so you can tell a cheap query from an expensive one before you run it. The course focuses on the skills of the people who query logs in container and Kubernetes environments, not on running a Loki installation.

Course Prerequisites

Working knowledge of containers with Docker or Podman. Experience with Kubernetes, Grafana or another query language helps but is not required, and no prior knowledge of Loki or LogQL is assumed.

Outline

This practical, hands-on course equips developers and second-line support engineers with the skills to use Grafana Loki for log analysis and troubleshooting in production environments. Unlike training that focuses on installation and administration, this course concentrates on what matters most to the people who use Loki every day: querying, analysing and extracting insights from their log data.

We do not cover administrative tasks, but we do explain how Loki works under the hood: how it indexes labels rather than log contents, how it stores and reads chunks, and which component does what. That knowledge is what makes the difference between a query that answers in a second and one that Loki refuses.

The course is designed for teams already familiar with log aggregation who now need to become proficient at querying their logs. For a broader view of observability, consider pairing it with our planned courses on Prometheus for metrics and on distributed tracing.

We cover both log queries and metric queries in LogQL, in Grafana and on the command line, and finish with what makes a query expensive: reading Loki's query statistics, measuring a query in bytes before you run it, understanding why Loki refuses a query, and rewriting it to read less.

Introduction

  • Introduce Grafana Loki as the L in the LGTM stack
  • Contrast Loki's push model with the way Prometheus pulls
  • Understand why Loki indexes labels and never the log contents
  • Appreciate how streams group the lines that share the same labels
  • Understand chunks, compressed and then flushed to object storage
  • Discuss what a query costs, and why the labels decide it
  • Distinguish good label candidates from poor ones
  • Appreciate what makes good logs from a developer's point of view

Log Collection

  • Understand why every log source needs a client that pushes its logs
  • Compare the clients for Kubernetes, Docker and Podman hosts, applications and existing pipelines
  • Send logs over OTLP, the OpenTelemetry protocol
  • Appreciate why Promtail reached its end of life
  • Introduce Grafana Alloy, a vendor-neutral distribution of the OpenTelemetry Collector
  • Understand Alloy's pipeline of components
  • Configure Alloy with blocks, attributes and expressions
  • Use discovery to find dynamic targets
  • Choose what identifies a source with relabel rules
  • Transform entries in a processing stage
  • Batch and retry pushes to Loki with loki.write
  • Collect from Kubernetes pods and the systemd journal with Alloy

The Log Entry

  • Understand the data model of a log entry
  • Inspect the wire format a client pushes to Loki
  • Distinguish the kinds of labels and where each comes from
  • Use structured metadata for values that should not become labels
  • Understand how Loki derives detected_level at write time

Loki Architecture

  • Understand Loki's microservices mode: one binary, a different target per component
  • Follow the write path through the distributor and the ingester
  • Appreciate how the ring decides which ingester owns a stream
  • Understand when an ingester flushes a chunk
  • Follow the read path from the query frontend to the queriers
  • Understand how the query frontend splits, shards and caches a query
  • Introduce the ruler and where its alerts go
  • Discuss Loki's single object store
  • Understand the compactor, which merges indexes and applies retention
  • Appreciate how each component scales independently

LogQL and logcli

  • Introduce LogQL, the query language of Grafana Loki
  • Compare LogQL with PromQL
  • Distinguish log queries from metric queries
  • Understand streams and series as the two kinds of result
  • Use logcli, Loki's command-line interface
  • Explore a new application's labels, series and detected fields
  • Run log and metric queries with query and instant-query
  • Control the size of a result with --limit and --batch
  • Follow new lines as they arrive with --tail

Log Queries

Selecting Streams and Lines

  • Write a log stream selector with label matching operators
  • Appreciate how the selector decides what a query reads
  • Filter lines with line filter expressions
  • Choose between a regular expression and a pattern

Parsing

  • Extract labels with the json and logfmt parsers
  • Parse unstructured lines with pattern and regexp
  • Understand why unpack still exists
  • Choose the right parser for a log format
  • Appreciate how a parser changes the result of a query

Filtering and Formatting

  • Refine results with label filter expressions
  • Understand the value types Loki infers in a label filter
  • Handle pipeline errors
  • Rewrite lines with line_format and its templates
  • Rename and compute labels with label_format
  • Strip ANSI colour codes with decolorize
  • Regroup results with drop and keep

Metric Queries

Time

  • Distinguish instant queries from range queries
  • Understand how the range and the step decide what each point counts
  • Use $__auto to match the range to the step in Grafana
  • Shift a query back in time with offset

Aggregations

  • Count lines and bytes with log range aggregations
  • Calculate a per-second rate
  • Detect a service that went quiet with absent_over_time
  • Combine series with vector aggregations
  • Group results with by and without
  • Rank series with topk
  • Relate each aggregation to its SQL equivalent

Values From the Line

  • Turn extracted values into samples with unwrap
  • Calculate latency quantiles
  • Understand decomposable aggregates and when Loki groups inside a range

Combining Results

  • Combine vectors with binary operators
  • Match series with on and ignoring
  • Join many-to-one with group_left
  • Filter or flag samples with comparison operators
  • Use the logical and set operators
  • Distinguish no data from zero, and repair the gaps
  • Investigate incidents with metric queries

Query Performance

  • Measure the cost of a query in bytes processed
  • Read Loki's query statistics in Grafana and logcli
  • Follow a query through the summary, index, querier and ingester statistics
  • Understand why Loki reads whole blocks before it applies the time range
  • Estimate a query's cost with logcli stats before running it
  • Find the loudest application with logcli volume
  • Recognise high cardinality and what it costs
  • Distinguish recent data in the ingesters from flushed chunks
  • Understand why max_query_series refuses a query
  • Move an aggregation to the command line when Loki refuses it
  • Decide between a line filter and a parser
  • Appreciate when CPU time matters as well as bytes
  • Understand the results cache and the chunks cache
  • Appreciate why late lines make a recent result unsafe to cache
  • Explain cache freshness, and why the latest minutes are always computed again

Frequently asked questions

How long is the Core Loki for Log Analysis and Observability course?

2 days, on-site or online. Sessions can run on consecutive days or be spread out to fit your team's schedule.

What are the prerequisites?

Working knowledge of containers with Docker or Podman. Experience with Kubernetes, Grafana or another query language helps but is not required, and no prior knowledge of Loki or LogQL is assumed.

How large are the groups?

Deliberately small so the trainer can adapt to every participant: at most 10 on-site and 7 online.

This Core Loki for Log Analysis and Observability course looks very interesting, I do however have a question

Related courses