Monitoring Docker Containers with cAdvisor and InfluxDB

Container workloads have quietly become the backbone of modern application delivery, and the teams running those workloads often discover the hard way that visibility is the difference between a calm Friday afternoon and a frantic Sunday night. Whether you are running a small SaaS in Surry Hills or a microservices platform in Brisbane, the same principle applies: if you cannot see what is happening inside your containers, you cannot defend their performance or your reputation.

cAdvisor and InfluxDB form a lightweight, dependable pairing for capturing container-level metrics and storing them as time-series data. cAdvisor exposes resource usage and historical statistics from each running container, while InfluxDB ingests that data and makes it queryable. The combination is straightforward enough for a solo operator yet robust enough to back a serious production environment.

The following walkthrough covers the rationale, the configuration, and the operational habits that turn raw metrics into useful signal. Australian operators have a few extra considerations around latency, bandwidth, and data sovereignty, and those will be woven through the discussion where they actually matter.

Why container observability deserves attention in Australia

Australian data centres tend to be concentrated in Sydney and Melbourne, and the rest of the country often connects to those hubs over the NBN or interstate fibre. A small team in Perth pushing metrics to a SaaS backend in Frankfurt can find themselves burning bandwidth they would rather use for actual application traffic. Self-hosting your monitoring stack in-region, or at least within an Australian cloud region, keeps the round-trip times low and protects the budget.

There is also a regulatory angle worth taking seriously. The Australian Privacy Principles and the Notifiable Data Breaches scheme set expectations about how telemetry is handled, especially when containers process customer records. Keeping metrics in the AWS Sydney region, or in an equivalent sovereign zone, helps meet those obligations. Operators who need a deeper view of these compliance considerations often consult SMB Research for current analysis.

Finally, the cost model for cloud resources in Australia is denominated in Australian dollars, and oversize monitoring infrastructure shows up directly on the monthly invoice from local providers. A right-sized cAdvisor and InfluxDB stack gives you the visibility you need without the runaway spend.

Preparing the host for cAdvisor

cAdvisor runs as a single container with privileged access to the Docker socket, so the first decision is where to place it. Most operators run it on every Docker host that they want to observe, treating each instance as a local agent rather than a centralised scraper. That approach scales well because each agent buffers its own metrics and the network only carries the data that InfluxDB actually needs.

On the host itself, make sure the kernel exposes cgroup v1 or v2 metrics in a way that cAdvisor can read. Modern Ubuntu 22.04 LTS and the current Amazon Linux images used in ap-southeast-2 instances both work out of the box. If you are running on older hosts, particularly some of the older on-premises kit still common in regional Australian offices, you may need to enable additional cgroup controllers before cAdvisor will report memory pressure correctly.

Persistent storage for cAdvisor is rarely needed because the tool keeps only a short rolling window in memory. The real persistence question belongs to InfluxDB, which is the next piece of the pipeline.

Building the InfluxDB time-series backend

InfluxDB is purpose-built for the kind of high-cardinality, timestamped data that container metrics generate. A single node running InfluxDB 2.x handles thousands of series without complaint, which suits most Australian teams that are not running hyperscale fleets. The official Docker image is the easiest path, and binding the data directory to a host path keeps the database safe across container restarts.

A few configuration choices are worth thinking about ahead of time. Retention policies should match how long you genuinely need the data, because the storage cost of keeping a full year of per-container metrics in ap-southeast-2 will eventually surprise you. A common pattern is raw data at full resolution for seven days, downsampled five-minute data for thirty days, and hourly aggregates for a year.

Authentication tokens deserve careful handling, particularly in environments where developers may come and go. Create a dedicated write token for cAdvisor and a separate read token for dashboards, and rotate them on a schedule. Treat the tokens the same way you would treat AWS access keys, like credentials that someone could absolutely abuse.

Wiring the data path together

Once both pieces are running, the data path is mostly a matter of telling cAdvisor where to write. The container exposes a configuration flag for the InfluxDB endpoint, and the storage driver within cAdvisor takes care of translating raw cgroup readings into InfluxDB line protocol. The connection should use TLS wherever the InfluxDB instance is reachable over an untrusted network, and within a single host that is often just a UNIX socket or a localhost bind.

Network segmentation is worth a moment of thought. Keeping cAdvisor and InfluxDB on the same Docker bridge network is convenient, but it does mean that any container with a route to that bridge can potentially reach the database port. A dedicated monitoring network that only the agent and the database share is a cleaner pattern, and it maps neatly onto how teams in Melbourne and Sydney typically structure their overlay networks.

If you are running the stack in Kubernetes rather than plain Docker, cAdvisor is usually replaced by the kubelet's metrics endpoint. The InfluxDB side of the equation stays the same, which is one of the reasons this combination has aged well.

Querying metrics and building useful dashboards

Raw metrics are not the same thing as a dashboard, and the difference is where the real work happens. InfluxDB's query language lets you group by container name, image, or label, which makes it straightforward to spot the noisy neighbour dragging down a host. A common starter query is the rolling five-minute mean of container CPU usage, grouped by container, sorted descending, and that single view often surfaces the actual culprit when someone reports that the website is slow.

Grafana sits comfortably on top of InfluxDB for visualising the data, and the integration is well documented. For teams that prefer a lighter approach, the InfluxDB UI itself is enough for ad-hoc investigation. The point is to keep the number of dashboards manageable so that they stay useful; a wall of unused panels is worse than no dashboard at all.

For sites that have a mixed workload, including a PHP front-end talking to backend services, container metrics can be cross-referenced with application-level timing. A guide to optimising PHP-FPM performance for high-traffic WordPress sites walks through the application side of that problem, and the two perspectives together paint a much fuller picture than either one alone.

Common pitfalls and operational gotchas

cAdvisor will happily report whatever the kernel gives it, which is occasionally misleading. Short-lived containers that start and stop between scrapes produce gaps that look like downtime, and they will inflate certain percentiles if you are not careful with your queries. Tagging every container with a stable identifier, rather than relying on container IDs, makes the data far easier to work with over time.

Time zone handling is another recurring trap. InfluxDB stores everything in UTC, which is correct, but the moment a developer in Adelaide opens a dashboard in their local time, the displayed values shift. Configuring the dashboard tool to render in Australian Eastern time, accounting for daylight saving, prevents a lot of confused Slack messages at the start of October.

Finally, remember that cAdvisor is a metrics tool, not a logging tool. If you need to debug a specific request, the answer lives in application logs, not in cAdvisor graphs. The two systems complement each other, and a serious operation keeps them separate.

Hardening the stack for production

Production deployments benefit from a few extra layers of discipline. Running InfluxDB on a dedicated host, or at least a dedicated VM, isolates the database from the noisy neighbour problem it is supposed to be measuring. Backups of the InfluxDB data directory should be automated and tested, because the loss of a metrics archive is rarely noticed until someone needs the data three months later.

Updates to cAdvisor are infrequent but worth tracking, particularly when Docker engine releases change the cgroup layout. Subscribing to the upstream release notes, or following a curated feed, saves the cost of being surprised by a behaviour change during a weekend incident.

For teams that want to push the stack further, adding a Kapacitor or Flux task layer turns the time-series data into alerts. A simple rule that pages on container restart count, or on sustained memory pressure, catches most of the issues that would otherwise wake an on-call engineer in the middle of the night.

If you are looking to extend the tooling with custom automation, the F# and Mono on OSX guide shows how to set up a small functional language environment for scripting tasks like metric parsing and alert dispatch, and the same patterns translate directly to Linux hosts.

A reliable monitoring pipeline pays for itself the first time it tells you about a problem before a customer does. Run cAdvisor on every host that matters, point it at an InfluxDB instance you actually control, and spend an afternoon building the dashboards your team will actually use. The combination is modest in cost, generous in payoff, and a good fit for the way Australian operators tend to work: small teams, real workloads, and a strong preference for tools that do exactly what they say on the tin. If your team needs a hand designing a monitoring stack that fits within a sovereign cloud footprint or an NBN-constrained site, get in touch and we can map out an approach that matches the realities on the ground.

Experience

Information Technology Consulting

Independent Practice

Provides IT consulting services focused on infrastructure planning, cloud migration strategy, and systems architecture. Engagements draw on years of hands-on sysadmin and development experience across Linux, Windows, and hybrid environments.

K9 Search & Rescue Volunteer

Ongoing

Active participant in K9 Search & Rescue operations, combining technical logistics skills with field support for canine search teams.

Karl Katzke's Blog

October 2006 – May 2014

Published a long-running personal technology blog covering cloud vs. in-house infrastructure, F# and Mono on OSX, hardware vendor critiques, RAID card performance analysis, and sysadmin storytelling. Notable posts include "When Sysadmins Ruled the Earth" (May 15, 2014) and "Getting Started with F# and Mono on OSX" (December 22, 2012).

Credentials

A small badge icon with a shield shape in muted blue tones on a light background

Systems Administration

Deep experience with Linux (RHEL, SLES, CentOS), high-availability clusters, and STONITH configurations.

A small badge icon with a gear shape in muted blue tones on a light background

Cloud Infrastructure

Practical knowledge of AWS EC2, reserved instances, and cost analysis for cloud vs. on-premises deployments.

A small badge icon with a code symbol in muted blue tones on a light background

Development

Proficient in F#, PHP (Symfony), and cross-platform tooling including Mono and MonoDevelop on OSX.

Studies

F# & Functional Programming

Self-directed, 2012

Explored strongly typed functional programming with F# on OSX using the Mono runtime. Published a detailed getting-started guide covering toolchain setup and cross-platform game development research.

High-Availability & Cluster Management

Professional Development, 2009

Configured and documented crm_mon email alerting for STONITH events on SLES11-HAE clusters, integrating with Nagios monitoring for production environments.

Hardware & Storage Performance

Ongoing

Conducted hands-on benchmarking of SATA/SAS RAID controllers including HighPoint RocketRaid 2740 and LSI/SuperMicro AOC-USASLP2-H8iR, comparing against software RAID configurations.

Skills

A small icon representing a server with clean geometric lines in slate blue

Linux Administration

RHEL, SLES, CentOS — package management, kernel tuning, HA clustering, and monitoring integration.

A small icon representing a cloud shape with clean geometric lines in slate blue

Cloud Architecture

AWS EC2, reserved-instance planning, cost modeling, and hybrid infrastructure strategy.

A small icon representing code brackets with clean geometric lines in slate blue

F# & .NET/Mono

Functional programming on OSX, MonoDevelop toolchain, and cross-platform game-dev exploration.

A small icon representing a database cylinder with clean geometric lines in slate blue

PHP & Symfony

Web application development with the Symfony framework and the broader PHP ecosystem.

A small icon representing a storage drive with clean geometric lines in slate blue

Storage & RAID

SATA/SAS controller evaluation, md RAID configuration, and performance benchmarking.

A small icon representing a shield with clean geometric lines in slate blue

High Availability

Pacemaker, STONITH, crm_mon alerting, and Nagios integration for production cluster monitoring.