1 The Problem
We want a log analyser: read a server log file, count how many entries are errors versus normal, find which hour had the most traffic, and list the most frequent error messages. It teaches parsing semi-structured text at scale and aggregating it into useful insight — a daily task in operations.
2 How to Think About It
Three passes over the same idea: parse each line into a structured Entry, tally by hour and by error message, then rank the tallies.
3 The Build — explained part by part
Here is the complete analyser. Like word-counter before it, this project’s C++ version gets a real hash table for free where the C version had to hand-roll one.
#pragma once
#include <optional>
#include <string>
#include <utility>
#include <vector>
constexpr int TOP_N_ERRORS = 3;
struct Entry {
std::string hour; // "00".."23"
std::string level;
std::string message;
};
struct Report {
int total = 0;
int errors = 0;
std::optional<std::pair<std::string, int>> busiest_hour; // hour, count
std::vector<std::pair<std::string, int>> top_errors;
};
// Parses one line shaped like "2026-06-24 14:30:00 ERROR Database timeout".
// Splits on whitespace but caps it at 4 fields so the message (which may
// itself contain spaces) stays whole as the last field. Lines that do not
// look like "date time LEVEL message" return std::nullopt, the same
// "this might not produce a value" pattern the basic-tier temperature-
// converter project used for an unknown conversion.
std::optional<Entry> parse_line(const std::string &line);
// Counts how often each key appears, using a real std::unordered_map hash
// table internally -- unlike C, which has no hash map at all and has to
// hand-roll a linear-search table for the same job.
// Sorts by count descending (ties broken alphabetically) and keeps the
// top `n`, the same std::vector-plus-std::sort pattern the basic-tier
// word-counter project used for its own top_n.
std::vector<std::pair<std::string, int>> top_n(
const std::vector<std::pair<std::string, int>> &counts, int n);
// Runs parse_line over every line and builds the full report: total valid
// entries, error count, the busiest hour, and the top 3 most common error
// messages.
Report analyse(const std::vector<std::string> &lines);
#include "LogAnalyser.hpp"
#include <algorithm>
#include <unordered_map>
std::optional<Entry> parse_line(const std::string &line) {
auto p1 = line.find(' ');
if (p1 == std::string::npos) return std::nullopt;
auto p2 = line.find(' ', p1 + 1);
if (p2 == std::string::npos) return std::nullopt;
auto p3 = line.find(' ', p2 + 1);
if (p3 == std::string::npos) return std::nullopt;
std::string date = line.substr(0, p1);
std::string time_s = line.substr(p1 + 1, p2 - p1 - 1);
std::string level = line.substr(p2 + 1, p3 - p2 - 1);
if (time_s.size() < 2) return std::nullopt;
if (date.find('-') == std::string::npos || time_s.find(':') == std::string::npos) {
return std::nullopt; // not a "date time" pair
}
std::string message = line.substr(p3 + 1);
while (!message.empty() && (message.back() == '\n' || message.back() == '\r')) {
message.pop_back();
}
return Entry{time_s.substr(0, 2), level, message};
}
std::vector<std::pair<std::string, int>> top_n(
const std::vector<std::pair<std::string, int>> &counts, int n) {
std::vector<std::pair<std::string, int>> sorted = counts;
std::sort(sorted.begin(), sorted.end(), [](const auto &a, const auto &b) {
if (a.second != b.second) return a.second > b.second;
return a.first < b.first;
});
if (static_cast<int>(sorted.size()) > n) sorted.resize(n);
return sorted;
}
Report analyse(const std::vector<std::string> &lines) {
Report report;
std::unordered_map<std::string, int> by_hour;
std::unordered_map<std::string, int> by_message;
for (const auto &line : lines) {
auto entry = parse_line(line);
if (!entry) continue;
report.total++;
by_hour[entry->hour]++;
if (entry->level == "ERROR") {
report.errors++;
by_message[entry->message]++;
}
}
std::vector<std::pair<std::string, int>> hour_vec(by_hour.begin(), by_hour.end());
auto busiest = top_n(hour_vec, 1);
if (!busiest.empty()) report.busiest_hour = busiest[0];
std::vector<std::pair<std::string, int>> message_vec(by_message.begin(), by_message.end());
report.top_errors = top_n(message_vec, TOP_N_ERRORS);
return report;
}
#include "LogAnalyser.hpp"
#include <fstream>
#include <iostream>
#include <sstream>
int main(int argc, char **argv) {
if (argc < 2) {
std::cerr << "Usage: log_analyser <logfile>\n";
return 1;
}
std::ifstream file(argv[1]);
if (!file) {
std::cerr << "Could not open " << argv[1] << "\n";
return 1;
}
std::vector<std::string> lines;
std::string line;
while (std::getline(file, line)) lines.push_back(line);
Report r = analyse(lines);
std::cout << "Total entries: " << r.total << "\n";
std::cout << "Errors: " << r.errors << "\n";
if (r.busiest_hour) {
std::cout << "Busiest hour: " << r.busiest_hour->first << ":00 ("
<< r.busiest_hour->second << " entries)\n";
}
std::cout << "Top error messages:\n";
for (const auto &[msg, count] : r.top_errors) {
std::cout << " " << count << "x " << msg << "\n";
}
return 0;
}
std::nullopt instead of an out-parameter plus a boolean return code, the same “this might not produce a value” pattern the basic-tier temperature-converter project used for an unrecognised conversion choice. The caller in analyse just writes if (!entry) continue;.std::unordered_map<std::string, int> by_hour, by_message; — real hash tables, replacing the two hand-rolled linear-search
Counted arrays the C version needed (C has no hash map at all). Counting becomes by_hour[entry->hour]++; — one line, average O(1), instead of a manual scan-and-insert loop.top_n — the exact same vector-plus-
std::sort pattern word-counter used in the basic tier: copy the map’s entries into a std::vector<std::pair<...>>, sort by count descending with alphabetical ties, and keep the first n. One small function, reused for both the busiest-hour ranking (n=1) and the top-3 error ranking.std::optional<std::pair<std::string, int>> busiest_hour in
Report — an empty log produces a report with no busiest hour at all, rather than a fabricated "00" or a separate boolean flag the C version needed (has_busiest_hour).
std::unordered_map directly to find the busiest hour — its order is hash-bucket order, not count order, so the first entry you see is not necessarily the busiest.std::vector and sort, as top_n does here, exactly the lesson word-counter’s basic-tier project already taught.std::nullopt from parse_line for anything that does not look right, and skip it in analyse, as this project does.4 Test & Prove Each Part
Seven checks: correct parsing of a well-formed line, rejecting a malformed one, the sort-with-ties behaviour of top_n, and the full analyse pipeline’s totals, busiest hour, top errors, and empty-input edge case.
#include "LogAnalyser.hpp"
#include <cassert>
#include <iostream>
#define RUN(name) do { name(); std::cout << "PASS: " << #name << "\n"; } while (0)
static void parses_a_well_formed_line() {
auto e = parse_line("2026-06-24 14:30:00 ERROR Database timeout");
assert(e.has_value());
assert(e->hour == "14");
assert(e->level == "ERROR");
assert(e->message == "Database timeout");
}
static void rejects_a_line_with_no_date_time_pair() {
assert(!parse_line("this is not a log line").has_value());
assert(!parse_line("notadate notatime LEVEL message").has_value());
}
static void top_n_sorts_by_count_descending_with_alphabetical_ties() {
std::vector<std::pair<std::string, int>> counts = {
{"zebra", 2}, {"apple", 2}, {"mango", 5}
};
auto top = top_n(counts, 2);
assert(top[0].first == "mango" && top[0].second == 5);
assert(top[1].first == "apple" && top[1].second == 2); // tie broken alphabetically
}
static void analyse_counts_totals_and_errors_correctly() {
std::vector<std::string> lines = {
"2026-06-24 09:00:00 INFO Server started",
"2026-06-24 09:05:00 ERROR Database timeout",
"2026-06-24 09:07:00 ERROR Database timeout",
"2026-06-24 10:00:00 ERROR Disk full",
"not a real log line at all",
};
Report r = analyse(lines);
assert(r.total == 4); // the garbage line is skipped
assert(r.errors == 3);
}
static void analyse_finds_the_busiest_hour() {
std::vector<std::string> lines = {
"2026-06-24 09:00:00 INFO a",
"2026-06-24 09:05:00 INFO b",
"2026-06-24 10:00:00 INFO c",
};
Report r = analyse(lines);
assert(r.busiest_hour.has_value());
assert(r.busiest_hour->first == "09");
assert(r.busiest_hour->second == 2);
}
static void analyse_ranks_the_top_error_messages() {
std::vector<std::string> lines = {
"2026-06-24 09:00:00 ERROR Database timeout",
"2026-06-24 09:01:00 ERROR Database timeout",
"2026-06-24 09:02:00 ERROR Disk full",
};
Report r = analyse(lines);
assert(!r.top_errors.empty());
assert(r.top_errors[0].first == "Database timeout");
assert(r.top_errors[0].second == 2);
}
static void an_empty_log_produces_a_zeroed_report_not_a_crash() {
Report r = analyse({});
assert(r.total == 0);
assert(r.errors == 0);
assert(!r.busiest_hour.has_value());
assert(r.top_errors.empty());
}
int main() {
RUN(parses_a_well_formed_line);
RUN(rejects_a_line_with_no_date_time_pair);
RUN(top_n_sorts_by_count_descending_with_alphabetical_ties);
RUN(analyse_counts_totals_and_errors_correctly);
RUN(analyse_finds_the_busiest_hour);
RUN(analyse_ranks_the_top_error_messages);
RUN(an_empty_log_produces_a_zeroed_report_not_a_crash);
std::cout << "All tests passed.\n";
return 0;
}
Compile and run with g++ -std=c++20 -Wall -Wextra -Wpedantic -o test_run LogAnalyser.cpp test_LogAnalyser.cpp && ./test_run.
5 The Interface
What it expects
$ ./loganalyser server.logWhat it returns
Total entries: 4
Errors: 3
Busiest hour: 09:00 (3 entries)
Top error messages:
2x Database timeout
1x Disk full6 Run It & Automate It
Save the code as LogAnalyser.hpp / LogAnalyser.cpp / main.cpp and compile it with g++ — that turns your source directly into a native executable for your machine. No separate runtime needed: the compiled binary runs on its own.
g++ -std=c++20 -o loganalyser main.cpp LogAnalyser.cpp && ./loganalyser server.logPoint it at any log file shaped like “2026-06-24 09:05:00 ERROR Database timeout”.
A CI tool like Jenkins runs the same compile-then-test-then-check-for-leaks steps automatically whenever the code changes — every line below has a plain explanation.
$ ./loganalyser server.log
Total entries: 4
Errors: 3
Busiest hour: 09:00 (3 entries)
Top error messages:
2x Database timeout
1x Disk fullparse_line — check the log actually uses a literal space between date, time, level, and message, and that the date contains a - and the time contains a :.std::stoi, so check any code you have added on top parses fields defensively.// Jenkinsfile — compiles, tests, and checks for leaks on every change.
pipeline {
agent any
stages {
stage('Get the code') {
// download the latest code
steps { checkout scm }
}
stage('Compile') {
steps {
// confirm a compiler is installed
sh 'g++ --version'
// compile with strict warnings on
sh 'g++ -std=c++20 -Wall -Wextra -o app *.cpp'
}
}
stage('Run the tests') {
steps {
// prints PASS/FAIL, exits non-zero on failure
sh './app'
}
}
stage('Check for memory leaks') {
steps {
// fails the build on any leak or invalid access
sh 'valgrind --error-exitcode=1 --leak-check=full ./app'
}
}
}
post {
success { echo 'All tests passed, no leaks found.' }
failure { echo 'A test or Valgrind check failed — see above.' }
}
}
- Report the busiest hour per level. Not just overall, but separately for ERROR, WARN, and INFO. (Teaches: a nested or compound map key.)
- Add a date range filter. Only analyse lines between two given dates. (Teaches: string comparison as a stand-in for date comparison, since the dates here sort correctly as plain text.)
- Stream instead of loading the whole file. Process line by line without holding a
std::vector<std::string>of the whole file in memory. (Teaches: whyanalysecurrently needs the whole file up front, and what would have to change.)
std::unordered_map a second time for real hash-table aggregation, reused the vector-plus-sort top_n pattern from word-counter, and used std::optional to report “no busiest hour” honestly for an empty log instead of a fabricated value or a separate boolean flag. Related: STL Containers, Modern C++ (C++11–C++23).