← thecodex.expert · The Codex Family of Knowledge
Tier 3 · Upper-Intermediate · C++ Project

Log Analyser

Parse a log file shaped like date time LEVEL message and report the busiest hour and the most common error messages. The interesting design choice, again, is what data structure does the counting.

🧠 Teaches how to think spoonfed, every age Last verified:

1 The Problem

We want a log analyser: read a server log file, count how many entries are errors versus normal, find which hour had the most traffic, and list the most frequent error messages. It teaches parsing semi-structured text at scale and aggregating it into useful insight — a daily task in operations.

Where this shows up: monitoring and observability, debugging production issues, security analysis, performance tuning. When something breaks at 3am, the person who can analyse the logs is the one who fixes it.

2 How to Think About It

Three passes over the same idea: parse each line into a structured Entry, tally by hour and by error message, then rank the tallies.

The plan — in plain English
1. Read the log file, one line at a time. → 2. Parse each line into hour, level, and message — skipping anything malformed rather than crashing on it. → 3. Tally counts by hour, and separately by error message. → 4. Rank both tallies and report the busiest hour and the top 3 errors.

Read log lines

Parse each: time, level, message

Count errors

Count entries per hour

Count message frequency

Report insights

3 The Build — explained part by part

Here is the complete analyser. Like word-counter before it, this project’s C++ version gets a real hash table for free where the C version had to hand-roll one.

C++LogAnalyser.hpp / LogAnalyser.cpp / main.cpp
#pragma once
#include <optional>
#include <string>
#include <utility>
#include <vector>

constexpr int TOP_N_ERRORS = 3;

struct Entry {
    std::string hour;    // "00".."23"
    std::string level;
    std::string message;
};

struct Report {
    int total = 0;
    int errors = 0;
    std::optional<std::pair<std::string, int>> busiest_hour; // hour, count
    std::vector<std::pair<std::string, int>> top_errors;
};

// Parses one line shaped like "2026-06-24 14:30:00 ERROR Database timeout".
// Splits on whitespace but caps it at 4 fields so the message (which may
// itself contain spaces) stays whole as the last field. Lines that do not
// look like "date time LEVEL message" return std::nullopt, the same
// "this might not produce a value" pattern the basic-tier temperature-
// converter project used for an unknown conversion.
std::optional<Entry> parse_line(const std::string &line);

// Counts how often each key appears, using a real std::unordered_map hash
// table internally -- unlike C, which has no hash map at all and has to
// hand-roll a linear-search table for the same job.
// Sorts by count descending (ties broken alphabetically) and keeps the
// top `n`, the same std::vector-plus-std::sort pattern the basic-tier
// word-counter project used for its own top_n.
std::vector<std::pair<std::string, int>> top_n(
    const std::vector<std::pair<std::string, int>> &counts, int n);

// Runs parse_line over every line and builds the full report: total valid
// entries, error count, the busiest hour, and the top 3 most common error
// messages.
Report analyse(const std::vector<std::string> &lines);

#include "LogAnalyser.hpp"
#include <algorithm>
#include <unordered_map>

std::optional<Entry> parse_line(const std::string &line) {
    auto p1 = line.find(' ');
    if (p1 == std::string::npos) return std::nullopt;
    auto p2 = line.find(' ', p1 + 1);
    if (p2 == std::string::npos) return std::nullopt;
    auto p3 = line.find(' ', p2 + 1);
    if (p3 == std::string::npos) return std::nullopt;

    std::string date = line.substr(0, p1);
    std::string time_s = line.substr(p1 + 1, p2 - p1 - 1);
    std::string level = line.substr(p2 + 1, p3 - p2 - 1);
    if (time_s.size() < 2) return std::nullopt;
    if (date.find('-') == std::string::npos || time_s.find(':') == std::string::npos) {
        return std::nullopt; // not a "date time" pair
    }

    std::string message = line.substr(p3 + 1);
    while (!message.empty() && (message.back() == '\n' || message.back() == '\r')) {
        message.pop_back();
    }

    return Entry{time_s.substr(0, 2), level, message};
}

std::vector<std::pair<std::string, int>> top_n(
    const std::vector<std::pair<std::string, int>> &counts, int n) {
    std::vector<std::pair<std::string, int>> sorted = counts;
    std::sort(sorted.begin(), sorted.end(), [](const auto &a, const auto &b) {
        if (a.second != b.second) return a.second > b.second;
        return a.first < b.first;
    });
    if (static_cast<int>(sorted.size()) > n) sorted.resize(n);
    return sorted;
}

Report analyse(const std::vector<std::string> &lines) {
    Report report;
    std::unordered_map<std::string, int> by_hour;
    std::unordered_map<std::string, int> by_message;

    for (const auto &line : lines) {
        auto entry = parse_line(line);
        if (!entry) continue;
        report.total++;
        by_hour[entry->hour]++;
        if (entry->level == "ERROR") {
            report.errors++;
            by_message[entry->message]++;
        }
    }

    std::vector<std::pair<std::string, int>> hour_vec(by_hour.begin(), by_hour.end());
    auto busiest = top_n(hour_vec, 1);
    if (!busiest.empty()) report.busiest_hour = busiest[0];

    std::vector<std::pair<std::string, int>> message_vec(by_message.begin(), by_message.end());
    report.top_errors = top_n(message_vec, TOP_N_ERRORS);
    return report;
}

#include "LogAnalyser.hpp"
#include <fstream>
#include <iostream>
#include <sstream>

int main(int argc, char **argv) {
    if (argc < 2) {
        std::cerr << "Usage: log_analyser <logfile>\n";
        return 1;
    }
    std::ifstream file(argv[1]);
    if (!file) {
        std::cerr << "Could not open " << argv[1] << "\n";
        return 1;
    }
    std::vector<std::string> lines;
    std::string line;
    while (std::getline(file, line)) lines.push_back(line);

    Report r = analyse(lines);
    std::cout << "Total entries: " << r.total << "\n";
    std::cout << "Errors: " << r.errors << "\n";
    if (r.busiest_hour) {
        std::cout << "Busiest hour: " << r.busiest_hour->first << ":00 ("
                   << r.busiest_hour->second << " entries)\n";
    }
    std::cout << "Top error messages:\n";
    for (const auto &[msg, count] : r.top_errors) {
        std::cout << "  " << count << "x " << msg << "\n";
    }
    return 0;
}
⚠ No in-browser playground here
C++ compiles to a real, native binary, so unlike the Python version of this project there is no editor above you can run in the browser. Copy the code below and run it on your own machine — it takes seconds once a C++17-or-newer compiler like g++ or clang++ is installed.
What each part does — in plain words
std::optional<Entry> parse_line(const std::string &line) — a malformed line returns std::nullopt instead of an out-parameter plus a boolean return code, the same “this might not produce a value” pattern the basic-tier temperature-converter project used for an unrecognised conversion choice. The caller in analyse just writes if (!entry) continue;.

std::unordered_map<std::string, int> by_hour, by_message; — real hash tables, replacing the two hand-rolled linear-search Counted arrays the C version needed (C has no hash map at all). Counting becomes by_hour[entry->hour]++; — one line, average O(1), instead of a manual scan-and-insert loop.

top_n — the exact same vector-plus-std::sort pattern word-counter used in the basic tier: copy the map’s entries into a std::vector<std::pair<...>>, sort by count descending with alphabetical ties, and keep the first n. One small function, reused for both the busiest-hour ranking (n=1) and the top-3 error ranking.

std::optional<std::pair<std::string, int>> busiest_hour in Report — an empty log produces a report with no busiest hour at all, rather than a fabricated "00" or a separate boolean flag the C version needed (has_busiest_hour).
Common mistakes — and how to avoid them
✗ Iterating a std::unordered_map directly to find the busiest hour — its order is hash-bucket order, not count order, so the first entry you see is not necessarily the busiest.
✓ Copy the entries into a std::vector and sort, as top_n does here, exactly the lesson word-counter’s basic-tier project already taught.
✗ Assuming every line in a real log file is well-formed — a truncated write, a stray blank line, or a line from a different logger entirely will break a parser that does not check its assumptions.
✓ Return std::nullopt from parse_line for anything that does not look right, and skip it in analyse, as this project does.

4 Test & Prove Each Part

Seven checks: correct parsing of a well-formed line, rejecting a malformed one, the sort-with-ties behaviour of top_n, and the full analyse pipeline’s totals, busiest hour, top errors, and empty-input edge case.

A well-formed "date time LEVEL message" line parses into the correct hour, level, and message
A line with no valid date/time pair is rejected, not mis-parsed
top_n sorts by count descending, breaking ties alphabetically
analyse counts total valid entries and errors correctly, skipping garbage lines
analyse identifies the busiest hour
analyse ranks the top error messages by frequency
An empty log produces a zeroed report, not a crash
C++test_LogAnalyser.cpp
#include "LogAnalyser.hpp"
#include <cassert>
#include <iostream>

#define RUN(name) do { name(); std::cout << "PASS: " << #name << "\n"; } while (0)

static void parses_a_well_formed_line() {
    auto e = parse_line("2026-06-24 14:30:00 ERROR Database timeout");
    assert(e.has_value());
    assert(e->hour == "14");
    assert(e->level == "ERROR");
    assert(e->message == "Database timeout");
}

static void rejects_a_line_with_no_date_time_pair() {
    assert(!parse_line("this is not a log line").has_value());
    assert(!parse_line("notadate notatime LEVEL message").has_value());
}

static void top_n_sorts_by_count_descending_with_alphabetical_ties() {
    std::vector<std::pair<std::string, int>> counts = {
        {"zebra", 2}, {"apple", 2}, {"mango", 5}
    };
    auto top = top_n(counts, 2);
    assert(top[0].first == "mango" && top[0].second == 5);
    assert(top[1].first == "apple" && top[1].second == 2); // tie broken alphabetically
}

static void analyse_counts_totals_and_errors_correctly() {
    std::vector<std::string> lines = {
        "2026-06-24 09:00:00 INFO Server started",
        "2026-06-24 09:05:00 ERROR Database timeout",
        "2026-06-24 09:07:00 ERROR Database timeout",
        "2026-06-24 10:00:00 ERROR Disk full",
        "not a real log line at all",
    };
    Report r = analyse(lines);
    assert(r.total == 4); // the garbage line is skipped
    assert(r.errors == 3);
}

static void analyse_finds_the_busiest_hour() {
    std::vector<std::string> lines = {
        "2026-06-24 09:00:00 INFO a",
        "2026-06-24 09:05:00 INFO b",
        "2026-06-24 10:00:00 INFO c",
    };
    Report r = analyse(lines);
    assert(r.busiest_hour.has_value());
    assert(r.busiest_hour->first == "09");
    assert(r.busiest_hour->second == 2);
}

static void analyse_ranks_the_top_error_messages() {
    std::vector<std::string> lines = {
        "2026-06-24 09:00:00 ERROR Database timeout",
        "2026-06-24 09:01:00 ERROR Database timeout",
        "2026-06-24 09:02:00 ERROR Disk full",
    };
    Report r = analyse(lines);
    assert(!r.top_errors.empty());
    assert(r.top_errors[0].first == "Database timeout");
    assert(r.top_errors[0].second == 2);
}

static void an_empty_log_produces_a_zeroed_report_not_a_crash() {
    Report r = analyse({});
    assert(r.total == 0);
    assert(r.errors == 0);
    assert(!r.busiest_hour.has_value());
    assert(r.top_errors.empty());
}

int main() {
    RUN(parses_a_well_formed_line);
    RUN(rejects_a_line_with_no_date_time_pair);
    RUN(top_n_sorts_by_count_descending_with_alphabetical_ties);
    RUN(analyse_counts_totals_and_errors_correctly);
    RUN(analyse_finds_the_busiest_hour);
    RUN(analyse_ranks_the_top_error_messages);
    RUN(an_empty_log_produces_a_zeroed_report_not_a_crash);
    std::cout << "All tests passed.\n";
    return 0;
}

Compile and run with g++ -std=c++20 -Wall -Wextra -Wpedantic -o test_run LogAnalyser.cpp test_LogAnalyser.cpp && ./test_run.

5 The Interface

INPUTINPUTa path to a log file, given as a command-line argument
What it expects
$ ./loganalyser server.log
OUTPUTOUTPUTtotal/error counts, busiest hour, top 3 errors
What it returns
Total entries: 4
Errors: 3
Busiest hour: 09:00 (3 entries)
Top error messages:
  2x Database timeout
  1x Disk full

6 Run It & Automate It

Save the code as LogAnalyser.hpp / LogAnalyser.cpp / main.cpp and compile it with g++ — that turns your source directly into a native executable for your machine. No separate runtime needed: the compiled binary runs on its own.

Run it locally
g++ -std=c++20 -o loganalyser main.cpp LogAnalyser.cpp && ./loganalyser server.log
Point it at any log file shaped like “2026-06-24 09:05:00 ERROR Database timeout”.

A CI tool like Jenkins runs the same compile-then-test-then-check-for-leaks steps automatically whenever the code changes — every line below has a plain explanation.

What you should see when it works
Terminala real run
$ ./loganalyser server.log
Total entries: 4
Errors: 3
Busiest hour: 09:00 (3 entries)
Top error messages:
  2x Database timeout
  1x Disk full
If it breaks — how to fix it
🚨 Total entries: 0
Every line is being rejected by parse_line — check the log actually uses a literal space between date, time, level, and message, and that the date contains a - and the time contains a :.
🚨 terminate called after throwing std::invalid_argument
Uncaught from somewhere expecting a number — this project itself never calls std::stoi, so check any code you have added on top parses fields defensively.
GroovyJenkinsfile
// Jenkinsfile — compiles, tests, and checks for leaks on every change.
pipeline {
    agent any

    stages {
        stage('Get the code') {
            // download the latest code
            steps { checkout scm }
        }
        stage('Compile') {
            steps {
                // confirm a compiler is installed
                sh 'g++ --version'
                // compile with strict warnings on
                sh 'g++ -std=c++20 -Wall -Wextra -o app *.cpp'
            }
        }
        stage('Run the tests') {
            steps {
                // prints PASS/FAIL, exits non-zero on failure
                sh './app'
            }
        }
        stage('Check for memory leaks') {
            steps {
                // fails the build on any leak or invalid access
                sh 'valgrind --error-exitcode=1 --leak-check=full ./app'
            }
        }
    }

    post {
        success { echo 'All tests passed, no leaks found.' }
        failure { echo 'A test or Valgrind check failed — see above.' }
    }
}
🎯 Try this next — make it yours
  1. Report the busiest hour per level. Not just overall, but separately for ERROR, WARN, and INFO. (Teaches: a nested or compound map key.)
  2. Add a date range filter. Only analyse lines between two given dates. (Teaches: string comparison as a stand-in for date comparison, since the dates here sort correctly as plain text.)
  3. Stream instead of loading the whole file. Process line by line without holding a std::vector<std::string> of the whole file in memory. (Teaches: why analyse currently needs the whole file up front, and what would have to change.)
What you learned
You learned to reach for std::unordered_map a second time for real hash-table aggregation, reused the vector-plus-sort top_n pattern from word-counter, and used std::optional to report “no busiest hour” honestly for an empty log instead of a fabricated value or a separate boolean flag. Related: STL Containers, Modern C++ (C++11–C++23).