1 The Problem
We want a log analyser: read a server log file, count how many entries are errors versus normal, find which hour had the most traffic, and list the most frequent error messages. It teaches parsing semi-structured text at scale and aggregating it into useful insight — a daily task in operations.
2 How to Think About It
Real logs have garbage lines mixed in. Design for that from the start: parsing a single line returns an Optional, never throws, and the rest of the program only ever sees successfully parsed entries.
3 The Build — explained part by part
Here is the complete analyser. parseLine returning Optional<Entry> instead of throwing is the load-bearing decision — it is what makes a malformed line an ordinary, expected outcome instead of a crash.
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
import java.util.Optional;
/**
* Log Analyser: parses "YYYY-MM-DD HH:MM:SS LEVEL message" log lines and
* reports level counts and the most frequent messages.
*/
public class LogAnalyser {
record Entry(String date, String time, String level, String message) {}
/** Returns empty for a line that doesn't structurally look like a log line, rather than guessing. */
static Optional<Entry> parseLine(String line) {
String[] parts = line.split(" ", 4);
if (parts.length < 4) return Optional.empty();
String date = parts[0], time = parts[1], level = parts[2], message = parts[3];
if (!date.contains("-") || !time.contains(":")) return Optional.empty();
return Optional.of(new Entry(date, time, level, message));
}
static List<Entry> parseAll(List<String> lines) {
return lines.stream().flatMap(l -> parseLine(l).stream()).toList();
}
static Map<String, Long> countByLevel(List<Entry> entries) {
Map<String, Long> counts = new LinkedHashMap<>();
for (Entry e : entries) {
counts.merge(e.level(), 1L, Long::sum);
}
return counts;
}
/** Reusable for any string-keyed frequency table, not just log messages. */
static <T> List<Map.Entry<T, Long>> topN(Map<T, Long> counts, int n) {
return counts.entrySet().stream()
.sorted((a, b) -> Long.compare(b.getValue(), a.getValue()))
.limit(n)
.toList();
}
static Map<String, Long> countByMessage(List<Entry> entries) {
Map<String, Long> counts = new LinkedHashMap<>();
for (Entry e : entries) {
counts.merge(e.message(), 1L, Long::sum);
}
return counts;
}
public static void main(String[] args) throws Exception {
if (args.length != 1) {
System.out.println("Usage: java LogAnalyser <file>");
return;
}
List<Entry> entries = parseAll(Files.readAllLines(Path.of(args[0])));
System.out.println("Parsed " + entries.size() + " entries.");
System.out.println("By level:");
countByLevel(entries).forEach((level, count) -> System.out.printf(" %-6s %d%n", level, count));
System.out.println("Top messages:");
for (var entry : topN(countByMessage(entries), 3)) {
System.out.printf(" %-30s %d%n", entry.getKey(), entry.getValue());
}
}
}
date.contains("-") and time.contains(":") catches lines that split into 4 parts by accident but are not actually log lines — the exact gap a naive splitn-based check would miss.lines.stream().flatMap(l -> parseLine(l).stream()) — an
Optional has a .stream() method that yields zero or one element, which is exactly what flatMap needs to turn “a stream of maybe-parsed lines” into “a stream of only the entries that parsed” in one line, with no explicit filtering step.static <T> List<Map.Entry<T, Long>> topN(Map<T, Long> counts, int n) — a generic method from the course’s Collections and Generics lesson: it works identically whether counting log levels, messages, or anything else keyed by any type
T, because nothing in its body assumes what T actually is.counts.merge(key, 1L, Long::sum) — the same one-call insert-or-increment pattern the word-counter project used, here counting by
Long instead of Integer since a log file can have far more lines than a word-counted text file.
parseLine throw on a malformed line — a single bad line anywhere in a large log file would crash the entire analysis.Optional.empty() for anything that does not structurally look right, as parseLine does, and skip it downstream.parts.length == 4 to validate a line — "not a real log line" also splits into exactly 4 parts by split(" ", 4), so length alone lets garbage through.parseLine does.4 Test & Prove Each Part
We test the parser on both well-formed and deliberately malformed input, and the counting logic in isolation.
import org.junit.Test;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import static org.junit.Assert.assertEquals;
import static org.junit.Assert.assertTrue;
public class LogAnalyserTest {
@Test
public void parsesAWellFormedLine() {
Optional<LogAnalyser.Entry> e = LogAnalyser.parseLine("2026-09-28 10:00:01 INFO Server started");
assertTrue(e.isPresent());
assertEquals("INFO", e.get().level());
assertEquals("Server started", e.get().message());
}
@Test
public void rejectsALineThatDoesNotLookLikeALogLine() {
assertTrue(LogAnalyser.parseLine("not a real log line").isEmpty());
assertTrue(LogAnalyser.parseLine("too short").isEmpty());
}
@Test
public void malformedLinesAreSkippedNotCrashed() {
List<String> lines = List.of(
"2026-09-28 10:00:01 INFO ok",
"garbage",
"2026-09-28 10:00:02 ERROR also ok");
List<LogAnalyser.Entry> entries = LogAnalyser.parseAll(lines);
assertEquals(2, entries.size());
}
@Test
public void countsByLevelCorrectly() {
List<LogAnalyser.Entry> entries = LogAnalyser.parseAll(List.of(
"2026-01-01 00:00:00 INFO a",
"2026-01-01 00:00:01 INFO b",
"2026-01-01 00:00:02 ERROR c"));
Map<String, Long> counts = LogAnalyser.countByLevel(entries);
assertEquals(Long.valueOf(2), counts.get("INFO"));
assertEquals(Long.valueOf(1), counts.get("ERROR"));
}
@Test
public void topNReturnsMostFrequentFirst() {
Map<String, Long> counts = Map.of("a", 5L, "b", 1L, "c", 3L);
List<Map.Entry<String, Long>> top = LogAnalyser.topN(counts, 2);
assertEquals("a", top.get(0).getKey());
assertEquals("c", top.get(1).getKey());
}
}
Compile and run with javac -cp junit-4.13.2.jar and hamcrest-core-1.3.jar LogAnalyser.java LogAnalyserTest.java then java -cp .:junit-4.13.2.jar:hamcrest-core-1.3.jar org.junit.runner.JUnitCore LogAnalyserTest. The malformed-line test is the important one: it feeds a genuinely garbage line into parseAll alongside good ones and checks the good ones still come through, rather than only testing parseLine in isolation.
5 The Interface
What it expects
java LogAnalyser server.logWhat it returns
Parsed 6 entries.
By level:
INFO 3
ERROR 2
Top messages:
Request received 26 Run It & Automate It
Save the code as LogAnalyser.java and compile it with javac — that turns your source into .class bytecode files, which java then runs on the JVM. No separate install step: any real JDK ships both tools.
javac LogAnalyser.java && java LogAnalyser sample.logCreate a sample.log with a few log lines first, in the same folder.
A CI tool like Jenkins runs the same compile-then-test steps automatically whenever the code changes — every line below has a plain explanation.
$ java LogAnalyser sample.log
Parsed 6 entries.
By level:
INFO 3
ERROR 2
WARN 1
Top messages:
Request received 2
Database timeout 2
Server started 1java from. Check your current directory or pass a full path.parseLine — open the file and check whether a line is missing its timestamp or level.// Jenkinsfile — runs the tests automatically every time the code changes.
pipeline {
agent any // run on any available machine
environment {
CP = 'junit-4.13.2.jar:hamcrest-core-1.3.jar' // JUnit + its one dependency
}
stages {
stage('Get the code') {
steps { checkout scm } // download the latest code
}
stage('Set up JDK') {
steps {
sh 'java -version' // confirm a JDK is installed
sh 'javac -cp "$CP" *.java' // compile the program and its tests together
}
}
stage('Run the tests') {
steps {
sh 'java -cp ".:$CP" org.junit.runner.JUnitCore LogAnalyserTest'
}
}
}
post {
success { echo 'All tests passed.' }
failure { echo 'A test failed — look above.' }
}
}
- Filter by time range. Only count entries between two timestamps. (Teaches: parsing the date/time fields into real
LocalDateTimevalues.) - Export to CSV. Write the level counts to a file instead of stdout. (Teaches: reusing the Files API for output, not just input.)
- Detect error bursts. Flag when 3+ ERROR lines happen within any 60-second window. (Teaches: a sliding-window algorithm over parsed timestamps.)
Optional for a parse that might legitimately fail, combining it with flatMap to filter a stream in one step, and writing a generic method that works for any key type without knowing what that type is. Related: Optional and Null Safety, Streams and Lambdas.