← thecodex.expert · The Codex Family of Knowledge
Tier 1 · Beginner · Java Project

Word Counter

Read a text file and report words, characters, lines, and the most frequent words. Teaches NIO's Files API, a record for structured results, and HashMap-based frequency counting.

🧠 Teaches how to think spoonfed, every age Last verified:

1 The Problem

We want a tool that takes some text — typed in or read from a file — and reports how many words, characters, and lines it has. It teaches the core string operations for breaking text into pieces and measuring them.

Where this shows up: word-count limits on forms and essays, reading-time estimates, search indexing, text analysis, validating input length. Measuring and slicing text is one of the most common jobs in software.

2 How to Think About It

Think of the file as one long string to slice up three different ways: by whitespace (words), by length (characters), and by line.

The plan — in plain English
1. Read the whole file into a String with Files.readString. → 2. Split on whitespace to count words and lines. → 3. Build a frequency table with a HashMap. → 4. Report the counts and the top words.

Get the text

Characters = length

Words = split on spaces, count

Lines = split on newlines, count

Show all three

3 The Build — explained part by part

Here is the complete counter. A small record bundles the four counts together so the method returns one clear value instead of four loose numbers.

JavaWordCounter.java
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;

/**
 * Word Counter: reads a text file and reports line, word, and character
 * counts, plus the most frequent words.
 */
public class WordCounter {

    record Counts(int lines, int words, int chars, int bytes) {}

    static Counts countAll(String text) {
        int chars = text.length();
        int bytes = text.getBytes().length;
        String[] lines = text.isEmpty() ? new String[0] : text.split("\n", -1);
        int lineCount = text.isEmpty() ? 0 : lines.length - (text.endsWith("\n") ? 1 : 0);
        String[] words = text.trim().isEmpty() ? new String[0] : text.trim().split("\\s+");
        return new Counts(lineCount, words.length, chars, bytes);
    }

    static Map<String, Integer> wordFrequency(String text) {
        Map<String, Integer> freq = new LinkedHashMap<>();
        for (String raw : text.toLowerCase().split("\\s+")) {
            String word = raw.replaceAll("[^a-z0-9']", "");
            if (word.isEmpty()) continue;
            freq.merge(word, 1, Integer::sum);
        }
        return freq;
    }

    static List<Map.Entry<String, Integer>> topN(Map<String, Integer> freq, int n) {
        return freq.entrySet().stream()
                .sorted((a, b) -> b.getValue() - a.getValue())
                .limit(n)
                .toList();
    }

    public static void main(String[] args) throws IOException {
        if (args.length != 1) {
            System.out.println("Usage: java WordCounter <file>");
            return;
        }
        String text = Files.readString(Path.of(args[0]));
        Counts c = countAll(text);
        System.out.printf("Lines: %d  Words: %d  Characters: %d  Bytes: %d%n",
                c.lines(), c.words(), c.chars(), c.bytes());
        System.out.println("Top words:");
        for (var entry : topN(wordFrequency(text), 5)) {
            System.out.printf("  %-12s %d%n", entry.getKey(), entry.getValue());
        }
    }
}
⚠ No in-browser playground here
Java compiles to JVM bytecode and needs a real JDK to run, so unlike the Python version of this project there is no editor above you can run in the browser. Copy the code below and run it on your own machine — it takes seconds once a JDK is installed.
What each part does — in plain words
record Counts(int lines, int words, int chars, int bytes) {} — from the course’s Records lesson: one line gets a constructor, accessors, equals, and toString for free, which is exactly enough structure for a method that needs to return four related numbers at once.

Files.readString(Path.of(args[0])) — the modern java.nio.file API from the course’s IO and NIO lesson; it reads an entire file into a String in one call, replacing the old BufferedReader-in-a-loop pattern for anything that comfortably fits in memory.

text.length() vs text.getBytes().length — a Java String is UTF-16 internally, so .length() counts chars (UTF-16 code units, close enough to characters for this project) while .getBytes() re-encodes to UTF-8 bytes — the two numbers only match for plain ASCII text, the same byte-vs-character trap every other language on this site documents.

freq.merge(word, 1, Integer::sum) — merge inserts 1 for a new key or adds 1 to an existing value via the given function, all in one call — Java’s answer to Python’s dict.get(k, 0) + 1 or Rust’s entry().or_insert(0).
Common mistakes — and how to avoid them
✗ Using text.length() as “the word count” — that counts every character in the file, not the number of words.
✓ Split on whitespace first (text.trim().split("\\s+")) and count the resulting array’s length.
✗ Forgetting to lowercase before counting frequency — "The" and "the" would be tallied as two different words.
✓ Call .toLowerCase() before splitting, as wordFrequency does.

4 Test & Prove Each Part

We test the counting logic directly against in-memory strings, without touching a real file for most of the checks.

Lines and words are counted correctly for a two-line sample
An empty string counts as zero lines and zero words, not a crash
Word frequency is case-insensitive and strips punctuation
The top-N list returns the most frequent words first
JavaWordCounterTest.java
import org.junit.Test;
import java.util.List;
import java.util.Map;
import static org.junit.Assert.assertEquals;

public class WordCounterTest {

    @Test
    public void countsLinesWordsAndChars() {
        WordCounter.Counts c = WordCounter.countAll("hello world\nsecond line\n");
        assertEquals(2, c.lines());
        assertEquals(4, c.words());
    }

    @Test
    public void emptyTextCountsAsZero() {
        WordCounter.Counts c = WordCounter.countAll("");
        assertEquals(0, c.lines());
        assertEquals(0, c.words());
    }

    @Test
    public void wordFrequencyIsCaseInsensitiveAndStripsPunctuation() {
        Map<String, Integer> freq = WordCounter.wordFrequency("The fox. THE Fox! the fox");
        assertEquals(Integer.valueOf(3), freq.get("the"));
        assertEquals(Integer.valueOf(3), freq.get("fox"));
    }

    @Test
    public void topNReturnsMostFrequentFirst() {
        Map<String, Integer> freq = WordCounter.wordFrequency("a a a b b c");
        List<Map.Entry<String, Integer>> top = WordCounter.topN(freq, 2);
        assertEquals("a", top.get(0).getKey());
        assertEquals("b", top.get(1).getKey());
    }
}

Compile and run with javac -cp junit-4.13.2.jar and hamcrest-core-1.3.jar WordCounter.java WordCounterTest.java then java -cp .:junit-4.13.2.jar:hamcrest-core-1.3.jar org.junit.runner.JUnitCore WordCounterTest. Notice these tests never touch the filesystem — they pass plain strings straight to countAll and wordFrequency, which is why splitting file-reading out of the counting logic mattered.

5 The Interface

INPUTINPUTa text file path
What it expects
java WordCounter sample.txt
OUTPUTOUTPUTcounts and top words
What it returns
Lines: 2  Words: 10  Characters: 48  Bytes: 48
Top words:
  the           3
  fox           2

6 Run It & Automate It

Save the code as WordCounter.java and compile it with javac — that turns your source into .class bytecode files, which java then runs on the JVM. No separate install step: any real JDK ships both tools.

Run it locally
javac WordCounter.java && java WordCounter sample.txt
Create a sample.txt with a few lines of text first, in the same folder.

A CI tool like Jenkins runs the same compile-then-test steps automatically whenever the code changes — every line below has a plain explanation.

What you should see when it works
Terminala real run
$ printf "the quick brown fox\nthe lazy dog the fox jumped\n" > sample.txt
$ java WordCounter sample.txt
Lines: 2  Words: 10  Characters: 48  Bytes: 48
Top words:
  the          3
  fox          2
  quick        1
  brown        1
  lazy         1
If it breaks — how to fix it
🚨 java.nio.file.NoSuchFileException: sample.txt
The file was not found relative to the folder you ran java from. Check your current directory, or pass a full path as the argument.
🚨 The word count looks too high or too low.
Check whether the text has multiple spaces or tabs between words — split("\\s+") handles runs of whitespace correctly, but a single literal-space split would not.
GroovyJenkinsfile
// Jenkinsfile — runs the tests automatically every time the code changes.
pipeline {
    agent any                          // run on any available machine
    environment {
        CP = 'junit-4.13.2.jar:hamcrest-core-1.3.jar'   // JUnit + its one dependency
    }

    stages {
        stage('Get the code') {
            steps { checkout scm }     // download the latest code
        }
        stage('Set up JDK') {
            steps {
                sh 'java -version'           // confirm a JDK is installed
                sh 'javac -cp "$CP" *.java'   // compile the program and its tests together
            }
        }
        stage('Run the tests') {
            steps {
                sh 'java -cp ".:$CP" org.junit.runner.JUnitCore WordCounterTest'
            }
        }
    }

    post {
        success { echo 'All tests passed.' }
        failure { echo 'A test failed — look above.' }
    }
}
🎯 Try this next — make it yours
  1. Ignore common stop words. Skip "the", "a", "and" from the top-words list. (Teaches: filtering a stream before collecting.)
  2. Read from stdin too. Support piping text in when no filename is given. (Teaches: System.in as a fallback input source.)
  3. Count sentences. Split on ., !, and ?. (Teaches: regex alternation.)
What you learned
You learned NIO’s Files.readString for one-call file reading, a record for bundling related results, the UTF-16-chars-vs-UTF-8-bytes distinction in Java strings, and the merge method for one-line frequency counting. Related: IO and NIO, Collections and Generics.