← thecodex.expert · The Codex Family of Knowledge
Tier 1 · Beginner · Rust Project

Word Counter

Read a text file and report words, characters, lines, and the most frequent words. Teaches file reading, HashMap-based counting, and the byte-vs-character distinction in Rust strings.

🧠 Teaches how to think spoonfed, every age Last verified:

1 The Problem

We want a tool that takes some text — typed in or read from a file — and reports how many words, characters, and lines it has. It teaches the core string operations for breaking text into pieces and measuring them.

Where this shows up: word-count limits on forms and essays, reading-time estimates, search indexing, text analysis, validating input length. Measuring and slicing text is one of the most common jobs in software.

2 How to Think About It

Think of the file as one long string to slice up three different ways: by whitespace (words), by character (length), and by line.

The plan — in plain English
1. Read the whole file into a String. → 2. Split on whitespace to count words. → 3. Count characters properly, not bytes. → 4. Build a frequency table with a HashMap and report the top words.

Get the text

Characters = length

Words = split on spaces, count

Lines = split on newlines, count

Show all three

3 The Build — explained part by part

Here is the complete counter. Rust strings are guaranteed valid UTF-8, which is exactly why counting characters needs a specific method rather than just .len().

Rustsrc/main.rs
use std::collections::HashMap;
use std::env;
use std::fs;

/// Counts words the way `wc -w` does: whitespace-separated tokens.
fn count_words(text: &str) -> usize {
    text.split_whitespace().count()
}

/// Counts characters, and separately notes the byte length, because a Rust
/// `String` is UTF-8 bytes: `text.len()` counts bytes, not characters, so a
/// word like "café" is 4 characters but 5 bytes. `.chars().count()` is what
/// actually answers "how many characters".
fn count_chars(text: &str) -> (usize, usize) {
    (text.chars().count(), text.len())
}

/// Builds a frequency table of lowercased words, stripping simple
/// punctuation from each token's edges so "word." and "word" count together.
fn word_frequency(text: &str) -> HashMap<String, u32> {
    let mut freq = HashMap::new();
    for raw in text.split_whitespace() {
        let word: String = raw
            .trim_matches(|c: char| !c.is_alphanumeric())
            .to_lowercase();
        if word.is_empty() {
            continue;
        }
        *freq.entry(word).or_insert(0) += 1;
    }
    freq
}

fn main() {
    let path = match env::args().nth(1) {
        Some(p) => p,
        None => {
            eprintln!("Usage: word_counter <file>");
            return;
        }
    };

    let text = match fs::read_to_string(&path) {
        Ok(t) => t,
        Err(e) => {
            eprintln!("Could not read {path}: {e}");
            return;
        }
    };

    let (chars, bytes) = count_chars(&text);
    println!("Words: {}", count_words(&text));
    println!("Characters: {chars} ({bytes} bytes)");
    println!("Lines: {}", text.lines().count());

    let freq = word_frequency(&text);
    let mut top: Vec<(&String, &u32)> = freq.iter().collect();
    top.sort_by(|a, b| b.1.cmp(a.1).then(a.0.cmp(b.0)));
    println!("Top 3 words:");
    for (word, count) in top.into_iter().take(3) {
        println!("  {word}: {count}");
    }
}

#[cfg(test)]
mod tests {
    use super::*;

    #[test]
    fn counts_words_by_whitespace() {
        assert_eq!(count_words("the quick brown fox"), 4);
        assert_eq!(count_words("  extra   spaces   here "), 3);
    }

    #[test]
    fn counts_characters_not_bytes_for_multibyte_text() {
        let (chars, bytes) = count_chars("café");
        assert_eq!(chars, 4);
        assert_eq!(bytes, 5); // é is 2 bytes in UTF-8
    }

    #[test]
    fn builds_a_case_insensitive_frequency_table() {
        let freq = word_frequency("The cat sat. The Cat ran!");
        assert_eq!(freq.get("the"), Some(&2));
        assert_eq!(freq.get("cat"), Some(&2));
        assert_eq!(freq.get("sat"), Some(&1));
    }

    #[test]
    fn counts_lines_correctly() {
        assert_eq!("one\ntwo\nthree".lines().count(), 3);
    }
}
⚠ No in-browser playground here
Rust compiles to a real binary, so unlike the Python version of this project there is no editor above you can run in the browser. Copy the code below and run it on your own machine — it takes seconds once Rust (via rustup) is installed.
What each part does — in plain words
text.split_whitespace().count() — splits on any run of whitespace (spaces, tabs, newlines) and collapses repeats automatically, unlike splitting on a single literal space.

text.len() vs text.chars().count() — len() returns the number of bytes in the string, and a multi-byte character like é is 2 bytes in UTF-8. chars().count() actually counts Unicode scalar values, which is what a person means by “characters.” This is the same trap Go’s len() sets, for the same UTF-8 reason.

HashMap<String, u32> used as a counter, with *freq.entry(word).or_insert(0) += 1 — the entry API looks up a key and, if it is missing, inserts a default first, all in one call — Rust’s answer to Python’s dict.get(k, 0) or Go’s comma-ok idiom, done atomically in a single expression.

trim_matches(|c: char| !c.is_alphanumeric()) — strips punctuation from both ends of a word using a closure as the trim predicate, so "word." and "word" count together.
Common mistakes — and how to avoid them
✗ Using text.len() to report “characters” — it silently reports bytes, which only matches character count for plain ASCII text.
✓ Use text.chars().count() whenever you mean actual characters.
✗ Indexing a String by byte position expecting a character, e.g. &text[0..1] — this panics if that byte boundary falls in the middle of a multi-byte character.
✓ Iterate with .chars() instead of slicing by raw byte index when you care about characters.
✗ Forgetting .to_lowercase() before counting — "The" and "the" would be counted as two different words.
✓ Normalize case before using a word as a HashMap key, as word_frequency does.

4 Test & Prove Each Part

We test each counting rule on small, known strings so the result can be checked by hand.

Words are counted by whitespace, collapsing extra spaces
Characters are counted correctly for multi-byte text like café
The frequency table is case-insensitive and strips punctuation
Lines are counted correctly
Rustsrc/main.rs (tests module)
#[cfg(test)]
mod tests {
    use super::*;

    #[test]
    fn counts_words_by_whitespace() {
        assert_eq!(count_words("the quick brown fox"), 4);
        assert_eq!(count_words("  extra   spaces   here "), 3);
    }

    #[test]
    fn counts_characters_not_bytes_for_multibyte_text() {
        let (chars, bytes) = count_chars("café");
        assert_eq!(chars, 4);
        assert_eq!(bytes, 5); // é is 2 bytes in UTF-8
    }

    #[test]
    fn builds_a_case_insensitive_frequency_table() {
        let freq = word_frequency("The cat sat. The Cat ran!");
        assert_eq!(freq.get("the"), Some(&2));
        assert_eq!(freq.get("cat"), Some(&2));
        assert_eq!(freq.get("sat"), Some(&1));
    }

    #[test]
    fn counts_lines_correctly() {
        assert_eq!("one\ntwo\nthree".lines().count(), 3);
    }
}

Run with cargo test. The multi-byte test is the important one: it asserts chars=4, bytes=5 for “café”, proving the byte/character distinction in code rather than just describing it.

5 The Interface

INPUTINPUTtext file path
What it expects
sample.txt (a plain text file, any length)
OUTPUTOUTPUTreport
What it returns
Words: 12
Characters: 60 (60 bytes)
Lines: 1
Top 3 words:
  the: 3
  dog: 2
  barks: 1

6 Run It & Automate It

Save the code as src/main.rs inside a Cargo project's src/ folder and run it with cargo run — Cargo compiles and executes in one step while you are experimenting, then cargo build --release gives you an optimized binary once you are done.

Run it locally
cargo run -- sample.txt
The -- tells Cargo everything after it is an argument to your program, not to Cargo itself.

A CI tool like Jenkins runs cargo test automatically whenever the code changes — every line below has a plain explanation.

What you should see when it works
Terminala real run
$ cargo run -- sample.txt
Words: 12
Characters: 60 (60 bytes)
Lines: 1
Top 3 words:
  the: 3
  dog: 2
  barks: 1
If it breaks — how to fix it
🚨 Could not read sample.txt: No such file or directory (os error 2)
Create the file in the same directory you run cargo run from, or pass a full path.
🚨 Usage: word_counter <file>
You ran the program with no arguments. Pass the file path after --: cargo run -- sample.txt.
GroovyJenkinsfile
// Jenkinsfile — runs the tests automatically every time the code changes.
pipeline {
    agent any                                  // run on any available machine

    stages {
        stage('Get the code') {
            steps { checkout scm }             // download the latest code
        }
        stage('Set up Rust') {
            steps {
                sh 'rustc --version'                // confirm Rust is installed
                sh 'cargo build'                     // compile, downloading any crates
            }
        }
        stage('Run the tests') {
            steps {
                sh 'cargo clippy -- -D warnings'     // catch obvious mistakes before running
                sh 'cargo test'                       // run every test, show each result
            }
        }
    }

    post {
        success { echo 'All tests passed.' }
        failure { echo 'A test failed — look above.' }
    }
}
🎯 Try this next — make it yours
  1. Report the longest word. Use .max_by_key(|w| w.chars().count()) on an iterator of words. (Teaches: iterator adapters.)
  2. Stream large files. Switch to BufReader and read line by line instead of loading the whole file. (Teaches: not everything needs to fit in memory.)
  3. Ignore common words. Skip “the”, “a”, “is” from the top-words report. (Teaches: a HashSet of stop words.)
What you learned
You learned why a Rust String’s .len() counts bytes, not characters, and how to use a HashMap’s entry API to build a frequency counter in one line. Related: String Handling, Collections.