1 The Problem
We want a tool that takes some text — typed in or read from a file — and reports how many words, characters, and lines it has. It teaches the core string operations for breaking text into pieces and measuring them.
2 How to Think About It
Think of the file as one long string to slice up three different ways: by whitespace (words), by character (length), and by line.
String. → 2. Split on whitespace to count words. → 3. Count characters properly, not bytes. → 4. Build a frequency table with a HashMap and report the top words.
3 The Build — explained part by part
Here is the complete counter. Rust strings are guaranteed valid UTF-8, which is exactly why counting characters needs a specific method rather than just .len().
use std::collections::HashMap;
use std::env;
use std::fs;
/// Counts words the way `wc -w` does: whitespace-separated tokens.
fn count_words(text: &str) -> usize {
text.split_whitespace().count()
}
/// Counts characters, and separately notes the byte length, because a Rust
/// `String` is UTF-8 bytes: `text.len()` counts bytes, not characters, so a
/// word like "café" is 4 characters but 5 bytes. `.chars().count()` is what
/// actually answers "how many characters".
fn count_chars(text: &str) -> (usize, usize) {
(text.chars().count(), text.len())
}
/// Builds a frequency table of lowercased words, stripping simple
/// punctuation from each token's edges so "word." and "word" count together.
fn word_frequency(text: &str) -> HashMap<String, u32> {
let mut freq = HashMap::new();
for raw in text.split_whitespace() {
let word: String = raw
.trim_matches(|c: char| !c.is_alphanumeric())
.to_lowercase();
if word.is_empty() {
continue;
}
*freq.entry(word).or_insert(0) += 1;
}
freq
}
fn main() {
let path = match env::args().nth(1) {
Some(p) => p,
None => {
eprintln!("Usage: word_counter <file>");
return;
}
};
let text = match fs::read_to_string(&path) {
Ok(t) => t,
Err(e) => {
eprintln!("Could not read {path}: {e}");
return;
}
};
let (chars, bytes) = count_chars(&text);
println!("Words: {}", count_words(&text));
println!("Characters: {chars} ({bytes} bytes)");
println!("Lines: {}", text.lines().count());
let freq = word_frequency(&text);
let mut top: Vec<(&String, &u32)> = freq.iter().collect();
top.sort_by(|a, b| b.1.cmp(a.1).then(a.0.cmp(b.0)));
println!("Top 3 words:");
for (word, count) in top.into_iter().take(3) {
println!(" {word}: {count}");
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn counts_words_by_whitespace() {
assert_eq!(count_words("the quick brown fox"), 4);
assert_eq!(count_words(" extra spaces here "), 3);
}
#[test]
fn counts_characters_not_bytes_for_multibyte_text() {
let (chars, bytes) = count_chars("café");
assert_eq!(chars, 4);
assert_eq!(bytes, 5); // é is 2 bytes in UTF-8
}
#[test]
fn builds_a_case_insensitive_frequency_table() {
let freq = word_frequency("The cat sat. The Cat ran!");
assert_eq!(freq.get("the"), Some(&2));
assert_eq!(freq.get("cat"), Some(&2));
assert_eq!(freq.get("sat"), Some(&1));
}
#[test]
fn counts_lines_correctly() {
assert_eq!("one\ntwo\nthree".lines().count(), 3);
}
}
rustup) is installed.text.len() vs text.chars().count() —
len() returns the number of bytes in the string, and a multi-byte character like é is 2 bytes in UTF-8. chars().count() actually counts Unicode scalar values, which is what a person means by “characters.” This is the same trap Go’s len() sets, for the same UTF-8 reason.HashMap<String, u32> used as a counter, with *freq.entry(word).or_insert(0) += 1 — the
entry API looks up a key and, if it is missing, inserts a default first, all in one call — Rust’s answer to Python’s dict.get(k, 0) or Go’s comma-ok idiom, done atomically in a single expression.trim_matches(|c: char| !c.is_alphanumeric()) — strips punctuation from both ends of a word using a closure as the trim predicate, so
"word." and "word" count together.
text.len() to report “characters” — it silently reports bytes, which only matches character count for plain ASCII text.text.chars().count() whenever you mean actual characters.String by byte position expecting a character, e.g. &text[0..1] — this panics if that byte boundary falls in the middle of a multi-byte character..chars() instead of slicing by raw byte index when you care about characters..to_lowercase() before counting — "The" and "the" would be counted as two different words.word_frequency does.4 Test & Prove Each Part
We test each counting rule on small, known strings so the result can be checked by hand.
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn counts_words_by_whitespace() {
assert_eq!(count_words("the quick brown fox"), 4);
assert_eq!(count_words(" extra spaces here "), 3);
}
#[test]
fn counts_characters_not_bytes_for_multibyte_text() {
let (chars, bytes) = count_chars("café");
assert_eq!(chars, 4);
assert_eq!(bytes, 5); // é is 2 bytes in UTF-8
}
#[test]
fn builds_a_case_insensitive_frequency_table() {
let freq = word_frequency("The cat sat. The Cat ran!");
assert_eq!(freq.get("the"), Some(&2));
assert_eq!(freq.get("cat"), Some(&2));
assert_eq!(freq.get("sat"), Some(&1));
}
#[test]
fn counts_lines_correctly() {
assert_eq!("one\ntwo\nthree".lines().count(), 3);
}
}
Run with cargo test. The multi-byte test is the important one: it asserts chars=4, bytes=5 for “café”, proving the byte/character distinction in code rather than just describing it.
5 The Interface
What it expects
sample.txt (a plain text file, any length)What it returns
Words: 12
Characters: 60 (60 bytes)
Lines: 1
Top 3 words:
the: 3
dog: 2
barks: 16 Run It & Automate It
Save the code as src/main.rs inside a Cargo project's src/ folder and run it with cargo run — Cargo compiles and executes in one step while you are experimenting, then cargo build --release gives you an optimized binary once you are done.
cargo run -- sample.txtThe
-- tells Cargo everything after it is an argument to your program, not to Cargo itself.A CI tool like Jenkins runs cargo test automatically whenever the code changes — every line below has a plain explanation.
$ cargo run -- sample.txt
Words: 12
Characters: 60 (60 bytes)
Lines: 1
Top 3 words:
the: 3
dog: 2
barks: 1cargo run from, or pass a full path.--: cargo run -- sample.txt.// Jenkinsfile — runs the tests automatically every time the code changes.
pipeline {
agent any // run on any available machine
stages {
stage('Get the code') {
steps { checkout scm } // download the latest code
}
stage('Set up Rust') {
steps {
sh 'rustc --version' // confirm Rust is installed
sh 'cargo build' // compile, downloading any crates
}
}
stage('Run the tests') {
steps {
sh 'cargo clippy -- -D warnings' // catch obvious mistakes before running
sh 'cargo test' // run every test, show each result
}
}
}
post {
success { echo 'All tests passed.' }
failure { echo 'A test failed — look above.' }
}
}
- Report the longest word. Use
.max_by_key(|w| w.chars().count())on an iterator of words. (Teaches: iterator adapters.) - Stream large files. Switch to
BufReaderand read line by line instead of loading the whole file. (Teaches: not everything needs to fit in memory.) - Ignore common words. Skip “the”, “a”, “is” from the top-words report. (Teaches: a HashSet of stop words.)
String’s .len() counts bytes, not characters, and how to use a HashMap’s entry API to build a frequency counter in one line. Related: String Handling, Collections.