← thecodex.expert · The Codex Family of Knowledge
Tier 3 · Upper-Intermediate · Rust Project

Web Scraper

Fetch a web page over a raw TCP socket and pull out every link, hand-written HTTP client and all. Teaches sockets, manual protocol handling, and dependency-free text parsing.

🧠 Teaches how to think spoonfed, every age Last verified:

1 The Problem

We want a scraper that downloads a web page and pulls out specific pieces — titles, links, prices. It teaches fetching HTML and extracting data from it, plus the responsibility that comes with scraping (respecting sites and their rules).

Where this shows up: price monitoring, news aggregation, research data collection, search engines, market analysis. When a site offers no API, scraping is how you get its data — carefully and respectfully.

2 How to Think About It

Strip away the crate that usually hides this: an HTTP GET is a socket connection, a few lines of text sent, and a response read back.

The plan — in plain English
1. Connect a TCP socket to the target host and port. → 2. Write a valid HTTP/1.1 request line and headers, ending in a blank line. → 3. Read the raw response text back. → 4. Split it into headers and body on the blank line HTTP itself defines. → 5. Scan the body for href="..." occurrences.

Download page HTML

Parse the HTML

Find target elements

Extract text or links

Use the data

3 The Build — explained part by part

Here is the complete scraper. This is the one project on this site where Rust’s sandbox constraint (no crates.io) turns into a genuine teaching opportunity rather than just a workaround: writing an HTTP client by hand shows you exactly what reqwest normally does for you.

Rustsrc/main.rs
use std::env;
use std::io::{Read, Write};
use std::net::TcpStream;

/// A minimal, hand-written HTTP/1.1 GET client over a raw `TcpStream`.
/// Idiomatic real-world Rust reaches for the `reqwest` crate here, but this
/// build environment cannot fetch crates.io, so this project does the thing
/// `reqwest` normally hides: open a socket, write the request line and
/// headers ourselves, and read the raw response back. It is more code, but
/// it is also a genuinely good look at what an HTTP client actually is
/// underneath — very much in the spirit of Rust's systems-programming roots.
fn fetch(host: &str, port: u16, path: &str) -> std::io::Result<String> {
    let mut stream = TcpStream::connect((host, port))?;
    let request = format!(
        "GET {path} HTTP/1.1\r\nHost: {host}\r\nUser-Agent: codex-scraper/1.0\r\nConnection: close\r\n\r\n"
    );
    stream.write_all(request.as_bytes())?;

    let mut raw = String::new();
    stream.read_to_string(&mut raw)?;
    Ok(raw)
}

/// Splits a raw HTTP/1.1 response into (status_line, headers, body). The
/// blank line (`\r\n\r\n`) is the boundary HTTP itself defines between
/// headers and body.
fn split_response(raw: &str) -> (&str, &str) {
    match raw.split_once("\r\n\r\n") {
        Some((head, body)) => (head, body),
        None => (raw, ""),
    }
}

/// Rust's standard library has no regex engine (that is the `regex` crate,
/// also unreachable here), so link extraction is a small hand-written
/// scanner: find every `href="..."` occurrence and pull out the quoted
/// text. It is less flexible than a real regex, but it is dependency-free,
/// entirely readable, and correct for well-formed HTML.
fn extract_links(html: &str) -> Vec<String> {
    let mut links = Vec::new();
    let mut rest = html;
    while let Some(start) = rest.find("href=\"") {
        rest = &rest[start + "href=\"".len()..];
        if let Some(end) = rest.find('"') {
            links.push(rest[..end].to_string());
            rest = &rest[end + 1..];
        } else {
            break;
        }
    }
    links
}

fn main() {
    let url = env::args().nth(1).unwrap_or_else(|| "http://127.0.0.1:8080/".to_string());
    let (host, port, path) = match parse_url(&url) {
        Some(parts) => parts,
        None => {
            eprintln!("Could not parse URL: {url} (expected http://host[:port]/path)");
            return;
        }
    };

    match fetch(&host, port, &path) {
        Ok(raw) => {
            let (head, body) = split_response(&raw);
            println!("{}", head.lines().next().unwrap_or(""));
            let links = extract_links(body);
            println!("Found {} link(s):", links.len());
            for link in &links {
                println!("  {link}");
            }
        }
        Err(e) => eprintln!("Request failed: {e}"),
    }
}

/// A tiny hand-written URL splitter — just enough for `http://host:port/path`.
/// A real project would use the `url` crate for this.
fn parse_url(url: &str) -> Option<(String, u16, String)> {
    let rest = url.strip_prefix("http://")?;
    let (authority, path) = match rest.find('/') {
        Some(i) => (&rest[..i], &rest[i..]),
        None => (rest, "/"),
    };
    let (host, port) = match authority.split_once(':') {
        Some((h, p)) => (h.to_string(), p.parse().ok()?),
        None => (authority.to_string(), 80),
    };
    Some((host, port, path.to_string()))
}

#[cfg(test)]
mod tests {
    use super::*;
    use std::io::BufReader;
    use std::net::TcpListener;
    use std::thread;

    #[test]
    fn extracts_links_from_html() {
        let html = r#"<a href="/about">About</a><a href="https://example.com">Ex</a>"#;
        let links = extract_links(html);
        assert_eq!(links, vec!["/about", "https://example.com"]);
    }

    #[test]
    fn returns_no_links_for_plain_text() {
        assert!(extract_links("just some text, no tags here").is_empty());
    }

    #[test]
    fn splits_headers_from_body_on_the_blank_line() {
        let raw = "HTTP/1.1 200 OK\r\nContent-Type: text/html\r\n\r\n<html>hi</html>";
        let (head, body) = split_response(raw);
        assert!(head.starts_with("HTTP/1.1 200 OK"));
        assert_eq!(body, "<html>hi</html>");
    }

    #[test]
    fn parses_a_simple_url() {
        let (host, port, path) = parse_url("http://example.com:9090/links").unwrap();
        assert_eq!(host, "example.com");
        assert_eq!(port, 9090);
        assert_eq!(path, "/links");
    }

    /// End-to-end: starts a real local TCP server on an OS-assigned port,
    /// serves one canned HTML response, and confirms `fetch` + `extract_links`
    /// pull the right links out of an actual network round trip, not just a
    /// hand-fed string.
    #[test]
    fn fetches_and_extracts_links_from_a_real_local_server() {
        let listener = TcpListener::bind("127.0.0.1:0").unwrap();
        let port = listener.local_addr().unwrap().port();

        let handle = thread::spawn(move || {
            let (stream, _) = listener.accept().unwrap();
            let mut reader = BufReader::new(stream.try_clone().unwrap());
            let mut request_line = String::new();
            std::io::BufRead::read_line(&mut reader, &mut request_line).unwrap();

            let body = r#"<html><body><a href="/one">One</a><a href="/two">Two</a></body></html>"#;
            let response = format!(
                "HTTP/1.1 200 OK\r\nContent-Length: {}\r\nConnection: close\r\n\r\n{}",
                body.len(),
                body
            );
            let mut stream = stream;
            stream.write_all(response.as_bytes()).unwrap();
        });

        let raw = fetch("127.0.0.1", port, "/").unwrap();
        handle.join().unwrap();

        let (_, body) = split_response(&raw);
        let links = extract_links(body);
        assert_eq!(links, vec!["/one", "/two"]);
    }
}
⚠ No in-browser playground here
Rust compiles to a real binary, so unlike the Python version of this project there is no editor above you can run in the browser. Copy the code below and run it on your own machine — it takes seconds once Rust (via rustup) is installed.
What each part does — in plain words
TcpStream::connect((host, port)) — opens a raw TCP socket, the same primitive every HTTP library in every language is eventually built on. There is no HTTP awareness at this layer at all — just bytes in, bytes out.

format!("GET {path} HTTP/1.1\r\nHost: {host}\r\n...\r\n\r\n") — a valid HTTP/1.1 request is a plain text protocol: a request line, headers each ending in \r\n, and a blank line marking the end of headers. Writing it out by hand is the entire “request” a crate like reqwest builds for you.

raw.split_once("\r\n\r\n") — that same blank line is how you find where headers end and the body begins in the response — HTTP defines this boundary explicitly, so no guessing is involved.

extract_links — Rust’s standard library ships no regex engine, so this is a hand-written scanner: repeatedly find href=", then find the closing quote, and slice out what is between them. Less flexible than a real regex, but dependency-free and fully readable.

the local-server test — rather than only testing extract_links on a hand-written string, one test starts a real TcpListener on an OS-assigned port, serves one canned response, and confirms fetch genuinely round-trips over a real socket — the same rigor Go’s version of this project used.
Common mistakes — and how to avoid them
✗ Forgetting Connection: close in the request headers — without it, a real server may keep the connection open waiting for another request, and read_to_string would then block forever waiting for the stream to end.
✓ Always send Connection: close for a one-shot client like this one, as the code above does.
✗ Assuming href="..." only ever uses double quotes — real-world HTML sometimes uses single quotes.
✓ A production-grade parser (or the regex/scraper crates, if reachable) would handle both; this project’s scanner is deliberately simple and documented as such.

4 Test & Prove Each Part

We test link extraction on known strings, the header/body split, the URL parser, and — the important one — a real fetch against a real local server.

Links are correctly extracted from a small HTML snippet
Plain text with no links returns an empty list
The response splitter correctly separates headers from body
A tiny URL like http://host:port/path parses into its three parts
fetch() against a real local TcpListener returns the links actually served
Rustsrc/main.rs (tests module)
#[cfg(test)]
mod tests {
    use super::*;
    use std::io::BufReader;
    use std::net::TcpListener;
    use std::thread;

    #[test]
    fn extracts_links_from_html() {
        let html = r#"<a href="/about">About</a><a href="https://example.com">Ex</a>"#;
        let links = extract_links(html);
        assert_eq!(links, vec!["/about", "https://example.com"]);
    }

    #[test]
    fn returns_no_links_for_plain_text() {
        assert!(extract_links("just some text, no tags here").is_empty());
    }

    #[test]
    fn splits_headers_from_body_on_the_blank_line() {
        let raw = "HTTP/1.1 200 OK\r\nContent-Type: text/html\r\n\r\n<html>hi</html>";
        let (head, body) = split_response(raw);
        assert!(head.starts_with("HTTP/1.1 200 OK"));
        assert_eq!(body, "<html>hi</html>");
    }

    #[test]
    fn parses_a_simple_url() {
        let (host, port, path) = parse_url("http://example.com:9090/links").unwrap();
        assert_eq!(host, "example.com");
        assert_eq!(port, 9090);
        assert_eq!(path, "/links");
    }

    /// End-to-end: starts a real local TCP server on an OS-assigned port,
    /// serves one canned HTML response, and confirms `fetch` + `extract_links`
    /// pull the right links out of an actual network round trip, not just a
    /// hand-fed string.
    #[test]
    fn fetches_and_extracts_links_from_a_real_local_server() {
        let listener = TcpListener::bind("127.0.0.1:0").unwrap();
        let port = listener.local_addr().unwrap().port();

        let handle = thread::spawn(move || {
            let (stream, _) = listener.accept().unwrap();
            let mut reader = BufReader::new(stream.try_clone().unwrap());
            let mut request_line = String::new();
            std::io::BufRead::read_line(&mut reader, &mut request_line).unwrap();

            let body = r#"<html><body><a href="/one">One</a><a href="/two">Two</a></body></html>"#;
            let response = format!(
                "HTTP/1.1 200 OK\r\nContent-Length: {}\r\nConnection: close\r\n\r\n{}",
                body.len(),
                body
            );
            let mut stream = stream;
            stream.write_all(response.as_bytes()).unwrap();
        });

        let raw = fetch("127.0.0.1", port, "/").unwrap();
        handle.join().unwrap();

        let (_, body) = split_response(&raw);
        let links = extract_links(body);
        assert_eq!(links, vec!["/one", "/two"]);
    }
}

Run with cargo test. The last test is the one worth reading closely: it spawns a thread running a real TcpListener, serves one hand-built HTTP response, and lets fetch connect to it exactly as it would to a real website — proving the socket code works over an actual network round trip, not just against a string.

5 The Interface

INPUTINPUTURL
What it expects
http://127.0.0.1:8099/index.html
OUTPUTOUTPUTstatus line + extracted links
What it returns
HTTP/1.0 200 OK
Found 2 link(s):
  /about
  https://example.com

6 Run It & Automate It

Save the code as src/main.rs inside a Cargo project's src/ folder and run it with cargo run — Cargo compiles and executes in one step while you are experimenting, then cargo build --release gives you an optimized binary once you are done.

Run it locally
cargo run -- http://127.0.0.1:8099/index.html
Point it at any plain HTTP (not HTTPS — this client has no TLS) server, including one you started locally with python3 -m http.server.

A CI tool like Jenkins runs cargo test automatically whenever the code changes — every line below has a plain explanation.

What you should see when it works
Terminala real run
$ python3 -m http.server 8099 &
$ cargo run -- http://127.0.0.1:8099/index.html
HTTP/1.0 200 OK
Found 2 link(s):
  /about
  https://example.com
If it breaks — how to fix it
🚨 Request failed: Connection refused (os error 111)
Nothing is listening on that host and port. Start a server there first, or check the port number.
🚨 Could not parse URL: ... (expected http://host[:port]/path)
This client only understands plain http:// URLs. An https:// URL needs TLS, which this hand-rolled client deliberately does not implement.
🚨 The program hangs and never prints anything.
Check the request includes Connection: close — without it, read_to_string can block forever waiting for a server that keeps the connection open.
GroovyJenkinsfile
// Jenkinsfile — runs the tests automatically every time the code changes.
pipeline {
    agent any                                  // run on any available machine

    stages {
        stage('Get the code') {
            steps { checkout scm }             // download the latest code
        }
        stage('Set up Rust') {
            steps {
                sh 'rustc --version'                // confirm Rust is installed
                sh 'cargo build'                     // compile, downloading any crates
            }
        }
        stage('Run the tests') {
            steps {
                sh 'cargo clippy -- -D warnings'     // catch obvious mistakes before running
                sh 'cargo test'                       // run every test, show each result
            }
        }
    }

    post {
        success { echo 'All tests passed.' }
        failure { echo 'A test failed — look above.' }
    }
}
🎯 Try this next — make it yours
  1. Follow redirects. Check for a 3xx status and a Location header, then fetch again. (Teaches: reading response headers, not just the body.)
  2. Use the real reqwest crate. If you have network access, compare how much of this code disappears. (Teaches: what a well-designed dependency buys you.)
  3. Add a timeout. Use TcpStream::set_read_timeout so a hung server cannot block forever. (Teaches: socket timeouts.)
What you learned
You learned what an HTTP request and response actually look like as raw bytes on a socket, by writing the client yourself instead of letting a crate hide it. You also saw a hand-written text scanner stand in for a regex engine, and tested it against a real local server rather than only against strings. Related: The Standard Library Tour, Testing.