← thecodex.expert · The Codex Family of Knowledge
Tier 2 · Intermediate · C++ Project

Web Scraper

Fetch a web page over a raw socket and pull out every link, hand-written HTTP client and all. C++ gives this project two genuine upgrades over its C counterpart: RAII-managed sockets and a real regex engine.

🧠 Teaches how to think spoonfed, every age Last verified:

1 The Problem

We want a scraper that downloads a web page and pulls out specific pieces — titles, links, prices. It teaches fetching HTML and extracting data from it, plus the responsibility that comes with scraping (respecting sites and their rules).

Where this shows up: price monitoring, news aggregation, research data collection, search engines, market analysis. When a site offers no API, scraping is how you get its data — carefully and respectfully.

2 How to Think About It

Strip away the library that usually hides this: an HTTP GET is a socket connection, a few lines of text sent, and a response read back. The two design questions worth pausing on are who closes the socket, and how links get pulled out of the HTML.

The plan — in plain English
1. Resolve the host with getaddrinfo and connect a TCP socket to it, wrapped in an RAII Socket. → 2. Write a valid HTTP/1.1 request line and headers, ending in a blank line. → 3. Read the raw response bytes back with recv. → 4. Split it into a status line and body on the blank line HTTP itself defines. → 5. Match every href="..." in the body with std::regex.

Download page HTML

Parse the HTML

Find target elements

Extract text or links

Use the data

3 The Build — explained part by part

Here is the complete scraper. C++ has no standard-library HTTP client either (unlike Go or Java) so writing one by hand over BSD sockets is still what an HTTP client looks like underneath any library — but two things below are genuinely nicer than the C version.

C++WebScraper.hpp / WebScraper.cpp / main.cpp
#pragma once
#include <optional>
#include <string>
#include <vector>

struct Url {
    std::string host;
    int port;
    std::string path;
};

// A tiny hand-written URL splitter -- just enough for
// "http://host[:port]/path". Returns std::nullopt on a malformed URL.
std::optional<Url> parse_url(const std::string &url);

// A small RAII wrapper around a raw BSD socket file descriptor. The
// destructor closes the socket automatically -- there is no equivalent in
// C, which must remember to call close() on every exit path by hand,
// including every early "return -1 on failure" branch. Move-only, since a
// socket handle should never be silently duplicated and closed twice.
class Socket {
public:
    Socket();
    explicit Socket(int fd);
    ~Socket();
    Socket(const Socket &) = delete;
    Socket &operator=(const Socket &) = delete;
    Socket(Socket &&other) noexcept;
    Socket &operator=(Socket &&other) noexcept;
    int fd() const { return fd_; }
    bool valid() const { return fd_ >= 0; }

private:
    int fd_ = -1;
};

// A minimal, hand-written HTTP/1.1 GET client over a raw socket. C++ has no
// standard-library HTTP client either (unlike Go's net/http or Java's
// java.net.http), so this project does by hand what any HTTP library does
// underneath: open a socket, write the request line and headers ourselves,
// and read the raw response back. Returns std::nullopt on failure.
std::optional<std::string> fetch(const std::string &host, int port, const std::string &path);

struct SplitResponse {
    std::string status_line;
    std::string body;
};

// Splits a raw HTTP/1.1 response into a status line and a body, at the
// blank line ("\r\n\r\n") HTTP itself defines as the boundary. Returns
// std::nullopt if no such boundary was found.
std::optional<SplitResponse> split_response(const std::string &raw);

// Unlike C, C++'s standard library does ship a real regex engine
// (<regex>), so link extraction can be a genuine pattern match for every
// href="..." occurrence instead of a hand-written character scanner.
std::vector<std::string> extract_links(const std::string &html);

#define _POSIX_C_SOURCE 200809L
#include "WebScraper.hpp"
#include <cstring>
#include <netdb.h>
#include <regex>
#include <sstream>
#include <sys/socket.h>
#include <unistd.h>
#include <utility>

// ---------- Socket ----------

Socket::Socket() : fd_(-1) {}
Socket::Socket(int fd) : fd_(fd) {}

Socket::~Socket() {
    if (fd_ >= 0) close(fd_);
}

Socket::Socket(Socket &&other) noexcept : fd_(other.fd_) {
    other.fd_ = -1;
}

Socket &Socket::operator=(Socket &&other) noexcept {
    if (this != &other) {
        if (fd_ >= 0) close(fd_);
        fd_ = other.fd_;
        other.fd_ = -1;
    }
    return *this;
}

// ---------- parse_url ----------

std::optional<Url> parse_url(const std::string &url) {
    std::string rest = url;
    const std::string prefix = "http://";
    if (rest.rfind(prefix, 0) != 0) return std::nullopt;
    rest = rest.substr(prefix.size());

    auto slash = rest.find('/');
    std::string host_port = slash == std::string::npos ? rest : rest.substr(0, slash);
    std::string path = slash == std::string::npos ? "/" : rest.substr(slash);
    if (host_port.empty()) return std::nullopt;

    int port = 80;
    std::string host = host_port;
    auto colon = host_port.find(':');
    if (colon != std::string::npos) {
        host = host_port.substr(0, colon);
        try {
            port = std::stoi(host_port.substr(colon + 1));
        } catch (const std::exception &) {
            return std::nullopt;
        }
    }
    if (host.empty()) return std::nullopt;
    return Url{host, port, path};
}

// ---------- fetch ----------

std::optional<std::string> fetch(const std::string &host, int port, const std::string &path) {
    addrinfo hints{};
    hints.ai_family = AF_INET;
    hints.ai_socktype = SOCK_STREAM;
    addrinfo *result = nullptr;

    std::string port_str = std::to_string(port);
    if (getaddrinfo(host.c_str(), port_str.c_str(), &hints, &result) != 0) {
        return std::nullopt;
    }

    Socket sock(socket(result->ai_family, result->ai_socktype, result->ai_protocol));
    if (!sock.valid()) {
        freeaddrinfo(result);
        return std::nullopt;
    }
    if (connect(sock.fd(), result->ai_addr, result->ai_addrlen) < 0) {
        freeaddrinfo(result);
        return std::nullopt;
    }
    freeaddrinfo(result);

    std::ostringstream request;
    request << "GET " << path << " HTTP/1.1\r\n"
             << "Host: " << host << "\r\n"
             << "Connection: close\r\n"
             << "\r\n";
    std::string req = request.str();
    if (send(sock.fd(), req.c_str(), req.size(), 0) < 0) return std::nullopt;

    std::string response;
    char buf[4096];
    ssize_t n;
    while ((n = recv(sock.fd(), buf, sizeof(buf), 0)) > 0) {
        response.append(buf, static_cast<size_t>(n));
    }
    return response;
}

// ---------- split_response ----------

std::optional<SplitResponse> split_response(const std::string &raw) {
    auto boundary = raw.find("\r\n\r\n");
    if (boundary == std::string::npos) return std::nullopt;
    auto line_end = raw.find("\r\n");
    std::string status_line = raw.substr(0, line_end);
    std::string body = raw.substr(boundary + 4);
    return SplitResponse{status_line, body};
}

// ---------- extract_links ----------

std::vector<std::string> extract_links(const std::string &html) {
    static const std::regex href_re(R"re(href="([^"]*)")re");
    std::vector<std::string> links;
    auto begin = std::sregex_iterator(html.begin(), html.end(), href_re);
    auto end = std::sregex_iterator();
    for (auto it = begin; it != end; ++it) {
        links.push_back((*it)[1].str());
    }
    return links;
}

#include "WebScraper.hpp"
#include <iostream>

int main(int argc, char **argv) {
    if (argc < 2) {
        std::cerr << "Usage: web_scraper <url>\n";
        return 1;
    }
    auto url = parse_url(argv[1]);
    if (!url) {
        std::cerr << "Could not parse URL (expected http://host[:port]/path)\n";
        return 1;
    }
    auto raw = fetch(url->host, url->port, url->path);
    if (!raw) {
        std::cerr << "Could not fetch " << argv[1] << "\n";
        return 1;
    }
    auto split = split_response(*raw);
    if (!split) {
        std::cerr << "Malformed HTTP response\n";
        return 1;
    }
    std::cout << split->status_line << "\n";
    auto links = extract_links(split->body);
    std::cout << "Found " << links.size() << " link(s):\n";
    for (const auto &link : links) {
        std::cout << "  " << link << "\n";
    }
    return 0;
}
⚠ No in-browser playground here
C++ compiles to a real, native binary, so unlike the Python version of this project there is no editor above you can run in the browser. Copy the code below and run it on your own machine — it takes seconds once a C++17-or-newer compiler like g++ or clang++ is installed.
What each part does — in plain words
class Socket — a small RAII wrapper around the raw file descriptor. Its destructor calls close() automatically, on every exit path, including an early return std::nullopt from fetch. C’s version of this project has to remember to call close(fd) by hand on every one of those same early-return branches — miss one, and that is a leaked file descriptor. Socket is move-only (copy is deleted) since a socket handle should never be silently duplicated and closed twice.

std::regex href_re(R"re(href=\"([^\"]*)\")re"); — unlike C, C++’s standard library ships a real regex engine in <regex>. extract_links can match every href="..." occurrence and capture the quoted text directly, instead of hand-writing a strstr/strchr scanner the way the C version had to (even though C’s own libc separately ships POSIX <regex.h>, this project’s C version chose not to use it, for reasons explained on that page).

getaddrinfo(host.c_str(), port_str.c_str(), &hints, &result) — the same modern, IPv4/IPv6-agnostic hostname resolution C uses, called identically from C++ since it is a POSIX C API with no C++ standard-library equivalent.

raw.find("\r\n\r\n") — std::string::find locates the same blank-line boundary HTTP itself defines between headers and body, the C++ equivalent of C’s strstr.
Common mistakes — and how to avoid them
✗ Letting a Socket be copied instead of moved — two copies with the same fd would both try to close it, the second call operating on an already-closed (or worse, reused) descriptor.
✓ Delete the copy constructor and copy-assignment operator, as this project does, so the compiler catches an accidental copy at build time.
✗ Writing a raw string literal like R"(href="([^"]*)")" for a pattern that itself ends in )" — the raw string terminates at the first )" it finds, silently truncating the pattern and breaking the build with a confusing error far from the real cause.
✓ Use a custom raw-string delimiter, as this project does with R"re(...)re", whenever the pattern itself might contain )".

4 Test & Prove Each Part

Eight checks, including a real end-to-end test that starts an actual HTTP server on a background std::thread bound to an OS-assigned port, then has fetch() talk to it over a real socket — no mocking, the same discipline the C version used with a background pthread.

parse_url splits a URL into host, port, and path correctly
parse_url defaults to port 80 and path "/" when they are omitted
parse_url rejects a non-http:// scheme or garbage input
split_response finds the blank-line boundary between headers and body
split_response rejects a response with no blank line
extract_links finds every href, including a mix of relative and absolute URLs
extract_links returns an empty list for HTML with no links
fetch and extract_links work end-to-end against a real local server
C++test_WebScraper.cpp
#define _POSIX_C_SOURCE 200809L
#include "WebScraper.hpp"
#include <arpa/inet.h>
#include <cassert>
#include <cstring>
#include <iostream>
#include <netinet/in.h>
#include <sys/socket.h>
#include <thread>
#include <unistd.h>

#define RUN(name) do { name(); std::cout << "PASS: " << #name << "\n"; } while (0)

static void parse_url_splits_host_port_and_path() {
    auto u = parse_url("http://example.com:8080/path/to/page");
    assert(u.has_value());
    assert(u->host == "example.com");
    assert(u->port == 8080);
    assert(u->path == "/path/to/page");
}

static void parse_url_defaults_to_port_80_and_root_path() {
    auto u = parse_url("http://example.com");
    assert(u.has_value());
    assert(u->port == 80);
    assert(u->path == "/");
}

static void parse_url_rejects_a_non_http_scheme() {
    assert(!parse_url("ftp://example.com/").has_value());
    assert(!parse_url("not a url at all").has_value());
}

static void split_response_finds_the_blank_line_boundary() {
    std::string raw = "HTTP/1.1 200 OK\r\nContent-Type: text/html\r\n\r\n<html>body</html>";
    auto s = split_response(raw);
    assert(s.has_value());
    assert(s->status_line == "HTTP/1.1 200 OK");
    assert(s->body == "<html>body</html>");
}

static void split_response_rejects_a_response_with_no_blank_line() {
    assert(!split_response("HTTP/1.1 200 OK\r\nno blank line here").has_value());
}

static void extract_links_finds_every_href() {
    std::string html = R"(<a href="/one">One</a><a href="https://two.example/">Two</a>)";
    auto links = extract_links(html);
    assert(links.size() == 2);
    assert(links[0] == "/one");
    assert(links[1] == "https://two.example/");
}

static void extract_links_returns_empty_for_html_with_no_links() {
    assert(extract_links("<html><body>No links here</body></html>").empty());
}

// A real end-to-end test: start a tiny HTTP server on a background thread
// bound to an OS-assigned port, then have fetch() talk to it over an
// actual socket -- no mocking, the same discipline the C version of this
// project used with a background pthread.
static void fetch_and_extract_links_work_against_a_real_local_server() {
    int server_fd = socket(AF_INET, SOCK_STREAM, 0);
    assert(server_fd >= 0);
    int opt = 1;
    setsockopt(server_fd, SOL_SOCKET, SO_REUSEADDR, &opt, sizeof(opt));

    sockaddr_in addr{};
    addr.sin_family = AF_INET;
    addr.sin_addr.s_addr = INADDR_ANY;
    addr.sin_port = 0; // let the OS pick a free port
    assert(bind(server_fd, reinterpret_cast<sockaddr *>(&addr), sizeof(addr)) == 0);
    assert(listen(server_fd, 1) == 0);

    socklen_t len = sizeof(addr);
    getsockname(server_fd, reinterpret_cast<sockaddr *>(&addr), &len);
    int port = ntohs(addr.sin_port);

    std::thread server([server_fd]() {
        int client = accept(server_fd, nullptr, nullptr);
        if (client < 0) return;
        char buf[1024];
        recv(client, buf, sizeof(buf), 0); // drain the request, ignore its contents
        std::string body = "<html><body><a href=\"/page1\">One</a><a href=\"/page2\">Two</a></body></html>";
        std::string response = "HTTP/1.1 200 OK\r\nContent-Type: text/html\r\nContent-Length: " +
                                std::to_string(body.size()) + "\r\n\r\n" + body;
        send(client, response.c_str(), response.size(), 0);
        close(client);
    });

    auto raw = fetch("127.0.0.1", port, "/");
    server.join();
    close(server_fd);

    assert(raw.has_value());
    auto split = split_response(*raw);
    assert(split.has_value());
    assert(split->status_line == "HTTP/1.1 200 OK");
    auto links = extract_links(split->body);
    assert(links.size() == 2);
    assert(links[0] == "/page1");
    assert(links[1] == "/page2");
}

int main() {
    RUN(parse_url_splits_host_port_and_path);
    RUN(parse_url_defaults_to_port_80_and_root_path);
    RUN(parse_url_rejects_a_non_http_scheme);
    RUN(split_response_finds_the_blank_line_boundary);
    RUN(split_response_rejects_a_response_with_no_blank_line);
    RUN(extract_links_finds_every_href);
    RUN(extract_links_returns_empty_for_html_with_no_links);
    RUN(fetch_and_extract_links_work_against_a_real_local_server);
    std::cout << "All tests passed.\n";
    return 0;
}

Compile and run with g++ -std=c++20 -Wall -Wextra -Wpedantic -pthread -o test_run WebScraper.cpp test_WebScraper.cpp && ./test_run. The -pthread flag is required because the end-to-end test spins up a real server on a std::thread.

5 The Interface

INPUTINPUTa URL, given as a command-line argument
What it expects
$ ./scraper http://example.com/page.html
OUTPUTOUTPUTthe HTTP status line and every link found
What it returns
HTTP/1.1 200 OK
Found 2 link(s):
  /foo
  /bar

6 Run It & Automate It

Save the code as WebScraper.hpp / WebScraper.cpp / main.cpp and compile it with g++ — that turns your source directly into a native executable for your machine. No separate runtime needed: the compiled binary runs on its own.

Run it locally
g++ -std=c++20 -pthread -o scraper main.cpp WebScraper.cpp && ./scraper http://example.com/
Point it at any plain HTTP (not HTTPS — this project does not implement TLS) URL.

A CI tool like Jenkins runs the same compile-then-test-then-check-for-leaks steps automatically whenever the code changes — every line below has a plain explanation.

What you should see when it works
Terminala real run
$ ./scraper http://127.0.0.1:8199/index.html
HTTP/1.0 200 OK
Found 2 link(s):
  /foo
  /bar
If it breaks — how to fix it
🚨 Could not fetch http://example.com/
This project speaks plain HTTP only, not HTTPS — a URL that redirects to https:// (almost every real site now does) will fail to connect on port 80 the way this expects.
🚨 Found 0 link(s), but the page clearly has links
Check the page actually reached the body: print split->body before calling extract_links. Many sites gzip-compress their response, and this project does not decompress Content-Encoding: gzip.
GroovyJenkinsfile
// Jenkinsfile — compiles, tests, and checks for leaks on every change.
pipeline {
    agent any

    stages {
        stage('Get the code') {
            // download the latest code
            steps { checkout scm }
        }
        stage('Compile') {
            steps {
                // confirm a compiler is installed
                sh 'g++ --version'
                // compile with strict warnings on
                sh 'g++ -std=c++20 -Wall -Wextra -o app *.cpp -pthread'
            }
        }
        stage('Run the tests') {
            steps {
                // prints PASS/FAIL, exits non-zero on failure
                sh './app'
            }
        }
        stage('Check for memory leaks') {
            steps {
                // fails the build on any leak or invalid access
                sh 'valgrind --error-exitcode=1 --leak-check=full ./app'
            }
        }
    }

    post {
        success { echo 'All tests passed, no leaks found.' }
        failure { echo 'A test or Valgrind check failed — see above.' }
    }
}
🎯 Try this next — make it yours
  1. Follow redirects. A 301/302 response includes a Location header — fetch that URL next. (Teaches: recursive or looped fetching, and a depth limit to avoid an infinite redirect loop.)
  2. Extract image sources too. Match src="..." on <img> tags with a second std::regex. (Teaches: composing more than one pattern over the same text.)
  3. Resolve relative links to absolute URLs. /foo found on http://example.com/page should become http://example.com/foo. (Teaches: URL-joining logic most HTTP libraries hide from you.)
What you learned
You learned to wrap a raw OS resource (a socket file descriptor) in a small RAII class so it closes itself on every exit path, and used std::regex for a genuine pattern match instead of a hand-written character scanner — both direct upgrades over this project’s C version. Related: Classes and RAII, STL Algorithms.