← thecodex.expert · The Codex Family of Knowledge
Tier 2 · Intermediate · TypeScript Project

Web Scraper

Fetch a web page and extract specific information from its HTML. Learn to pull structured data out of the messy web.

🧠 Teaches how to think spoonfed, every age Last verified:

1 The Problem

We want a scraper that downloads a web page and pulls out specific pieces — titles, links, prices. It teaches fetching HTML and extracting data from it, plus the responsibility that comes with scraping (respecting sites and their rules).

Where this shows up: price monitoring, news aggregation, research data collection, search engines, market analysis. When a site offers no API, scraping is how you get its data — carefully and respectfully.

2 How to Think About It

Think about fetch-then-extract, before any code:

The plan — in plain English
1. Download the page’s HTML. → 2. Search the raw text for the pattern you want. → 3. Extract the piece you need from each match. → Always: check the site allows it and do not hammer it with requests.

Download page HTML

Parse the HTML

Find target elements

Extract text or links

Use the data

3 The Build — explained part by part

Here is the complete scraper. Read each part’s note below — you should understand the whole thing from the notes alone.

TypeScriptscraper.ts
import * as https from "node:https";

// extractLinks finds every <a href="..."> in a chunk of HTML. We scan for
// opening "<a" tags by hand (the same idea as Python's html.parser) rather
// than pulling in a full HTML/DOM library — good enough for simple pages,
// and it keeps this project dependency-free.
export function extractLinks(html: string): string[] {
  const links: string[] = [];
  const tagPattern = /<a\b[^>]*>/gi;
  const hrefPattern = /href\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s>]+))/i;

  for (const tagMatch of html.matchAll(tagPattern)) {
    const hrefMatch = hrefPattern.exec(tagMatch[0]);
    if (hrefMatch) {
      links.push(hrefMatch[1] ?? hrefMatch[2] ?? hrefMatch[3]);
    }
  }
  return links;
}

// fetch downloads a page's HTML. A polite scraper identifies itself with a
// User-Agent header.
export function fetch(url: string): Promise<string> {
  return new Promise((resolve, reject) => {
    const request = https.get(url, { headers: { "User-Agent": "CodexBot/1.0" } }, (response) => {
      let data = "";
      response.on("data", (chunk) => (data += chunk));
      response.on("end", () => resolve(data));
    });
    request.on("error", reject);
  });
}

async function main(): Promise<void> {
  const html = await fetch("https://example.com");
  for (const link of extractLinks(html)) {
    console.log(link);
  }
}

if (require.main === module) {
  main();
}
⚠ No in-browser playground here
Running real, type-checked TypeScript in the browser needs either a full copy of the compiler or a third-party CDN script — the same kind of external dependency this site avoids relying on for a core teaching example. Copy the code below and run it with Node on your own machine instead; the “Run It” section explains exactly how.
What each part does — in plain words
export function extractLinks(html: string): string[] — rather than pulling in a full HTML-parsing library, we scan for <a ...> tags with one regular expression and the href inside each match with a second — the same “good enough for simple pages” spirit as the Python version’s hand-rolled HTMLParser subclass, just without a class.

html.matchAll(tagPattern) — returns every match of the pattern, not just the first (unlike .match()), which is exactly what we need to collect every link on the page.

export function fetch(url: string): Promise<string> — wraps Node’s callback-based https.get in a Promise so it can be awaited like any other async call. Promise<string> is a generic type: “a promise that, when it resolves, hands you a string”.

headers: { "User-Agent": "CodexBot/1.0" } — a polite scraper identifies itself, exactly as the Python version does.
Common mistakes — and how to avoid them
✗ Trying to parse HTML with a single regular expression that tries to match the whole <a href="...">link text</a> structure at once.
✓ Split the problem: one pattern finds each opening <a> tag, a second pulls the href out of just that tag’s text. Trying to do it in one pattern gets fragile fast once attributes appear in a different order.
✗ Forgetting that an href can be single-quoted, double-quoted, or even unquoted.
✓ The hrefPattern above matches all three forms and picks whichever capture group actually matched.
✗ Wrapping https.get in a Promise but forgetting to handle its 'error' event.
✓ A network failure fires 'error', not the response callback — without request.on("error", reject), a failed request hangs forever instead of rejecting the promise.

4 Test & Prove Each Part

How do we know this works? We pull the real logic into small, plain functions and check each one against cases we already know the answer to.

A single anchor’s href is extracted
A page with no anchors returns empty
All anchor hrefs are collected
TypeScriptscraper.test.ts
import { test } from "node:test";
import assert from "node:assert/strict";
import { extractLinks } from "./web-scraper";

test("a single anchor's href is extracted", () => {
  assert.deepEqual(extractLinks('<a href="/page">Link</a>'), ["/page"]);
});

test("a page with no anchors returns empty", () => {
  assert.deepEqual(extractLinks("<p>No links here</p>"), []);
});

test("all anchor hrefs are collected", () => {
  const html = '<a href="/a">A</a><a href="/b">B</a>';
  assert.deepEqual(extractLinks(html), ["/a", "/b"]);
});

Compile with npx tsc then run node --test scraper.test.js. extractLinks takes a plain HTML string and returns a plain array, so the tests never make a real network request — only fetch (a thin, separately-reviewed wrapper around https.get) touches the network.

5 The Interface

INPUTa URLthe page to scrape
What it expects
fetch("https://example.com")
OUTPUTextracted linksevery href on the page
What it returns
/one
/two

6 Run It & Automate It

Save the code as scraper.ts, compile with npx tsc, and run with node scraper.js — or run it directly with npx tsx scraper.ts. It needs an internet connection and a real URL to fetch.

Run it locally
npx tsc scraper.ts && node scraper.js
Fetches the configured page and prints every link it finds, one per line.

A CI tool like Jenkins runs the type-checker and tests automatically whenever the code changes — every line below has a plain explanation.

What you should see when it works
Terminala real run
/
/domains/1
/about/
/index.html
If it breaks — how to fix it
🚨 getaddrinfo ENOTFOUND
The URL does not resolve — check for a typo, or that you actually have an internet connection from wherever this is running.
🚨 No links print, even though the page clearly has some
The hrefPattern only looks inside <a> tags. Some pages link out through other elements (like <link> or JavaScript-inserted content) that a simple HTML scraper will not see — a real crawler would need a full HTML parser or a headless browser for those.
GroovyJenkinsfile
// Jenkinsfile &mdash; runs the type-checker and tests automatically every time the code changes.
pipeline {
    agent any                                  // run on any available machine

    stages {
        stage('Get the code') {
            steps { checkout scm }             // download the latest code
        }
        stage('Set up Node') {
            steps {
                sh 'node --version'                              // confirm Node is installed
                sh 'npm install -D typescript @types/node'       // zero runtime deps &mdash; just the compiler and its Node types
            }
        }
        stage('Type-check and test') {
            steps {
                sh 'npx tsc --noEmit'                 // catch type errors before anything runs
                sh 'npx tsc'                           // compile to plain JavaScript
                sh 'node --test scraper.test.js'            // Node's built-in test runner, no extra install needed
            }
        }
    }

    post {
        success { echo 'All tests passed.' }
        failure { echo 'A test failed &mdash; look above.' }
    }
}
🎯 Try this next — make it yours

You have a working web scraper. Extend it:

  1. Follow the links. Fetch each discovered link too, up to some depth. (Teaches: recursion with async functions.)
  2. Filter to one domain. Use the built-in URL class to ignore external links. (Teaches: resolving relative URLs against a base.)
  3. Extract more than links. Pull out every <img src="..."> too, with a second pattern. (Teaches: generalising the extraction approach.)
  4. Respect robots.txt. Fetch and check it before scraping anything else. (Teaches: a real-world scraping courtesy, and more fetch calls.)
What you learned
You learned to wrap a callback-based Node API in a Promise<T> so it can be awaited, why a network wrapper must handle its 'error' event, and a pragmatic regex-based alternative to a full HTML parser for simple link-scraping. Related reference: Typing Async Code, Generics.