1 The Problem
We want a scraper that downloads a web page and pulls out specific pieces — titles, links, prices. It teaches fetching HTML and extracting data from it, plus the responsibility that comes with scraping (respecting sites and their rules).
2 How to Think About It
Think about fetch-then-extract, before any code:
3 The Build — explained part by part
Here is the complete scraper. Read each part’s note below — you should understand the whole thing from the notes alone.
import * as https from "node:https";
// extractLinks finds every <a href="..."> in a chunk of HTML. We scan for
// opening "<a" tags by hand (the same idea as Python's html.parser) rather
// than pulling in a full HTML/DOM library — good enough for simple pages,
// and it keeps this project dependency-free.
export function extractLinks(html: string): string[] {
const links: string[] = [];
const tagPattern = /<a\b[^>]*>/gi;
const hrefPattern = /href\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s>]+))/i;
for (const tagMatch of html.matchAll(tagPattern)) {
const hrefMatch = hrefPattern.exec(tagMatch[0]);
if (hrefMatch) {
links.push(hrefMatch[1] ?? hrefMatch[2] ?? hrefMatch[3]);
}
}
return links;
}
// fetch downloads a page's HTML. A polite scraper identifies itself with a
// User-Agent header.
export function fetch(url: string): Promise<string> {
return new Promise((resolve, reject) => {
const request = https.get(url, { headers: { "User-Agent": "CodexBot/1.0" } }, (response) => {
let data = "";
response.on("data", (chunk) => (data += chunk));
response.on("end", () => resolve(data));
});
request.on("error", reject);
});
}
async function main(): Promise<void> {
const html = await fetch("https://example.com");
for (const link of extractLinks(html)) {
console.log(link);
}
}
if (require.main === module) {
main();
}<a ...> tags with one regular expression and the href inside each match with a second — the same “good enough for simple pages” spirit as the Python version’s hand-rolled HTMLParser subclass, just without a class.html.matchAll(tagPattern) — returns every match of the pattern, not just the first (unlike
.match()), which is exactly what we need to collect every link on the page.export function fetch(url: string): Promise<string> — wraps Node’s callback-based
https.get in a Promise so it can be awaited like any other async call. Promise<string> is a generic type: “a promise that, when it resolves, hands you a string”.headers: { "User-Agent": "CodexBot/1.0" } — a polite scraper identifies itself, exactly as the Python version does.
<a href="...">link text</a> structure at once.<a> tag, a second pulls the href out of just that tag’s text. Trying to do it in one pattern gets fragile fast once attributes appear in a different order.href can be single-quoted, double-quoted, or even unquoted.hrefPattern above matches all three forms and picks whichever capture group actually matched.https.get in a Promise but forgetting to handle its 'error' event.'error', not the response callback — without request.on("error", reject), a failed request hangs forever instead of rejecting the promise.4 Test & Prove Each Part
How do we know this works? We pull the real logic into small, plain functions and check each one against cases we already know the answer to.
import { test } from "node:test";
import assert from "node:assert/strict";
import { extractLinks } from "./web-scraper";
test("a single anchor's href is extracted", () => {
assert.deepEqual(extractLinks('<a href="/page">Link</a>'), ["/page"]);
});
test("a page with no anchors returns empty", () => {
assert.deepEqual(extractLinks("<p>No links here</p>"), []);
});
test("all anchor hrefs are collected", () => {
const html = '<a href="/a">A</a><a href="/b">B</a>';
assert.deepEqual(extractLinks(html), ["/a", "/b"]);
});Compile with npx tsc then run node --test scraper.test.js. extractLinks takes a plain HTML string and returns a plain array, so the tests never make a real network request — only fetch (a thin, separately-reviewed wrapper around https.get) touches the network.
5 The Interface
What it expects
fetch("https://example.com")What it returns
/one
/two6 Run It & Automate It
Save the code as scraper.ts, compile with npx tsc, and run with node scraper.js — or run it directly with npx tsx scraper.ts. It needs an internet connection and a real URL to fetch.
npx tsc scraper.ts && node scraper.jsFetches the configured page and prints every link it finds, one per line.
A CI tool like Jenkins runs the type-checker and tests automatically whenever the code changes — every line below has a plain explanation.
/
/domains/1
/about/
/index.htmlgetaddrinfo ENOTFOUNDhrefPattern only looks inside <a> tags. Some pages link out through other elements (like <link> or JavaScript-inserted content) that a simple HTML scraper will not see — a real crawler would need a full HTML parser or a headless browser for those.// Jenkinsfile — runs the type-checker and tests automatically every time the code changes.
pipeline {
agent any // run on any available machine
stages {
stage('Get the code') {
steps { checkout scm } // download the latest code
}
stage('Set up Node') {
steps {
sh 'node --version' // confirm Node is installed
sh 'npm install -D typescript @types/node' // zero runtime deps — just the compiler and its Node types
}
}
stage('Type-check and test') {
steps {
sh 'npx tsc --noEmit' // catch type errors before anything runs
sh 'npx tsc' // compile to plain JavaScript
sh 'node --test scraper.test.js' // Node's built-in test runner, no extra install needed
}
}
}
post {
success { echo 'All tests passed.' }
failure { echo 'A test failed — look above.' }
}
}
You have a working web scraper. Extend it:
- Follow the links. Fetch each discovered link too, up to some depth. (Teaches: recursion with async functions.)
- Filter to one domain. Use the built-in
URLclass to ignore external links. (Teaches: resolving relative URLs against a base.) - Extract more than links. Pull out every
<img src="...">too, with a second pattern. (Teaches: generalising the extraction approach.) - Respect
robots.txt. Fetch and check it before scraping anything else. (Teaches: a real-world scraping courtesy, and morefetchcalls.)
Promise<T> so it can be awaited, why a network wrapper must handle its 'error' event, and a pragmatic regex-based alternative to a full HTML parser for simple link-scraping. Related reference: Typing Async Code, Generics.