1 The Problem
We want a scraper that downloads a web page and pulls out specific pieces — titles, links, prices. It teaches fetching HTML and extracting data from it, plus the responsibility that comes with scraping (respecting sites and their rules).
2 How to Think About It
Two separate jobs, kept in two separate methods: getting the raw HTML over the network, and pulling links out of text you already have. Only the first one needs a network at all.
HttpClient. → 2. Check the status code before trusting the body. → 3. Extract every href with a regular expression. → 4. Print what was found.
3 The Build — explained part by part
Here is the complete scraper. java.net.http, standard since Java 11, needs no external HTTP library at all — a genuine advantage over languages whose standard library ships no HTTP client.
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.util.ArrayList;
import java.util.List;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* Web Scraper: fetches a page over real HTTP with java.net.http and extracts
* every link from its HTML.
*/
public class WebScraper {
private static final Pattern HREF = Pattern.compile("href=\"([^\"]+)\"");
static String fetch(HttpClient client, String url) throws Exception {
HttpRequest request = HttpRequest.newBuilder(URI.create(url)).GET().build();
HttpResponse<String> response = client.send(request, HttpResponse.BodyHandlers.ofString());
if (response.statusCode() != 200) {
throw new RuntimeException("Unexpected status: " + response.statusCode());
}
return response.body();
}
static List<String> extractLinks(String html) {
List<String> links = new ArrayList<>();
Matcher m = HREF.matcher(html);
while (m.find()) {
links.add(m.group(1));
}
return links;
}
public static void main(String[] args) throws Exception {
if (args.length != 1) {
System.out.println("Usage: java WebScraper <url>");
return;
}
HttpClient client = HttpClient.newHttpClient();
String html = fetch(client, args[0]);
List<String> links = extractLinks(html);
System.out.println("Found " + links.size() + " links:");
links.forEach(link -> System.out.println(" " + link));
}
}
java.net.http, part of the standard library since Java 11, so this project needs no third-party HTTP library at all, unlike languages whose standard library ships no HTTP client of its own.response.statusCode() != 200 checked before returning the body — a 404 or 500 response still arrives as a normal, successful network exchange with no exception thrown; checking the status code explicitly is the only way to know the page actually loaded.
Pattern.compile("href=\"([^\"]+)\"") — a hand-written regular expression rather than a full HTML parser, which is a reasonable, honest trade for a learning project: it works for simple well-formed pages but is not a substitute for a real parser like Jsoup on messy real-world HTML.
fetch and extractLinks are separate, independently testable methods — the test suite proves the extraction logic works on plain strings with zero network involved, and separately proves the network call works against a real local server, rather than one large method doing both at once.
statusCode() before trusting body(), as fetch does.HttpServer inside the test itself, as fetchesAndExtractsLinksFromARealLocalServer does — still a genuine network round trip, just not a dependency on the outside world.4 Test & Prove Each Part
We test link extraction on plain strings, and separately prove the network call works, against a real server this test starts and stops itself.
import com.sun.net.httpserver.HttpServer;
import org.junit.Test;
import java.net.InetSocketAddress;
import java.net.http.HttpClient;
import java.util.List;
import static org.junit.Assert.assertEquals;
import static org.junit.Assert.assertTrue;
public class WebScraperTest {
@Test
public void extractsHrefsFromHtml() {
String html = "<a href=\"/one\">One</a><a href=\"/two\">Two</a>";
List<String> links = WebScraper.extractLinks(html);
assertEquals(List.of("/one", "/two"), links);
}
@Test
public void noLinksMeansEmptyListNotNull() {
assertTrue(WebScraper.extractLinks("<p>no links here</p>").isEmpty());
}
@Test
public void fetchesAndExtractsLinksFromARealLocalServer() throws Exception {
HttpServer server = HttpServer.create(new InetSocketAddress("localhost", 0), 0);
server.createContext("/", exchange -> {
byte[] body = "<a href=\"/about\">About</a><a href=\"/contact\">Contact</a>"
.getBytes();
exchange.sendResponseHeaders(200, body.length);
exchange.getResponseBody().write(body);
exchange.close();
});
server.start();
try {
int port = server.getAddress().getPort();
HttpClient client = HttpClient.newHttpClient();
String html = WebScraper.fetch(client, "http://localhost:" + port + "/");
List<String> links = WebScraper.extractLinks(html);
assertEquals(List.of("/about", "/contact"), links);
} finally {
server.stop(0);
}
}
@Test(expected = RuntimeException.class)
public void nonOkStatusThrows() throws Exception {
HttpServer server = HttpServer.create(new InetSocketAddress("localhost", 0), 0);
server.createContext("/", exchange -> {
exchange.sendResponseHeaders(404, -1);
exchange.close();
});
server.start();
try {
int port = server.getAddress().getPort();
HttpClient client = HttpClient.newHttpClient();
WebScraper.fetch(client, "http://localhost:" + port + "/");
} finally {
server.stop(0);
}
}
}
Compile and run with javac -cp junit-4.13.2.jar and hamcrest-core-1.3.jar WebScraper.java WebScraperTest.java then java -cp .:junit-4.13.2.jar:hamcrest-core-1.3.jar org.junit.runner.JUnitCore WebScraperTest. The last two tests use com.sun.net.httpserver.HttpServer — the same JDK-bundled server class the REST API project builds on — to start a real, ephemeral-port server inside the test itself, so the network call is genuine but the test never depends on any address outside localhost.
5 The Interface
What it expects
java WebScraper http://localhost:8099/index.htmlWhat it returns
Found 3 links:
/about
/contact
https://example.com6 Run It & Automate It
Save the code as WebScraper.java and compile it with javac — that turns your source into .class bytecode files, which java then runs on the JVM. No separate install step: any real JDK ships both tools.
javac WebScraper.java && java WebScraper http://localhost:8099/index.htmlPoint it at any page you are allowed to fetch, local or on the real internet.
A CI tool like Jenkins runs the same compile-then-test steps automatically whenever the code changes — every line below has a plain explanation.
$ python3 -m http.server 8099 &
$ java WebScraper http://localhost:8099/index.html
Found 3 links:
/about
/contact
https://example.comfetch working as designed — the URL you gave it does not exist on that server. Double-check the path.// Jenkinsfile — runs the tests automatically every time the code changes.
pipeline {
agent any // run on any available machine
environment {
CP = 'junit-4.13.2.jar:hamcrest-core-1.3.jar' // JUnit + its one dependency
}
stages {
stage('Get the code') {
steps { checkout scm } // download the latest code
}
stage('Set up JDK') {
steps {
sh 'java -version' // confirm a JDK is installed
sh 'javac -cp "$CP" *.java' // compile the program and its tests together
}
}
stage('Run the tests') {
steps {
sh 'java -cp ".:$CP" org.junit.runner.JUnitCore WebScraperTest'
}
}
}
post {
success { echo 'All tests passed.' }
failure { echo 'A test failed — look above.' }
}
}
- Follow links one level deep. Fetch every extracted link and report how many are reachable. (Teaches: recursive or iterative crawling, and cycle avoidance.)
- Extract more than links. Pull out
<title>and<img src>too. (Teaches: generalizing a regex-based extractor, and where it starts to strain.) - Add a real HTML parser. If you have network access, swap the regex for Jsoup. (Teaches: what a purpose-built parser handles that a regex cannot — nested tags, malformed HTML.)
java.net.http’s HttpClient/HttpRequest/HttpResponse trio built into the standard library, why a status code must be checked explicitly, and how to test real network code against a real local server instead of the live internet. Related: Standard Library, IO and NIO.