1 The Problem
We want a scraper that downloads a web page and pulls out specific pieces — titles, links, prices. It teaches fetching HTML and extracting data from it, plus the responsibility that comes with scraping (respecting sites and their rules).
2 How to Think About It
Think about fetch-then-extract, before any code:
3 The Build — explained part by part
Here is the complete scraper. Read each part’s note below — you should understand the whole thing from the notes alone.
import java.net.URI
import java.net.http.HttpClient
import java.net.http.HttpRequest
import java.net.http.HttpResponse
private val HREF = Regex("href=\"([^\"]+)\"")
/** Fetches [url] over real HTTP and returns the response body as text. */
fun fetch(client: HttpClient, url: String): String {
val request = HttpRequest.newBuilder(URI.create(url)).GET().build()
val response = client.send(request, HttpResponse.BodyHandlers.ofString())
check(response.statusCode() == 200) { "Unexpected status: ${response.statusCode()}" }
return response.body()
}
/** Every href="..." target found in [html], in the order they appear. */
fun extractLinks(html: String): List<String> =
HREF.findAll(html).map { it.groupValues[1] }.toList()
fun main(args: Array<String>) {
if (args.size != 1) {
println("Usage: kotlin WebScraperKt <url>")
return
}
val client = HttpClient.newHttpClient()
val html = fetch(client, args[0])
val links = extractLinks(html)
println("Found ${links.size} links:")
links.forEach { println(" $it") }
}kotlinc on your own machine instead; the “Run It” section explains exactly how.HttpRequest.newBuilder(URI.create(url)).GET().build() — the client is built once; each request is its own small, immutable object describing exactly one call.
check(response.statusCode() == 200) { "..." } — Kotlin's
check throws an IllegalStateException with the given message if the condition is false, a concise one-line alternative to a multi-line if (...) throw ....val HREF = Regex("href=\"([^\"]+)\"") — a simple, pragmatic pattern, not a full HTML parser;
HREF.findAll(html).map { it.groupValues[1] } walks every match and pulls out just the captured URL.
HttpClient.send; always check response.statusCode() before trusting the body.href="..." (as here), but a real crawler handling arbitrary pages should use a proper HTML parser — a regex cannot reliably handle nested tags or attributes split across lines.HttpClient.send throws on a connection failure (DNS, refused connection, timeout).try/catch around the call and a status check right after it.4 Test & Prove Each Part
How do we know this works? We pull the real logic into small, plain functions and check each one against cases we already know the answer to.
import kotlin.test.Test
import kotlin.test.assertEquals
class WebScraperTest {
@Test
fun findsEveryLink() {
val html = """<a href="https://a.com">A</a> <a href="/b">B</a>"""
assertEquals(listOf("https://a.com", "/b"), extractLinks(html))
}
@Test
fun noLinksIsAnEmptyList() {
assertEquals(emptyList(), extractLinks("<p>nothing here</p>"))
}
@Test
fun ignoresOtherAttributes() {
val html = """<img src="pic.png"> <a href="/ok">ok</a>"""
assertEquals(listOf("/ok"), extractLinks(html))
}
}Compile with kotlinc webScraper.kt webScraperTest.kt -include-runtime -d webScraper.jar and run with JUnit's own runner. extractLinks takes a plain String of HTML, so the tests check known snippets directly — no real network call or running server involved.
5 The Interface
What it expects
fetch("https://example.com")What it returns
/one
/two6 Run It & Automate It
Save the code as webScraper.kt and compile it with kotlinc webScraper.kt -include-runtime -d webScraper.jar. It takes the URL to fetch as a command-line argument.
kotlinc webScraper.kt -include-runtime -d webScraper.jar && java -jar webScraper.jar https://example.comFetches the given URL and prints every link found on the page.
A CI tool like Jenkins compiles and tests automatically whenever the code changes — every line below has a plain explanation.
Found 3 links:
https://example.com
/about
/contactUsage: kotlin WebScraperKt <url>java -jar webScraper.jar https://example.com, with the full URL including https://.java.net.ConnectException or the program hangs for a long time// Jenkinsfile — compiles and tests automatically every time the code changes.
pipeline {
agent any // run on any available machine
stages {
stage('Get the code') {
steps { checkout scm } // download the latest code
}
stage('Set up Kotlin') {
steps {
sh 'kotlinc -version' // confirm the compiler is installed
}
}
stage('Compile and test') {
steps {
sh 'kotlinc webScraper.kt webScraperTest.kt -include-runtime -d build.jar' // one real JVM jar, no build tool required
sh 'java -cp build.jar:kotlin-test-junit.jar:junit.jar org.junit.runner.JUnitCore WebScraperTest'
}
}
}
post {
success { echo 'All tests passed.' }
failure { echo 'A test failed — look above.' }
}
}
You have a working web scraper. Extend it:
- Follow the links. Fetch each discovered link too, up to some depth. (Teaches: recursion, or an explicit work queue.)
- Filter to one domain. Use
java.net.URIto resolve relative links and ignore external ones. (Teaches: resolving a relative URL against a base.) - Extract more than links. Pull out every
<img src="...">too, with a second pattern. (Teaches: generalising the extraction approach.) - Respect
robots.txt. Fetch and check it before scraping anything else. (Teaches: a real-world scraping courtesy, and one more HTTP call.)
HttpClient (no extra dependency needed), Kotlin’s check for a concise invariant check, and a pragmatic regex-based alternative to a full HTML parser for simple link-scraping. Related reference: Exception Handling in Kotlin, Kotlin & Java Interop.