sdxc

Type to search, or start from one of these:

[ Content & feeds ]

Read other people's pages

Save links for later by fetching each page under bounds in a job, pulling out its article, and building preview cards.

Last updated 2026-09-29

A "save for later" button looks like one line of work: fetch the URL, keep the article. The URL is one a stranger chose, though, so the fetch has to refuse your own network, stop at a byte cap and a deadline, and treat whatever comes back as hostile markup. This guide builds that feature: an endpoint accepts a link, a background job reads the article out of the page, and your app stores a sanitized copy and a plain-text excerpt.

@sdxc/distill fetches under bounds, scores the page to find the article and sanitizes it. @sdxc/html parses markup and answers questions about it: the visible text, a meta tag, a canonical link. @sdxc/robots lets a publisher say no, and @sdxc/jobs keeps the fetch off the request.

npm add @sdxc/distill @sdxc/html @sdxc/robots @sdxc/jobs @sdxc/result \
	@sdxc/validate @sdxc/http remix

Declare the job

Reading a page takes up to eight seconds, and nobody should wait on that when they press save. The endpoint stores the link and enqueues a job that reads it:

app/jobs/index.ts
import { job, jobs } from "@sdxc/jobs";
import * as s from "remix/data-schema";

export default jobs({
	links: {
		read: job({ input: s.object({ linkId: s.string() }) }),
	},
});

The message carries the link's id rather than its URL, so the job always reads the row as it is now, and a link deleted before the queue gets to it ends the run early.

addressable(url) is the check the fetch itself makes before any request: HTTP(S) only, and no literal IP address, loopback name or .local host. Running it in the endpoint turns a link the job would refuse into an immediate 422, instead of a row that fails later:

app/http/controllers/links/create.ts
import { addressable } from "@sdxc/distill";
import { accepted, unprocessableEntity } from "@sdxc/http/response/json";
import { isFailure } from "@sdxc/result";
import { validate } from "@sdxc/validate";
import * as s from "remix/data-schema";
import { createAction } from "remix/router";

import jobs from "~/app/jobs";
import { Links } from "~/app/repositories/links";
import routes from "~/routes/web";

const SAVE_LINK = s.object({ url: s.string() });

export default createAction(routes.links.create, async (ctx) => {
	let body = await validate(ctx.request, SAVE_LINK);
	if (isFailure(body)) return unprocessableEntity({ issues: body.error.issues });

	let url = addressable(body.data.url);
	if (isFailure(url)) return unprocessableEntity({ error: url.error.message });

	let link = await Links.save(ctx.db, { url: url.data.href });
	await ctx.jobs.enqueue(jobs.links.read, { linkId: link.id });
	return accepted({ id: link.id, status: link.status });
});

Links is your own repository; save inserts a row with a pending status. The route is a post("/api/links"), and validate reads a JSON or form body alike, so a browser extension and a plain form can share it.

ctx.jobs comes from jobEnqueuer(queue) in the router's middleware, over the queue your dispatcher delivers from. Background jobs and cron builds both:

bootstrap/app.ts
import { jobEnqueuer } from "@sdxc/jobs/router";
import { createRouter } from "remix/router";

import { queue } from "~/app/jobs/queue";

export const router = createRouter({ middleware: [jobEnqueuer(queue)] });

Read the article in the job

distill(url, options) walks the redirect chain itself, re-checking every hop with the same addressable rule, reads at most 2 MB off the stream, and gives the whole chain eight seconds. It scores the page's containers, keeps the one holding the article, and sanitizes it before answering:

app/jobs/links/read.ts
import { distill } from "@sdxc/distill";
import { createJobHandler } from "@sdxc/jobs";
import { isFailure } from "@sdxc/result";
import { fetchRobots } from "@sdxc/robots/fetch";

import jobs from "~/app/jobs";
import { excerptOf } from "~/app/lib/excerpt";
import { Links } from "~/app/repositories/links";

const USER_AGENT = "Shelf/1.0 (+https://shelf.example/about/bot)";

export default createJobHandler(jobs.links.read, async (ctx) => {
	let link = await Links.find(ctx.database, ctx.input.linkId);
	if (link === null) return ctx.exit("The link was deleted");

	let robots = await fetchRobots(link.url, { userAgent: USER_AGENT });
	let article = await distill(link.url, { userAgent: USER_AGENT, robots });

	if (isFailure(article)) {
		let { outcome } = article.error;
		ctx.log.set({ link: { outcome } });
		if (outcome === "timeout" && ctx.attempts < 3) {
			return ctx.retry({ delay: "15 minutes" });
		}
		await Links.markUnreadable(ctx.database, link.id, outcome);
		return;
	}

	let { byline, html, mayCache, title, url } = article.data;
	let excerpt = excerptOf(html);
	await Links.markRead(ctx.database, link.id, {
		url,
		title,
		byline,
		excerpt,
		html: mayCache ? html : null,
	});
	ctx.log.set({ link: { outcome: "extracted", bytes: article.data.bytes } });
});

The user agent is required, and it names your app and a page about it, so a publisher who wants to refuse you in particular can do it without refusing browsers. fetchRobots retrieves the origin's robots.txt, and a path it disallows is refused before distill sends a request. Its outcome carries a lifetimeMs, so once many links share an origin, cache it as JSON for that long instead of asking on every run.

Every failure carries an outcome your interface can have copy for: refused (the site said no, by status or by robots.txt), timeout (time, bytes or hops ran out) and empty (the page arrived with no article in it). Only a timeout is worth trying again, and only a few times; the other two would give the same answer. ctx.database is published by job middleware, the same way ctx.db is on a request.

On success, url is the article's canonical address, so two links to one article can be recognized as one. mayCache is false when the response's X-Robots-Tag says noarchive: that publisher is asking you not to keep a copy, so the row keeps the title and the excerpt and drops the body.

Keep a plain-text excerpt

A list of saved links needs a line or two under each title, and it must be text, not markup. HTML.parse reads the sanitized article back, and text is what a reader would see, with block boundaries turned into spaces:

app/lib/excerpt.ts
import { HTML } from "@sdxc/html";
import { isFailure } from "@sdxc/result";

const EXCERPT_LENGTH = 280;

export function excerptOf(html: string): string | null {
	let page = HTML.parse(html);
	if (isFailure(page)) return null;

	let text = page.data.text;
	if (text.length <= EXCERPT_LENGTH) return text;

	let cut = text.slice(0, EXCERPT_LENGTH - 1);
	let space = cut.lastIndexOf(" ");
	return `${space > 0 ? cut.slice(0, space) : cut}…`;
}

The stored html is already safe to put in your page. The sanitizer keeps prose, lists, tables, figures, links and images, drops scripts, iframes, forms, every on* handler, style and class, restricts URLs to http:, https: and mailto:, and resolves relative ones against the article's address. A Content-Security-Policy on the page that renders it is the second line behind that. An image still loads from the publisher, who sees the reader's address when it does; closing that takes an image proxy of your own.

Build a preview card from the head

Not every link needs its article. A URL pasted into a comment only needs a card: a title, a line of description, an image. @sdxc/distill/retrieve exports the same bounded fetch on its own, and @sdxc/html reads the head of what it returns:

app/services/link-preview.ts
import { addressable, readWithin, retrieve } from "@sdxc/distill/retrieve";
import { HTML } from "@sdxc/html";
import { isFailure, isSuccess } from "@sdxc/result";

const USER_AGENT = "Shelf/1.0 (+https://shelf.example/about/bot)";

export interface LinkPreview {
	url: string;
	title: string | null;
	description: string | null;
	image: string | null;
}

export async function previewOf(input: string): Promise<LinkPreview | null> {
	let url = addressable(input);
	if (isFailure(url)) return null;

	let retrieved = await retrieve(url.data, {
		userAgent: USER_AGENT,
		timeoutMs: 3_000,
	});
	if (isFailure(retrieved)) return null;

	let body = await readWithin(retrieved.data);
	let page = isSuccess(body) ? HTML.parse(body.data.text) : body;
	if (isFailure(page)) return null;

	let doc = page.data;
	let title = doc.meta("og:title");
	let description = doc.meta("og:description");
	let image = doc.meta("og:image");

	return {
		url: retrieved.data.url,
		title: isSuccess(title) ? title.data : (doc.title ?? null),
		description: isSuccess(description) ? description.data : null,
		image: isSuccess(image)
			? (URL.parse(image.data, retrieved.data.url)?.href ?? null)
			: null,
	};
}

retrieve answers only a 2xx response, reporting a refusing status such as 403 or 429 as a failure. retrieved.data.url is where the redirect chain ended, which is the address to show and the base every relative URL in the page resolves against. doc.meta matches name or property, so Open Graph tags and plain <meta name> tags are one lookup, and each answers a Result because a page is free to leave any of them out.

A card is optional, so every failure here becomes null and the comment renders without one. The shorter deadline is for the same reason: a card that takes eight seconds is not worth having. Call previewOf from a job, the same way the article is read.

Clean markup that arrives another way

Some markup reaches you without a fetch: the content of a feed item, or an HTML email. HTML.sanitize(source, policy) is the same allow-list distill applies, run on markup in hand:

app/lib/feed-content.ts
import { HTML } from "@sdxc/html";
import { isFailure } from "@sdxc/result";

export function cleanContent(source: string, itemUrl: string): string | null {
	let clean = HTML.sanitize(source, { baseUrl: itemUrl });
	return isFailure(clean) ? null : clean.data;
}

baseUrl is where the markup came from, so a relative src points at the publisher's site rather than yours. When you already hold a whole page, from a test fixture or a response something else retrieved, distillFrom(source, url) from @sdxc/distill runs the scoring and sanitizing without a fetch, which is how to test the job's extraction without the network.

Where to go next