Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,12 @@ All notable changes to kage are recorded here. The format follows

## [Unreleased]

### Fixed

- Saved pages no longer re-root relative URLs against the live origin via a leftover `<base href>`, and active content that escaped the script stripper is neutralized: `iframe` `srcdoc` and `data:text/html` sources, live remote frames, and HTML/SVG `object`/`embed` carriers.
- Relative links on pages that redirected are resolved against the browser's final URL (and any document `<base href>`), while the page is still written under the discovered URL so existing offline links keep working.
- Non-UTF-8 `<meta charset>` / Content-Type charset declarations are rewritten to `utf-8`, since kage always serialises pages as UTF-8 ([#16](https://github.com/tamnd/kage/issues/16)).

## [0.3.11] - 2026-08-01

### Fixed
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -259,7 +259,7 @@ cmd/kage/ thin main: pins the main thread, then hands off to cli.Execute
cli/ the cobra command tree and flag wiring
clone/ the crawl: frontier, render workers, asset workers, resume state
browser/ headless Chrome control and DOM snapshotting
sanitize/ strip scripts, handlers, and javascript: URLs from the DOM
sanitize/ strip scripts, handlers, active frames, and javascript: URLs
asset/ download and localise CSS, images, and fonts
urlx/ the deterministic URL-to-path mapping
zim/ a pure-Go ZIM reader and writer
Expand Down
59 changes: 58 additions & 1 deletion clone/cloner.go
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ import (
"github.com/tamnd/kage/sanitize"
"github.com/tamnd/kage/urlx"
"golang.org/x/net/html"
"golang.org/x/net/html/atom"
"golang.org/x/time/rate"
)

Expand Down Expand Up @@ -302,6 +303,13 @@ func (c *Cloner) processPage(ctx context.Context, j pageItem) {
return
}

// Resolve references against the post-redirect URL (and any <base href>),
// but keep writing the page under the discovered URL so existing offline
// links that pointed at /old still resolve. Cross-host redirects leave the
// resolve base as the final location for relative refs; scope checks still
// use that absolute URL.
resolveBase := pageResolveBase(j.u, res.FinalURL, root)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the good part of the PR and I want it. Resolving against the post redirect URL fixes a real bug: today a page fetched at /old that redirects to /new/ resolves href="next" as /next instead of /new/next, which breaks links and produces 404 asset fetches across the whole mirror.

I checked the ordering and it is correct. You read the <base href> here, RewriteHTML consumes it, and only then does CleanTree delete the element. Worth keeping that dependency in mind if anyone reorders processPage later.

One case to name in the CHANGELOG: on a cross host redirect (ex.com/old to other.com/new) the base becomes the other host, so every relative reference resolves out of scope, gets left absolute, and the page saved under ex.com/old ends up with nothing local in it at all. Your comment acknowledges the mechanism but not that outcome. Arguably the right answer is to enqueue the final URL as its own page when it is in scope, but I am happy to leave that for later as long as it is written down.


localFile := urlx.LocalPath(c.seedHost, j.u, urlx.Page, c.cfg.Reserved)
fileDir := urlx.Dir(localFile)

Expand All @@ -324,7 +332,7 @@ func (c *Cloner) processPage(ctx context.Context, j pageItem) {
}
}

asset.RewriteHTML(root, j.u, sink)
asset.RewriteHTML(root, resolveBase, sink)
sanitize.CleanTree(root, sanitize.Options{
KeepNoscript: c.cfg.KeepNoscript,
MobileReadable: c.cfg.MobileReadable,
Expand Down Expand Up @@ -354,6 +362,55 @@ func (c *Cloner) waitForCrawlDelay(ctx context.Context) bool {
return c.crawlLimiter.Wait(ctx) == nil
}

// pageResolveBase picks the URL against which relative references on a rendered
// page should resolve. Preference order:
// 1. A document <base href> (the live page's own base);
// 2. The browser's final URL after redirects;
// 3. The URL that was enqueued.
//
// The page is still written under the enqueued URL so offline links discovered
// as /old keep working when the server redirected /old → /new.
func pageResolveBase(enqueued *url.URL, finalURL string, root *html.Node) *url.URL {
base := enqueued
if finalURL != "" {
if u, err := url.Parse(finalURL); err == nil && u.Scheme != "" && u.Host != "" {
// Drop fragment; keep query/path as the browser shows them.
u.Fragment = ""
base = u
}
}
if href := documentBaseHref(root); href != "" {
if u, err := urlx.Normalize(base, href); err == nil {
return u
}
}
return base
}

// documentBaseHref returns the first <base href> in document order, or "".
func documentBaseHref(root *html.Node) string {

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This reimplements a first match tree walk that findElement in the sanitize package already does. Not blocking, but two nearly identical walkers in two packages is the kind of thing that drifts. A shared helper would be fine.

var found string
var walk func(*html.Node)
walk = func(n *html.Node) {
if found != "" || n == nil {
return
}
if n.Type == html.ElementNode && n.DataAtom == atom.Base {
for _, a := range n.Attr {
if strings.EqualFold(a.Key, "href") && strings.TrimSpace(a.Val) != "" {
found = strings.TrimSpace(a.Val)
return
}
}
}
for c := n.FirstChild; c != nil && found == ""; c = c.NextSibling {
walk(c)
}
}
walk(root)
return found
}

// processAsset downloads one asset, rewriting CSS references on the way, and
// writes it to its deterministic local path.
func (c *Cloner) processAsset(ctx context.Context, j assetItem) {
Expand Down
45 changes: 45 additions & 0 deletions clone/resolve_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
package clone

import (
"net/url"
"strings"
"testing"

"golang.org/x/net/html"
)

func TestPageResolveBaseUsesFinalURL(t *testing.T) {
enqueued, _ := url.Parse("https://ex.com/old")
root, err := html.Parse(strings.NewReader(`<html><head></head><body><a href="next">n</a></body></html>`))
if err != nil {
t.Fatal(err)
}
base := pageResolveBase(enqueued, "https://ex.com/new/", root)
if base.String() != "https://ex.com/new/" {
t.Fatalf("resolve base = %q, want final URL", base)
}
}

func TestPageResolveBasePrefersDocumentBase(t *testing.T) {
enqueued, _ := url.Parse("https://ex.com/page")
root, err := html.Parse(strings.NewReader(
`<html><head><base href="https://ex.com/dir/"></head><body></body></html>`))
if err != nil {
t.Fatal(err)
}
base := pageResolveBase(enqueued, "https://ex.com/page", root)
if base.String() != "https://ex.com/dir/" {
t.Fatalf("resolve base = %q, want document <base>", base)
}
}

func TestDocumentBaseHref(t *testing.T) {
root, err := html.Parse(strings.NewReader(
`<html><head><base href="/subdir/"><base href="https://ignored.example/"></head></html>`))
if err != nil {
t.Fatal(err)
}
if got := documentBaseHref(root); got != "/subdir/" {
t.Fatalf("documentBaseHref = %q, want first base", got)
}
}
5 changes: 5 additions & 0 deletions docs/content/reference/release-notes.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,11 @@ weight: 40

The authoritative, commit-level history lives in [`CHANGELOG.md`](https://github.com/tamnd/kage/blob/main/CHANGELOG.md) and on the [releases page](https://github.com/tamnd/kage/releases). This page summarises each version.

## Unreleased

- **Inert-snapshot gaps closed.** `<base href>`, `iframe` `srcdoc` / `data:text/html`, live remote frames, and HTML/SVG `object`/`embed` no longer reintroduce live or executable content.
- **Redirects resolve correctly.** Relative links use the post-redirect URL and any document `<base href>`, and non-UTF-8 charset metas are rewritten to UTF-8.

## v0.3.11

- **`go install ...@latest` works again.** The v0.3.9 antivirus fix replaced Rod's leakless dependency with a local stub. That kept the flagged helper out of `kage.exe`, but Go refuses versioned installation of a module containing a dependency-changing `replace` directive ([#72](https://github.com/tamnd/kage/issues/72)). Windows now launches Chrome through a small platform-specific launcher that never imports leakless. Other platforms keep Rod's launcher, the Windows binary remains free of the flagged helper, and the module no longer needs `replace`.
Expand Down
Loading