refactor(baoyu-url-to-markdown): replace custom pipeline with baoyu-fetch CLI

This commit is contained in:
Jim Liu 宝玉
2026-03-27 14:11:05 -05:00
parent d0764c2739
commit 2ff139112f
70 changed files with 7913 additions and 4369 deletions
Executable
BIN
View File
Binary file not shown.
+158 -151
View File
@@ -1,7 +1,7 @@
--- ---
name: baoyu-url-to-markdown name: baoyu-url-to-markdown
description: Fetch any URL and convert to markdown using Chrome CDP. Saves the rendered HTML snapshot alongside the markdown, uses an upgraded Defuddle pipeline with better web-component handling and YouTube transcript extraction, and automatically falls back to the pre-Defuddle HTML-to-Markdown pipeline when needed. If local browser capture fails entirely, it can fall back to the hosted defuddle.md API. Supports two modes - auto-capture on page load, or wait for user signal (for pages requiring login). Use when user wants to save a webpage as markdown. description: Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.
version: 1.59.0 version: 1.60.0
metadata: metadata:
openclaw: openclaw:
homepage: https://github.com/JimLiu/baoyu-skills#baoyu-url-to-markdown homepage: https://github.com/JimLiu/baoyu-skills#baoyu-url-to-markdown
@@ -13,29 +13,18 @@ metadata:
# URL to Markdown # URL to Markdown
Fetches any URL via Chrome CDP, saves the rendered HTML snapshot, and converts it to clean markdown. Fetches any URL via `baoyu-fetch` CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.
## Script Directory ## CLI Setup
**Important**: All scripts are located in the `scripts/` subdirectory of this skill. **Important**: The CLI source is vendored in the `scripts/vendor/baoyu-fetch/` subdirectory of this skill.
**Agent Execution Instructions**: **Agent Execution Instructions**:
1. Determine this SKILL.md file's directory path as `{baseDir}` 1. Determine this SKILL.md file's directory path as `{baseDir}`
2. Script path = `{baseDir}/scripts/<script-name>.ts` 2. CLI entry point = `{baseDir}/scripts/vendor/baoyu-fetch/src/cli.ts`
3. Resolve `${BUN_X}` runtime: if `bun` installed → `bun`; if `npx` available → `npx -y bun`; else suggest installing bun 3. Resolve `${BUN_X}` runtime: if `bun` installed → `bun`; if `npx` available → `npx -y bun`; else suggest installing bun
4. Replace all `{baseDir}` and `${BUN_X}` in this document with actual values 4. `${READER}` = `${BUN_X} {baseDir}/scripts/vendor/baoyu-fetch/src/cli.ts`
5. Replace all `${READER}` in this document with the resolved value
**Script Reference**:
| Script | Purpose |
|--------|---------|
| `scripts/main.ts` | CLI entry point for URL fetching |
| `scripts/html-to-markdown.ts` | Markdown conversion entry point and converter selection |
| `scripts/parsers/index.ts` | Unified parser entry: dispatches URL-specific rules before generic converters |
| `scripts/parsers/types.ts` | Unified parser interface shared by all rule files |
| `scripts/parsers/rules/*.ts` | One file per URL rule, for example X status and X article |
| `scripts/defuddle-converter.ts` | Defuddle-based conversion |
| `scripts/legacy-converter.ts` | Pre-Defuddle legacy extraction and markdown conversion |
| `scripts/markdown-conversion-shared.ts` | Shared metadata parsing and markdown document helpers |
## Preferences (EXTEND.md) ## Preferences (EXTEND.md)
@@ -56,23 +45,17 @@ if (Test-Path "$xdg/baoyu-skills/baoyu-url-to-markdown/EXTEND.md") { "xdg" }
if (Test-Path "$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md") { "user" } if (Test-Path "$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md") { "user" }
``` ```
┌────────────────────────────────────────────────────────┬───────────────────┐ | Path | Location |
│ Path │ Location │ |------|----------|
├────────────────────────────────────────────────────────┼───────────────────┤ | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Project directory |
.baoyu-skills/baoyu-url-to-markdown/EXTEND.md │ Project directory │ | `$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | User home |
├────────────────────────────────────────────────────────┼───────────────────┤
│ $HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md │ User home │
└────────────────────────────────────────────────────────┴───────────────────┘
┌───────────┬───────────────────────────────────────────────────────────────────────────┐ | Result | Action |
│ Result │ Action │ |--------|--------|
├───────────┼───────────────────────────────────────────────────────────────────────────┤ | Found | Read, parse, apply settings |
│ Found │ Read, parse, apply settings │ | Not found | **MUST** run first-time setup (see below) — do NOT silently create defaults |
├───────────┼───────────────────────────────────────────────────────────────────────────┤
│ Not found │ **MUST** run first-time setup (see below) — do NOT silently create defaults │
└───────────┴───────────────────────────────────────────────────────────────────────────┘
**EXTEND.md Supports**: Download media by default | Default output directory | Default capture mode | Timeout settings **EXTEND.md Supports**: Download media by default | Default output directory
### First-Time Setup (BLOCKING) ### First-Time Setup (BLOCKING)
@@ -107,54 +90,63 @@ Full reference: [references/config/first-time-setup.md](references/config/first-
**EXTEND.md → CLI mapping**: **EXTEND.md → CLI mapping**:
| EXTEND.md key | CLI argument | Notes | | EXTEND.md key | CLI argument | Notes |
|---------------|-------------|-------| |---------------|-------------|-------|
| `download_media: 1` | `--download-media` | | | `download_media: 1` | `--download-media` | Requires `--output` to be set |
| `default_output_dir: ./posts/` | `--output-dir ./posts/` | Directory path. Do NOT pass to `-o` (which expects a file path) | | `default_output_dir: ./posts/` | Agent constructs `--output ./posts/{domain}/{slug}.md` | Agent generates path, not a direct CLI flag |
**Value priority**: **Value priority**:
1. CLI arguments (`--download-media`, `-o`, `--output-dir`) 1. CLI arguments (`--download-media`, `--output`)
2. EXTEND.md 2. EXTEND.md
3. Skill defaults 3. Skill defaults
## Features ## Features
- Chrome CDP for full JavaScript rendering - Chrome CDP for full JavaScript rendering via `baoyu-fetch` CLI
- Browser strategy fallback: default headless first, then visible Chrome on technical failure - Site-specific adapters: X/Twitter, YouTube, Hacker News, generic (Defuddle)
- URL-specific parser layer for sites that need custom HTML rules before generic extraction - Automatic adapter selection based on URL, or force with `--adapter`
- Two capture modes: auto or wait-for-user - Interaction gate detection: Cloudflare, reCAPTCHA, hCAPTCHA, custom challenges
- Save rendered HTML as a sibling `-captured.html` file - Two capture modes: headless (default) or interactive with wait-for-interaction
- Clean markdown output with metadata - Clean markdown output with YAML front matter
- Upgraded Defuddle-first markdown conversion with automatic fallback to the pre-Defuddle extractor from git history - Structured JSON output available via `--format json`
- X/Twitter pages can use HTML-specific parsing for Tweets and Articles, which improves title/body/media extraction on `x.com` / `twitter.com` - X/Twitter: extracts tweets, threads, and X Articles with media
- `archive.ph` / related archive mirrors can restore the original URL from `input[name=q]` and prefer `#CONTENT` before falling back to the page body - YouTube: transcript/caption extraction, chapters, cover images
- Materializes shadow DOM content before conversion so web-component pages survive serialization better - Hacker News: threaded comment parsing with proper nesting
- YouTube pages can include transcript/caption text in the markdown when YouTube exposes a caption track - Generic: Defuddle extraction with Readability fallback
- If local browser capture fails completely, can fall back to `defuddle.md/<url>` and still save markdown
- Handles login-required pages via wait mode
- Download images and videos to local directories - Download images and videos to local directories
- Chrome profile persistence for authenticated sessions
- Debug artifact output for troubleshooting
## Usage ## Usage
```bash ```bash
# Auto mode (default) - capture when page loads # Default: headless capture, markdown to stdout
${BUN_X} {baseDir}/scripts/main.ts <url> ${READER} <url>
# Force headless only # Save to file
${BUN_X} {baseDir}/scripts/main.ts <url> --browser headless ${READER} <url> --output article.md
# Force visible browser # Save with media download
${BUN_X} {baseDir}/scripts/main.ts <url> --browser headed ${READER} <url> --output article.md --download-media
# Wait mode - wait for user signal before capture # Headless mode (explicit)
${BUN_X} {baseDir}/scripts/main.ts <url> --wait ${READER} <url> --headless --output article.md
# Save to specific file # Wait for interaction (login/CAPTCHA) — auto-detect and continue
${BUN_X} {baseDir}/scripts/main.ts <url> -o output.md ${READER} <url> --wait-for interaction --output article.md
# Save to a custom output directory (auto-generates filename) # Wait for interaction — manual control (Enter to continue)
${BUN_X} {baseDir}/scripts/main.ts <url> --output-dir ./posts/ ${READER} <url> --wait-for force --output article.md
# Download images and videos to local directories # JSON output
${BUN_X} {baseDir}/scripts/main.ts <url> --download-media ${READER} <url> --format json --output article.json
# Force specific adapter
${READER} <url> --adapter youtube --output transcript.md
# Connect to existing Chrome
${READER} <url> --cdp-url http://localhost:9222 --output article.md
# Debug artifacts
${READER} <url> --output article.md --debug-dir ./debug/
``` ```
## Options ## Options
@@ -162,112 +154,122 @@ ${BUN_X} {baseDir}/scripts/main.ts <url> --download-media
| Option | Description | | Option | Description |
|--------|-------------| |--------|-------------|
| `<url>` | URL to fetch | | `<url>` | URL to fetch |
| `-o <path>` | Output file path — must be a **file** path, not directory (default: auto-generated) | | `--output <path>` | Output file path (default: stdout) |
| `--output-dir <dir>` | Base output directory — auto-generates `{dir}/{domain}/{slug}.md` (default: `./url-to-markdown/`) | | `--format <type>` | Output format: `markdown` (default) or `json` |
| `--wait` | Wait for user signal before capturing | | `--json` | Shorthand for `--format json` |
| `--browser <mode>` | Browser strategy: `auto` (default), `headless`, or `headed` | | `--adapter <name>` | Force adapter: `x`, `youtube`, `hn`, or `generic` (default: auto-detect) |
| `--headless` | Shortcut for `--browser headless` | | `--headless` | Force headless Chrome (no visible window) |
| `--headed` | Shortcut for `--browser headed` | | `--wait-for <mode>` | Interaction wait mode: `none` (default), `interaction`, or `force` |
| `--wait-for-interaction` | Alias for `--wait-for interaction` |
| `--wait-for-login` | Alias for `--wait-for interaction` |
| `--timeout <ms>` | Page load timeout (default: 30000) | | `--timeout <ms>` | Page load timeout (default: 30000) |
| `--download-media` | Download image/video assets to local `imgs/` and `videos/`, and rewrite markdown links to local relative paths | | `--interaction-timeout <ms>` | Login/CAPTCHA wait timeout (default: 600000 = 10 min) |
| `--interaction-poll-interval <ms>` | Poll interval for interaction checks (default: 1500) |
| `--download-media` | Download images/videos to local `imgs/` and `videos/`, rewrite markdown links. Requires `--output` |
| `--media-dir <dir>` | Base directory for downloaded media (default: same as `--output` directory) |
| `--cdp-url <url>` | Reuse existing Chrome DevTools Protocol endpoint |
| `--browser-path <path>` | Custom Chrome/Chromium binary path |
| `--chrome-profile-dir <path>` | Chrome user data directory (default: `BAOYU_CHROME_PROFILE_DIR` env or `./baoyu-skills/chrome-profile`) |
| `--debug-dir <dir>` | Write debug artifacts (document.json, markdown.md, page.html, network.json) |
## Capture Modes ## Capture Modes
| Mode | Behavior | Use When | | Mode | Behavior | Use When |
|------|----------|----------| |------|----------|----------|
| Auto (default) | Try headless first, then retry in visible Chrome if needed | Public pages, static content, unknown pages | | Default | Headless Chrome, auto-extract on network idle | Public pages, static content |
| Wait (`--wait`) | User signals when ready | Login-required, lazy loading, paywalls | | `--headless` | Explicit headless (same as default) | Clarify intent |
| `--wait-for interaction` | Opens visible Chrome, auto-detects login/CAPTCHA gates, waits for them to clear, then continues | Login-required, CAPTCHA-protected |
| `--wait-for force` | Opens visible Chrome, auto-detects OR accepts Enter keypress to continue | Complex flows, lazy loading, paywalls |
**Wait mode workflow**: **Interaction gate auto-detection**:
1. Run with `--wait` → script outputs "Press Enter when ready" - Cloudflare Turnstile / "just a moment" pages
2. Ask user to confirm page is ready - Google reCAPTCHA
3. Send newline to stdin to trigger capture - hCaptcha
- Custom challenge / verification screens
**Default browser fallback**: **Wait-for-interaction workflow**:
1. Auto mode starts with headless Chrome and captures on network idle 1. Run with `--wait-for interaction` → Chrome opens visibly
2. If headless capture fails technically, retry with visible Chrome 2. CLI auto-detects login/CAPTCHA gates
3. If a shared Chrome session for this profile already exists, reuse it instead of launching a new browser 3. User completes login or solves CAPTCHA in the browser
4. The script does not hard-code login or paywall detection; the agent must inspect the captured markdown or HTML and decide whether to rerun with `--browser headed --wait` 4. CLI auto-detects gate cleared → captures page
5. If `--wait-for force` is used, user can also press Enter to trigger capture manually
## Agent Quality Gate ## Agent Quality Gate
**CRITICAL**: The agent must treat headless capture as provisional. Some sites render differently in headless mode and can silently return an error shell, partially hydrated page, or low-quality extraction **without** causing the CLI to fail. **CRITICAL**: The agent must treat default headless capture as provisional. Some sites render differently in headless mode and can silently return low-quality content without causing the CLI to fail.
After every run that used `--browser auto` or `--browser headless`, the agent **MUST** inspect the saved markdown first, and inspect the saved `-captured.html` when the markdown looks suspicious. After every headless run, the agent **MUST** inspect the saved markdown output.
### Quality checks the agent must perform ### Quality checks the agent must perform
1. Confirm the markdown title matches the target page, not a generic site shell 1. Confirm the markdown title matches the target page, not a generic site shell
2. Confirm the body contains the expected article or page content, not just navigation, footer, or a generic error 2. Confirm the body contains the expected article or page content, not just navigation, footer, or a generic error
3. Watch for obvious failure signs such as: 3. Watch for obvious failure signs:
- `Application error` - `Application error`
- `This page could not be found` - `This page could not be found`
- login, signup, subscribe, or verification shells - Login, signup, subscribe, or verification shells
- extremely short markdown for a page that should be long-form - Extremely short markdown for a page that should be long-form
- raw framework payloads or mostly boilerplate content - Raw framework payloads or mostly boilerplate content
4. If the result is low quality, incomplete, or clearly wrong, do **not** accept the run as successful just because the CLI exited with code 0 4. If the result is low quality, incomplete, or clearly wrong, do **not** accept the run as successful just because the CLI exited with code 0
**Tip**: Use `--format json` to get structured output including `status`, `login.state`, and `interaction` fields for programmatic quality assessment. A `"status": "needs_interaction"` response means the page requires manual interaction.
### Recovery workflow the agent must follow ### Recovery workflow the agent must follow
1. First run with default `auto` unless there is already a clear reason to use wait mode 1. First run headless (default) unless there is already a clear reason to use interaction mode
2. Review markdown quality immediately after the run 2. Review markdown quality immediately after the run
3. If the content is low quality, rerun locally with visible Chrome: 3. If the content is low quality or indicates login/CAPTCHA:
- `--browser headed` for ordinary rendering issues - `--wait-for interaction` for auto-detected gates (login, CAPTCHA, Cloudflare)
- `--browser headed --wait` when the page may need login, anti-bot interaction, cookie acceptance, or extra hydration time - `--wait-for force` when the page needs manual browsing, scroll loading, or complex interaction
4. If `--wait` is used, tell the user exactly what to do: 4. If `--wait-for` is used, tell the user exactly what to do:
- if login is required, ask them to sign in - If login is required, ask them to sign in in the browser
- if the page needs time to hydrate, ask them to wait until the full content is visible - If CAPTCHA appears, ask them to solve it
- once ready, ask them to press Enter so capture can continue - If the page needs time to load, ask them to wait until content is visible
5. Only fall back to hosted `defuddle.md` after the local browser strategies have failed or are clearly lower fidelity - For `--wait-for force`: tell them to press Enter when ready
5. If JSON output shows `"status": "needs_interaction"`, switch to `--wait-for interaction` automatically
## Output Path Generation
The agent must construct the output file path since `baoyu-fetch` does not auto-generate paths.
**Algorithm**:
1. Determine base directory from EXTEND.md `default_output_dir` or default `./url-to-markdown/`
2. Extract domain from URL (e.g., `example.com`)
3. Generate slug from URL path or page title (kebab-case, 2-6 words)
4. Construct: `{base_dir}/{domain}/{slug}/{slug}.md` — each URL gets its own directory so media files stay isolated
5. Conflict resolution: append timestamp `{slug}-YYYYMMDD-HHMMSS/{slug}-YYYYMMDD-HHMMSS.md`
Pass the constructed path to `--output`. Media files (`--download-media`) are saved into subdirectories next to the markdown file, keeping each URL's assets self-contained.
## Output Format ## Output Format
Each run saves two files side by side: Markdown output to stdout (or file with `--output`) as clean markdown text.
- Markdown: YAML front matter with `url`, `title`, `description`, `author`, `published`, optional `coverImage`, and `captured_at`, followed by converted markdown content JSON output (`--format json`) returns structured data including:
- HTML snapshot: `*-captured.html`, containing the rendered page HTML captured from Chrome - `adapter` — which adapter handled the URL
- `status``"ok"` or `"needs_interaction"`
When Defuddle or page metadata provides a language hint, the markdown front matter also includes `language`. - `login` — login state detection (`logged_in`, `logged_out`, `unknown`)
- `interaction` — interaction gate details (kind, provider, prompt)
The HTML snapshot is saved before any markdown media localization, so it stays a faithful capture of the page DOM used for conversion. - `document` — structured content (url, title, author, publishedAt, content blocks, metadata)
If the hosted `defuddle.md` API fallback is used, markdown is still saved, but there is no local `-captured.html` snapshot for that run. - `media` — collected media assets with url, kind, role
- `markdown` — converted markdown text
## Output Directory - `downloads` — media download results (when `--download-media` used)
Default: `url-to-markdown/<domain>/<slug>.md`
With `--output-dir ./posts/`: `./posts/<domain>/<slug>.md`
HTML snapshot path uses the same basename:
- `url-to-markdown/<domain>/<slug>-captured.html`
- `./posts/<domain>/<slug>-captured.html`
- `<slug>`: From page title or URL path (kebab-case, 2-6 words)
- Conflict resolution: Append timestamp `<slug>-YYYYMMDD-HHMMSS.md`
When `--download-media` is enabled: When `--download-media` is enabled:
- Images are saved to `imgs/` next to the markdown file - Images are saved to `imgs/` next to the output file (or in `--media-dir`)
- Videos are saved to `videos/` next to the markdown file - Videos are saved to `videos/` next to the output file (or in `--media-dir`)
- Markdown media links are rewritten to local relative paths - Markdown media links are rewritten to local relative paths
## Conversion Fallback ## Built-in Adapters
Conversion order: | Adapter | URLs | Key Features |
|---------|------|-------------|
| `x` | x.com, twitter.com | Tweets, threads, X Articles, media, login detection |
| `youtube` | youtube.com, youtu.be | Transcript/captions, chapters, cover image, metadata |
| `hn` | news.ycombinator.com | Threaded comments, story metadata, nested replies |
| `generic` | Any URL (fallback) | Defuddle extraction, Readability fallback, auto-scroll, network idle detection |
1. Try the URL-specific parser layer first when a site rule matches Adapter is auto-selected based on URL. Use `--adapter <name>` to override.
2. If no specialized parser matches, try Defuddle
3. For rich pages such as YouTube, prefer Defuddle's extractor-specific output (including transcripts when available) instead of replacing it with the legacy pipeline
4. If Defuddle throws, cannot load, returns obviously incomplete markdown, or captures lower-quality content than the legacy pipeline, automatically fall back to the pre-Defuddle extractor
5. If the agent determines the captured result is a login screen, verification screen, or paywall shell, rerun locally with `--browser headed --wait` and ask the user to complete access before capture
6. If the entire local browser capture flow still fails before markdown can be produced, try the hosted `https://defuddle.md/<url>` API and save its markdown output directly
7. The legacy fallback path uses the older Readability/selector/Next.js-data based HTML-to-Markdown implementation recovered from git history
CLI output will show:
- `Converter: parser:...` when a URL-specific parser succeeded
- `Converter: defuddle` when Defuddle succeeds
- `Converter: legacy:...` plus `Fallback used: ...` when fallback was needed
- `Converter: defuddle-api` when local browser capture failed and the hosted API was used instead
## Media Download Workflow ## Media Download Workflow
@@ -275,42 +277,47 @@ Based on `download_media` setting in EXTEND.md:
| Setting | Behavior | | Setting | Behavior |
|---------|----------| |---------|----------|
| `1` (always) | Run script with `--download-media` flag | | `1` (always) | Run CLI with `--download-media --output <path>` |
| `0` (never) | Run script without `--download-media` flag | | `0` (never) | Run CLI with `--output <path>` (no media download) |
| `ask` (default) | Follow the ask-each-time flow below | | `ask` (default) | Follow the ask-each-time flow below |
### Ask-Each-Time Flow ### Ask-Each-Time Flow
1. Run script **without** `--download-media` → markdown saved 1. Run CLI **without** `--download-media` with `--output <path>` → markdown saved
2. Check saved markdown for remote media URLs (`https://` in image/video links) 2. Check saved markdown for remote media URLs (`https://` in image/video links)
3. **If no remote media found** → done, no prompt needed 3. **If no remote media found** → done, no prompt needed
4. **If remote media found** → use `AskUserQuestion`: 4. **If remote media found** → use `AskUserQuestion`:
- header: "Media", question: "Download N images/videos to local files?" - header: "Media", question: "Download N images/videos to local files?"
- "Yes" — Download to local directories - "Yes" — Download to local directories
- "No" — Keep remote URLs - "No" — Keep remote URLs
5. If user confirms → run script **again** with `--download-media` (overwrites markdown with localized links) 5. If user confirms → run CLI **again** with `--download-media --output <same-path>` (overwrites markdown with localized links)
## Environment Variables ## Environment Variables
| Variable | Description | | Variable | Description |
|----------|-------------| |----------|-------------|
| `URL_CHROME_PATH` | Custom Chrome executable path | | `BAOYU_CHROME_PROFILE_DIR` | Chrome user data directory (can also use `--chrome-profile-dir`) |
| `URL_DATA_DIR` | Custom data directory |
| `URL_CHROME_PROFILE_DIR` | Custom Chrome profile directory |
**Troubleshooting**: Chrome not found → set `URL_CHROME_PATH`. Timeout → increase `--timeout`. Complex pages → try `--wait` mode. If markdown quality is poor, inspect the saved `-captured.html` and check whether the run logged a legacy fallback. **Troubleshooting**: Chrome not found → use `--browser-path`. Timeout → increase `--timeout`. Login/CAPTCHA pages → use `--wait-for interaction`. Debug → use `--debug-dir` to inspect captured HTML and network logs.
### YouTube Notes ### YouTube Notes
- The upgraded Defuddle path uses async extractors, so YouTube pages can include transcript text directly in the markdown body. - YouTube adapter extracts transcripts/captions automatically when available
- Transcript availability depends on YouTube exposing a caption track. Videos with captions disabled, restricted playback, or blocked regional access may still produce description-only output. - Transcript format: `[MM:SS] Text segment` with chapter headings
- If the page needs time to finish loading descriptions, chapters, or player metadata, prefer `--wait` and capture after the watch page is fully hydrated. - Transcript availability depends on YouTube exposing a caption track. Videos with captions disabled or restricted playback may produce description-only output
- Use `--wait-for force` if the page needs time to finish loading player metadata
### Hosted API Fallback ### X/Twitter Notes
- The hosted fallback endpoint is `https://defuddle.md/<url>`. In shell form: `curl https://defuddle.md/stephango.com` - Extracts single tweets, threads, and X Articles
- Use it only when the local Chrome/CDP capture path fails outright. The local path still has higher fidelity because it can save the captured HTML and handle authenticated pages. - Auto-detects login state; if logged out and content requires auth, JSON output will show `"status": "needs_interaction"`
- The hosted API already returns Markdown with YAML frontmatter, so save that response as-is and then apply the normal media-localization step if requested. - Use `--wait-for interaction` for login-protected content
### Hacker News Notes
- Parses threaded comments with proper nesting and reply hierarchy
- Includes story metadata (title, URL, author, score, comment count)
- Shows comment deletion/dead status
## Extension Support ## Extension Support
+346 -40
View File
@@ -4,19 +4,49 @@
"": { "": {
"name": "baoyu-url-to-markdown-scripts", "name": "baoyu-url-to-markdown-scripts",
"dependencies": { "dependencies": {
"@mozilla/readability": "^0.6.0", "baoyu-fetch": "file:./vendor/baoyu-fetch",
"baoyu-chrome-cdp": "file:./vendor/baoyu-chrome-cdp",
"defuddle": "^0.14.0",
"jsdom": "^24.1.3",
"linkedom": "^0.18.12",
"turndown": "^7.2.2",
"turndown-plugin-gfm": "^1.0.2",
}, },
}, },
}, },
"packages": { "packages": {
"@asamuzakjp/css-color": ["@asamuzakjp/css-color@3.2.0", "", { "dependencies": { "@csstools/css-calc": "^2.1.3", "@csstools/css-color-parser": "^3.0.9", "@csstools/css-parser-algorithms": "^3.0.4", "@csstools/css-tokenizer": "^3.0.3", "lru-cache": "^10.4.3" } }, "sha512-K1A6z8tS3XsmCMM86xoWdn7Fkdn9m6RSVtocUrJYIwZnFVkng/PvkEoWtOWmP+Scc6saYWHWZYbndEEXxl24jw=="], "@asamuzakjp/css-color": ["@asamuzakjp/css-color@3.2.0", "", { "dependencies": { "@csstools/css-calc": "^2.1.3", "@csstools/css-color-parser": "^3.0.9", "@csstools/css-parser-algorithms": "^3.0.4", "@csstools/css-tokenizer": "^3.0.3", "lru-cache": "^10.4.3" } }, "sha512-K1A6z8tS3XsmCMM86xoWdn7Fkdn9m6RSVtocUrJYIwZnFVkng/PvkEoWtOWmP+Scc6saYWHWZYbndEEXxl24jw=="],
"@babel/runtime": ["@babel/runtime@7.29.2", "", {}, "sha512-JiDShH45zKHWyGe4ZNVRrCjBz8Nh9TMmZG1kh4QTK8hCBTWBi8Da+i7s1fJw7/lYpM4ccepSNfqzZ/QvABBi5g=="],
"@changesets/apply-release-plan": ["@changesets/apply-release-plan@7.1.0", "", { "dependencies": { "@changesets/config": "^3.1.3", "@changesets/get-version-range-type": "^0.4.0", "@changesets/git": "^3.0.4", "@changesets/should-skip-package": "^0.1.2", "@changesets/types": "^6.1.0", "@manypkg/get-packages": "^1.1.3", "detect-indent": "^6.0.0", "fs-extra": "^7.0.1", "lodash.startcase": "^4.4.0", "outdent": "^0.5.0", "prettier": "^2.7.1", "resolve-from": "^5.0.0", "semver": "^7.5.3" } }, "sha512-yq8ML3YS7koKQ/9bk1PqO0HMzApIFNwjlwCnwFEXMzNe8NpzeeYYKCmnhWJGkN8g7E51MnWaSbqRcTcdIxUgnQ=="],
"@changesets/assemble-release-plan": ["@changesets/assemble-release-plan@6.0.9", "", { "dependencies": { "@changesets/errors": "^0.2.0", "@changesets/get-dependents-graph": "^2.1.3", "@changesets/should-skip-package": "^0.1.2", "@changesets/types": "^6.1.0", "@manypkg/get-packages": "^1.1.3", "semver": "^7.5.3" } }, "sha512-tPgeeqCHIwNo8sypKlS3gOPmsS3wP0zHt67JDuL20P4QcXiw/O4Hl7oXiuLnP9yg+rXLQ2sScdV1Kkzde61iSQ=="],
"@changesets/changelog-git": ["@changesets/changelog-git@0.2.1", "", { "dependencies": { "@changesets/types": "^6.1.0" } }, "sha512-x/xEleCFLH28c3bQeQIyeZf8lFXyDFVn1SgcBiR2Tw/r4IAWlk1fzxCEZ6NxQAjF2Nwtczoen3OA2qR+UawQ8Q=="],
"@changesets/cli": ["@changesets/cli@2.30.0", "", { "dependencies": { "@changesets/apply-release-plan": "^7.1.0", "@changesets/assemble-release-plan": "^6.0.9", "@changesets/changelog-git": "^0.2.1", "@changesets/config": "^3.1.3", "@changesets/errors": "^0.2.0", "@changesets/get-dependents-graph": "^2.1.3", "@changesets/get-release-plan": "^4.0.15", "@changesets/git": "^3.0.4", "@changesets/logger": "^0.1.1", "@changesets/pre": "^2.0.2", "@changesets/read": "^0.6.7", "@changesets/should-skip-package": "^0.1.2", "@changesets/types": "^6.1.0", "@changesets/write": "^0.4.0", "@inquirer/external-editor": "^1.0.2", "@manypkg/get-packages": "^1.1.3", "ansi-colors": "^4.1.3", "enquirer": "^2.4.1", "fs-extra": "^7.0.1", "mri": "^1.2.0", "package-manager-detector": "^0.2.0", "picocolors": "^1.1.0", "resolve-from": "^5.0.0", "semver": "^7.5.3", "spawndamnit": "^3.0.1", "term-size": "^2.1.0" }, "bin": { "changeset": "bin.js" } }, "sha512-5D3Nk2JPqMI1wK25pEymeWRSlSMdo5QOGlyfrKg0AOufrUcjEE3RQgaCpHoBiM31CSNrtSgdJ0U6zL1rLDDfBA=="],
"@changesets/config": ["@changesets/config@3.1.3", "", { "dependencies": { "@changesets/errors": "^0.2.0", "@changesets/get-dependents-graph": "^2.1.3", "@changesets/logger": "^0.1.1", "@changesets/should-skip-package": "^0.1.2", "@changesets/types": "^6.1.0", "@manypkg/get-packages": "^1.1.3", "fs-extra": "^7.0.1", "micromatch": "^4.0.8" } }, "sha512-vnXjcey8YgBn2L1OPWd3ORs0bGC4LoYcK/ubpgvzNVr53JXV5GiTVj7fWdMRsoKUH7hhhMAQnsJUqLr21EncNw=="],
"@changesets/errors": ["@changesets/errors@0.2.0", "", { "dependencies": { "extendable-error": "^0.1.5" } }, "sha512-6BLOQUscTpZeGljvyQXlWOItQyU71kCdGz7Pi8H8zdw6BI0g3m43iL4xKUVPWtG+qrrL9DTjpdn8eYuCQSRpow=="],
"@changesets/get-dependents-graph": ["@changesets/get-dependents-graph@2.1.3", "", { "dependencies": { "@changesets/types": "^6.1.0", "@manypkg/get-packages": "^1.1.3", "picocolors": "^1.1.0", "semver": "^7.5.3" } }, "sha512-gphr+v0mv2I3Oxt19VdWRRUxq3sseyUpX9DaHpTUmLj92Y10AGy+XOtV+kbM6L/fDcpx7/ISDFK6T8A/P3lOdQ=="],
"@changesets/get-release-plan": ["@changesets/get-release-plan@4.0.15", "", { "dependencies": { "@changesets/assemble-release-plan": "^6.0.9", "@changesets/config": "^3.1.3", "@changesets/pre": "^2.0.2", "@changesets/read": "^0.6.7", "@changesets/types": "^6.1.0", "@manypkg/get-packages": "^1.1.3" } }, "sha512-Q04ZaRPuEVZtA+auOYgFaVQQSA98dXiVe/yFaZfY7hoSmQICHGvP0TF4u3EDNHWmmCS4ekA/XSpKlSM2PyTS2g=="],
"@changesets/get-version-range-type": ["@changesets/get-version-range-type@0.4.0", "", {}, "sha512-hwawtob9DryoGTpixy1D3ZXbGgJu1Rhr+ySH2PvTLHvkZuQ7sRT4oQwMh0hbqZH1weAooedEjRsbrWcGLCeyVQ=="],
"@changesets/git": ["@changesets/git@3.0.4", "", { "dependencies": { "@changesets/errors": "^0.2.0", "@manypkg/get-packages": "^1.1.3", "is-subdir": "^1.1.1", "micromatch": "^4.0.8", "spawndamnit": "^3.0.1" } }, "sha512-BXANzRFkX+XcC1q/d27NKvlJ1yf7PSAgi8JG6dt8EfbHFHi4neau7mufcSca5zRhwOL8j9s6EqsxmT+s+/E6Sw=="],
"@changesets/logger": ["@changesets/logger@0.1.1", "", { "dependencies": { "picocolors": "^1.1.0" } }, "sha512-OQtR36ZlnuTxKqoW4Sv6x5YIhOmClRd5pWsjZsddYxpWs517R0HkyiefQPIytCVh4ZcC5x9XaG8KTdd5iRQUfg=="],
"@changesets/parse": ["@changesets/parse@0.4.3", "", { "dependencies": { "@changesets/types": "^6.1.0", "js-yaml": "^4.1.1" } }, "sha512-ZDmNc53+dXdWEv7fqIUSgRQOLYoUom5Z40gmLgmATmYR9NbL6FJJHwakcCpzaeCy+1D0m0n7mT4jj2B/MQPl7A=="],
"@changesets/pre": ["@changesets/pre@2.0.2", "", { "dependencies": { "@changesets/errors": "^0.2.0", "@changesets/types": "^6.1.0", "@manypkg/get-packages": "^1.1.3", "fs-extra": "^7.0.1" } }, "sha512-HaL/gEyFVvkf9KFg6484wR9s0qjAXlZ8qWPDkTyKF6+zqjBe/I2mygg3MbpZ++hdi0ToqNUF8cjj7fBy0dg8Ug=="],
"@changesets/read": ["@changesets/read@0.6.7", "", { "dependencies": { "@changesets/git": "^3.0.4", "@changesets/logger": "^0.1.1", "@changesets/parse": "^0.4.3", "@changesets/types": "^6.1.0", "fs-extra": "^7.0.1", "p-filter": "^2.1.0", "picocolors": "^1.1.0" } }, "sha512-D1G4AUYGrBEk8vj8MGwf75k9GpN6XL3wg8i42P2jZZwFLXnlr2Pn7r9yuQNbaMCarP7ZQWNJbV6XLeysAIMhTA=="],
"@changesets/should-skip-package": ["@changesets/should-skip-package@0.1.2", "", { "dependencies": { "@changesets/types": "^6.1.0", "@manypkg/get-packages": "^1.1.3" } }, "sha512-qAK/WrqWLNCP22UDdBTMPH5f41elVDlsNyat180A33dWxuUDyNpg6fPi/FyTZwRriVjg0L8gnjJn2F9XAoF0qw=="],
"@changesets/types": ["@changesets/types@6.1.0", "", {}, "sha512-rKQcJ+o1nKNgeoYRHKOS07tAMNd3YSN0uHaJOZYjBAgxfV7TUE7JE+z4BzZdQwb5hKaYbayKN5KrYV7ODb2rAA=="],
"@changesets/write": ["@changesets/write@0.4.0", "", { "dependencies": { "@changesets/types": "^6.1.0", "fs-extra": "^7.0.1", "human-id": "^4.1.1", "prettier": "^2.7.1" } }, "sha512-CdTLvIOPiCNuH71pyDu3rA+Q0n65cmAbXnwWH84rKGiFumFzkmHNT8KHTMEchcxN+Kl8I54xGUhJ7l3E7X396Q=="],
"@csstools/color-helpers": ["@csstools/color-helpers@5.1.0", "", {}, "sha512-S11EXWJyy0Mz5SYvRmY8nJYTFFd1LCNV+7cXyAgQtOOuzb4EsgfqDufL+9esx72/eLhsRdGZwaldu/h+E4t4BA=="], "@csstools/color-helpers": ["@csstools/color-helpers@5.1.0", "", {}, "sha512-S11EXWJyy0Mz5SYvRmY8nJYTFFd1LCNV+7cXyAgQtOOuzb4EsgfqDufL+9esx72/eLhsRdGZwaldu/h+E4t4BA=="],
"@csstools/css-calc": ["@csstools/css-calc@2.1.4", "", { "peerDependencies": { "@csstools/css-parser-algorithms": "^3.0.5", "@csstools/css-tokenizer": "^3.0.4" } }, "sha512-3N8oaj+0juUw/1H3YwmDDJXCgTB1gKU6Hc/bB502u9zR0q2vd786XJH9QfrKIEgFlZmhZiq6epXl4rHqhzsIgQ=="], "@csstools/css-calc": ["@csstools/css-calc@2.1.4", "", { "peerDependencies": { "@csstools/css-parser-algorithms": "^3.0.5", "@csstools/css-tokenizer": "^3.0.4" } }, "sha512-3N8oaj+0juUw/1H3YwmDDJXCgTB1gKU6Hc/bB502u9zR0q2vd786XJH9QfrKIEgFlZmhZiq6epXl4rHqhzsIgQ=="],
@@ -27,26 +57,76 @@
"@csstools/css-tokenizer": ["@csstools/css-tokenizer@3.0.4", "", {}, "sha512-Vd/9EVDiu6PPJt9yAh6roZP6El1xHrdvIVGjyBsHR0RYwNHgL7FJPyIIW4fANJNG6FtyZfvlRPpFI4ZM/lubvw=="], "@csstools/css-tokenizer": ["@csstools/css-tokenizer@3.0.4", "", {}, "sha512-Vd/9EVDiu6PPJt9yAh6roZP6El1xHrdvIVGjyBsHR0RYwNHgL7FJPyIIW4fANJNG6FtyZfvlRPpFI4ZM/lubvw=="],
"@inquirer/external-editor": ["@inquirer/external-editor@1.0.3", "", { "dependencies": { "chardet": "^2.1.1", "iconv-lite": "^0.7.0" }, "peerDependencies": { "@types/node": ">=18" }, "optionalPeers": ["@types/node"] }, "sha512-RWbSrDiYmO4LbejWY7ttpxczuwQyZLBUyygsA9Nsv95hpzUWwnNTVQmAq3xuh7vNwCp07UTmE5i11XAEExx4RA=="],
"@manypkg/find-root": ["@manypkg/find-root@1.1.0", "", { "dependencies": { "@babel/runtime": "^7.5.5", "@types/node": "^12.7.1", "find-up": "^4.1.0", "fs-extra": "^8.1.0" } }, "sha512-mki5uBvhHzO8kYYix/WRy2WX8S3B5wdVSc9D6KcU5lQNglP2yt58/VfLuAK49glRXChosY8ap2oJ1qgma3GUVA=="],
"@manypkg/get-packages": ["@manypkg/get-packages@1.1.3", "", { "dependencies": { "@babel/runtime": "^7.5.5", "@changesets/types": "^4.0.1", "@manypkg/find-root": "^1.1.0", "fs-extra": "^8.1.0", "globby": "^11.0.0", "read-yaml-file": "^1.1.0" } }, "sha512-fo+QhuU3qE/2TQMQmbVMqaQ6EWbMhi4ABWP+O4AM1NqPBuy0OrApV5LO6BrrgnhtAHS2NH6RrVk9OL181tTi8A=="],
"@mixmark-io/domino": ["@mixmark-io/domino@2.2.0", "", {}, "sha512-Y28PR25bHXUg88kCV7nivXrP2Nj2RueZ3/l/jdx6J9f8J4nsEGcgX0Qe6lt7Pa+J79+kPiJU3LguR6O/6zrLOw=="], "@mixmark-io/domino": ["@mixmark-io/domino@2.2.0", "", {}, "sha512-Y28PR25bHXUg88kCV7nivXrP2Nj2RueZ3/l/jdx6J9f8J4nsEGcgX0Qe6lt7Pa+J79+kPiJU3LguR6O/6zrLOw=="],
"@mozilla/readability": ["@mozilla/readability@0.6.0", "", {}, "sha512-juG5VWh4qAivzTAeMzvY9xs9HY5rAcr2E4I7tiSSCokRFi7XIZCAu92ZkSTsIj1OPceCifL3cpfteP3pDT9/QQ=="], "@mozilla/readability": ["@mozilla/readability@0.6.0", "", {}, "sha512-juG5VWh4qAivzTAeMzvY9xs9HY5rAcr2E4I7tiSSCokRFi7XIZCAu92ZkSTsIj1OPceCifL3cpfteP3pDT9/QQ=="],
"@nodelib/fs.scandir": ["@nodelib/fs.scandir@2.1.5", "", { "dependencies": { "@nodelib/fs.stat": "2.0.5", "run-parallel": "^1.1.9" } }, "sha512-vq24Bq3ym5HEQm2NKCr3yXDwjc7vTsEThRDnkp2DK9p1uqLR+DHurm/NOTo0KG7HYHU7eppKZj3MyqYuMBf62g=="],
"@nodelib/fs.stat": ["@nodelib/fs.stat@2.0.5", "", {}, "sha512-RkhPPp2zrqDAQA/2jNhnztcPAlv64XdhIp7a7454A5ovI7Bukxgt7MX7udwAu3zg1DcpPU0rz3VV1SeaqvY4+A=="],
"@nodelib/fs.walk": ["@nodelib/fs.walk@1.2.8", "", { "dependencies": { "@nodelib/fs.scandir": "2.1.5", "fastq": "^1.6.0" } }, "sha512-oGB+UxlgWcgQkgwo8GcEGwemoTFt3FIO9ababBmaGwXIoBKZ+GTy0pP185beGg7Llih/NSHSV2XAs1lnznocSg=="],
"@types/bun": ["@types/bun@1.3.11", "", { "dependencies": { "bun-types": "1.3.11" } }, "sha512-5vPne5QvtpjGpsGYXiFyycfpDF2ECyPcTSsFBMa0fraoxiQyMJ3SmuQIGhzPg2WJuWxVBoxWJ2kClYTcw/4fAg=="],
"@types/debug": ["@types/debug@4.1.13", "", { "dependencies": { "@types/ms": "*" } }, "sha512-KSVgmQmzMwPlmtljOomayoR89W4FynCAi3E8PPs7vmDVPe84hT+vGPKkJfThkmXs0x0jAaa9U8uW8bbfyS2fWw=="],
"@types/jsdom": ["@types/jsdom@21.1.7", "", { "dependencies": { "@types/node": "*", "@types/tough-cookie": "*", "parse5": "^7.0.0" } }, "sha512-yOriVnggzrnQ3a9OKOCxaVuSug3w3/SbOj5i7VwXWZEyUNl3bLF9V3MfxGbZKuwqJOQyRfqXyROBB1CoZLFWzA=="],
"@types/mdast": ["@types/mdast@4.0.4", "", { "dependencies": { "@types/unist": "*" } }, "sha512-kGaNbPh1k7AFzgpud/gMdvIm5xuECykRR+JnWKQno9TAXVa6WIVCGTPvYGekIDL4uwCZQSYbUxNBSb1aUo79oA=="],
"@types/ms": ["@types/ms@2.1.0", "", {}, "sha512-GsCCIZDE/p3i96vtEqx+7dBUGXrc7zeSK3wwPHIaRThS+9OhWIXRqzs4d6k1SVU8g91DrNRWxWUGhp5KXQb2VA=="],
"@types/node": ["@types/node@25.5.0", "", { "dependencies": { "undici-types": "~7.18.0" } }, "sha512-jp2P3tQMSxWugkCUKLRPVUpGaL5MVFwF8RDuSRztfwgN1wmqJeMSbKlnEtQqU8UrhTmzEmZdu2I6v2dpp7XIxw=="],
"@types/tough-cookie": ["@types/tough-cookie@4.0.5", "", {}, "sha512-/Ad8+nIOV7Rl++6f1BdKxFSMgmoqEoYbHRpPcx3JEfv8VRsQe9Z4mCXeJBzxs7mbHY/XOZZuXlRNfhpVPbs6ZA=="],
"@types/unist": ["@types/unist@3.0.3", "", {}, "sha512-ko/gIFJRv177XgZsZcBwnqJN5x/Gien8qNOn0D5bQU/zAzVf9Zt3BlcUiLqhV9y4ARk0GbT3tnUiPNgnTXzc/Q=="],
"@types/ws": ["@types/ws@8.18.1", "", { "dependencies": { "@types/node": "*" } }, "sha512-ThVF6DCVhA8kUGy+aazFQ4kXQ7E1Ty7A3ypFOe0IcJV8O/M511G99AW24irKrW56Wt44yG9+ij8FaqoBGkuBXg=="],
"@xmldom/xmldom": ["@xmldom/xmldom@0.8.11", "", {}, "sha512-cQzWCtO6C8TQiYl1ruKNn2U6Ao4o4WBBcbL61yJl84x+j5sOWWFU9X7DpND8XZG3daDppSsigMdfAIl2upQBRw=="], "@xmldom/xmldom": ["@xmldom/xmldom@0.8.11", "", {}, "sha512-cQzWCtO6C8TQiYl1ruKNn2U6Ao4o4WBBcbL61yJl84x+j5sOWWFU9X7DpND8XZG3daDppSsigMdfAIl2upQBRw=="],
"agent-base": ["agent-base@7.1.4", "", {}, "sha512-MnA+YT8fwfJPgBx3m60MNqakm30XOkyIoH1y6huTQvC0PwZG7ki8NacLBcrPbNoo8vEZy7Jpuk7+jMO+CUovTQ=="], "agent-base": ["agent-base@7.1.4", "", {}, "sha512-MnA+YT8fwfJPgBx3m60MNqakm30XOkyIoH1y6huTQvC0PwZG7ki8NacLBcrPbNoo8vEZy7Jpuk7+jMO+CUovTQ=="],
"asynckit": ["asynckit@0.4.0", "", {}, "sha512-Oei9OH4tRh0YqU3GxhX79dM/mwVgvbZJaSNaRk+bshkj0S5cfHcgYakreBjrHwatXKbz+IoIdYLxrKim2MjW0Q=="], "ansi-colors": ["ansi-colors@4.1.3", "", {}, "sha512-/6w/C21Pm1A7aZitlI5Ni/2J6FFQN8i1Cvz3kHABAAbw93v/NlvKdVOqz7CCWz/3iv/JplRSEEZ83XION15ovw=="],
"baoyu-chrome-cdp": ["baoyu-chrome-cdp@file:vendor/baoyu-chrome-cdp", {}], "ansi-regex": ["ansi-regex@5.0.1", "", {}, "sha512-quJQXlTSUGL2LH9SUXo8VwsY4soanhgo6LNSm84E1LBcE8s3O0wpdiRzyR9z/ZZJMlMWv37qOOb9pdJlMUEKFQ=="],
"argparse": ["argparse@2.0.1", "", {}, "sha512-8+9WqebbFzpX9OR+Wa6O29asIogeRMzcGtAINdpMHHyAg10f05aSFVBbcEqGf/PXw1EjAZ+q2/bEBg3DvurK3Q=="],
"array-union": ["array-union@2.1.0", "", {}, "sha512-HGyxoOTYUyCM6stUe6EJgnd4EoewAI7zMdfqO+kGjnlZmBDz/cR5pf8r/cR4Wq60sL/p0IkcjUEEPwS3GFrIyw=="],
"bail": ["bail@2.0.2", "", {}, "sha512-0xO6mYd7JB2YesxDKplafRpsiOzPt9V02ddPCLbY1xYGPOX24NTyN50qnUxgCPcSoYMhKpAuBTjQoRZCAkUDRw=="],
"baoyu-fetch": ["baoyu-fetch@file:vendor/baoyu-fetch", { "dependencies": { "@mozilla/readability": "^0.6.0", "chrome-launcher": "^1.2.1", "defuddle": "^0.14.0", "jsdom": "^26.0.0", "remark-gfm": "^4.0.1", "remark-parse": "^11.0.0", "turndown": "^7.2.0", "turndown-plugin-gfm": "^1.0.2", "unified": "^11.0.5", "ws": "^8.18.3" }, "devDependencies": { "@changesets/cli": "^2.30.0", "@types/bun": "^1.2.23", "@types/jsdom": "^21.1.7", "@types/ws": "^8.18.1", "typescript": "^5.9.2" }, "bin": { "baoyu-fetch": "./src/cli.ts" } }],
"better-path-resolve": ["better-path-resolve@1.0.0", "", { "dependencies": { "is-windows": "^1.0.0" } }, "sha512-pbnl5XzGBdrFU/wT4jqmJVPn2B6UHPBOhzMQkY/SPUPB6QtUXtmBHBIwCbXJol93mOpGMnQyP/+BB19q04xj7g=="],
"boolbase": ["boolbase@1.0.0", "", {}, "sha512-JZOSA7Mo9sNGB8+UjSgzdLtokWAky1zbztM3WRLCbZ70/3cTANmQmOdR7y2g+J0e2WXywy1yS468tY+IruqEww=="], "boolbase": ["boolbase@1.0.0", "", {}, "sha512-JZOSA7Mo9sNGB8+UjSgzdLtokWAky1zbztM3WRLCbZ70/3cTANmQmOdR7y2g+J0e2WXywy1yS468tY+IruqEww=="],
"call-bind-apply-helpers": ["call-bind-apply-helpers@1.0.2", "", { "dependencies": { "es-errors": "^1.3.0", "function-bind": "^1.1.2" } }, "sha512-Sp1ablJ0ivDkSzjcaJdxEunN5/XvksFJ2sMBFfq6x0ryhQV/2b/KwFe21cMpmHtPOSij8K99/wSfoEuTObmuMQ=="], "braces": ["braces@3.0.3", "", { "dependencies": { "fill-range": "^7.1.1" } }, "sha512-yQbXgO/OSZVD2IsiLlro+7Hf6Q18EJrKSEsdoMzKePKXct3gvD8oLcOQdIzGupr5Fj+EDe8gO/lxc1BzfMpxvA=="],
"combined-stream": ["combined-stream@1.0.8", "", { "dependencies": { "delayed-stream": "~1.0.0" } }, "sha512-FQN4MRfuJeHf7cBbBMJFXhKSDq+2kAArBlmRBvcvFE5BB1HZKXtSFASDhdlz9zOYwxh8lDdnvmMOe/+5cdoEdg=="], "bun-types": ["bun-types@1.3.11", "", { "dependencies": { "@types/node": "*" } }, "sha512-1KGPpoxQWl9f6wcZh57LvrPIInQMn2TQ7jsgxqpRzg+l0QPOFvJVH7HmvHo/AiPgwXy+/Thf6Ov3EdVn1vOabg=="],
"ccount": ["ccount@2.0.1", "", {}, "sha512-eyrF0jiFpY+3drT6383f1qhkbGsLSifNAjA61IUjZjmLCWjItY6LB9ft9YhoDgwfmclB2zhu51Lc7+95b8NRAg=="],
"character-entities": ["character-entities@2.0.2", "", {}, "sha512-shx7oQ0Awen/BRIdkjkvz54PnEEI/EjwXDSIZp86/KKdbafHh1Df/RYGBhn4hbe2+uKC9FnT5UCEdyPz3ai9hQ=="],
"chardet": ["chardet@2.1.1", "", {}, "sha512-PsezH1rqdV9VvyNhxxOW32/d75r01NY7TQCmOqomRo15ZSOKbpTFVsfjghxo6JloQUCGnH4k1LGu0R4yCLlWQQ=="],
"chrome-launcher": ["chrome-launcher@1.2.1", "", { "dependencies": { "@types/node": "*", "escape-string-regexp": "^4.0.0", "is-wsl": "^2.2.0", "lighthouse-logger": "^2.0.1" }, "bin": { "print-chrome-path": "bin/print-chrome-path.cjs" } }, "sha512-qmFR5PLMzHyuNJHwOloHPAHhbaNglkfeV/xDtt5b7xiFFyU1I+AZZX0PYseMuhenJSSirgxELYIbswcoc+5H4A=="],
"commander": ["commander@12.1.0", "", {}, "sha512-Vw8qHK3bZM9y/P10u3Vib8o/DdkvA2OtPtZvD871QKjy74Wj1WSKFILMPRPSdUSx5RFK1arlJzEtA4PkFgnbuA=="], "commander": ["commander@12.1.0", "", {}, "sha512-Vw8qHK3bZM9y/P10u3Vib8o/DdkvA2OtPtZvD871QKjy74Wj1WSKFILMPRPSdUSx5RFK1arlJzEtA4PkFgnbuA=="],
"cross-spawn": ["cross-spawn@7.0.6", "", { "dependencies": { "path-key": "^3.1.0", "shebang-command": "^2.0.0", "which": "^2.0.1" } }, "sha512-uV2QOWP2nWzsy2aMp8aRibhi9dlzF5Hgh5SHaB9OiTGEyDTiJJyx0uy51QXdyWbtAHNua4XJzUKca3OzKUd3vA=="],
"css-select": ["css-select@5.2.2", "", { "dependencies": { "boolbase": "^1.0.0", "css-what": "^6.1.0", "domhandler": "^5.0.2", "domutils": "^3.0.1", "nth-check": "^2.0.1" } }, "sha512-TizTzUddG/xYLA3NXodFM0fSbNizXjOKhqiQQwvhlspadZokn1KDy0NZFS0wuEubIYAV5/c1/lAr0TaaFXEXzw=="], "css-select": ["css-select@5.2.2", "", { "dependencies": { "boolbase": "^1.0.0", "css-what": "^6.1.0", "domhandler": "^5.0.2", "domutils": "^3.0.1", "nth-check": "^2.0.1" } }, "sha512-TizTzUddG/xYLA3NXodFM0fSbNizXjOKhqiQQwvhlspadZokn1KDy0NZFS0wuEubIYAV5/c1/lAr0TaaFXEXzw=="],
"css-what": ["css-what@6.2.2", "", {}, "sha512-u/O3vwbptzhMs3L1fQE82ZSLHQQfto5gyZzwteVIEyeaY5Fc7R4dapF/BvRoSYFeqfBk4m0V1Vafq5Pjv25wvA=="], "css-what": ["css-what@6.2.2", "", {}, "sha512-u/O3vwbptzhMs3L1fQE82ZSLHQQfto5gyZzwteVIEyeaY5Fc7R4dapF/BvRoSYFeqfBk4m0V1Vafq5Pjv25wvA=="],
@@ -61,9 +141,17 @@
"decimal.js": ["decimal.js@10.6.0", "", {}, "sha512-YpgQiITW3JXGntzdUmyUR1V812Hn8T1YVXhCu+wO3OpS4eU9l4YdD3qjyiKdV6mvV29zapkMeD390UVEf2lkUg=="], "decimal.js": ["decimal.js@10.6.0", "", {}, "sha512-YpgQiITW3JXGntzdUmyUR1V812Hn8T1YVXhCu+wO3OpS4eU9l4YdD3qjyiKdV6mvV29zapkMeD390UVEf2lkUg=="],
"decode-named-character-reference": ["decode-named-character-reference@1.3.0", "", { "dependencies": { "character-entities": "^2.0.0" } }, "sha512-GtpQYB283KrPp6nRw50q3U9/VfOutZOe103qlN7BPP6Ad27xYnOIWv4lPzo8HCAL+mMZofJ9KEy30fq6MfaK6Q=="],
"defuddle": ["defuddle@0.14.0", "", { "dependencies": { "commander": "^12.1.0" }, "optionalDependencies": { "linkedom": "^0.18.12", "mathml-to-latex": "^1.5.0", "temml": "^0.13.1", "turndown": "^7.2.0" }, "bin": { "defuddle": "dist/cli.js" } }, "sha512-btavZGd1WgiVqrVM62WGRXMUi/aU7ckTZiq0xXWLZMHvzIqNZjwIFQEDRx8MarD7fIgsB90NXZ9xHJkKtapt2Q=="], "defuddle": ["defuddle@0.14.0", "", { "dependencies": { "commander": "^12.1.0" }, "optionalDependencies": { "linkedom": "^0.18.12", "mathml-to-latex": "^1.5.0", "temml": "^0.13.1", "turndown": "^7.2.0" }, "bin": { "defuddle": "dist/cli.js" } }, "sha512-btavZGd1WgiVqrVM62WGRXMUi/aU7ckTZiq0xXWLZMHvzIqNZjwIFQEDRx8MarD7fIgsB90NXZ9xHJkKtapt2Q=="],
"delayed-stream": ["delayed-stream@1.0.0", "", {}, "sha512-ZySD7Nf91aLB0RxL4KGrKHBXl7Eds1DAmEdcoVawXnLD7SDhpNgtuII2aAkg7a7QS41jxPSZ17p4VdGnMHk3MQ=="], "dequal": ["dequal@2.0.3", "", {}, "sha512-0je+qPKHEMohvfRTCEo3CrPG6cAzAYgmzKyxRiYSSDkS6eGJdyVJm7WaYA5ECaAD9wLB2T4EEeymA5aFVcYXCA=="],
"detect-indent": ["detect-indent@6.1.0", "", {}, "sha512-reYkTUJAZb9gUuZ2RvVCNhVHdg62RHnJ7WJl8ftMi4diZ6NWlciOzQN88pUhSELEwflJht4oQDv0F0BMlwaYtA=="],
"devlop": ["devlop@1.1.0", "", { "dependencies": { "dequal": "^2.0.0" } }, "sha512-RWmIqhcFf1lRYBvNmr7qTNuyCt/7/ns2jbpp1+PalgE/rDQcBT0fioSMUpJ93irlUhC5hrg4cYqe6U+0ImW0rA=="],
"dir-glob": ["dir-glob@3.0.1", "", { "dependencies": { "path-type": "^4.0.0" } }, "sha512-WkrWp9GR4KXfKGYzOLmTuGVi1UWFfws377n9cc55/tb6DuqyF6pcQ5AbiHEshaDpY9v6oaSr2XCDidGmMwdzIA=="],
"dom-serializer": ["dom-serializer@2.0.0", "", { "dependencies": { "domelementtype": "^2.3.0", "domhandler": "^5.0.2", "entities": "^4.2.0" } }, "sha512-wIkAryiqt/nV5EQKqQpo3SToSOV9J0DnbJqwK7Wv/Trc92zIAYZ4FlMu+JPFW1DfGFt81ZTCGgDEabffXeLyJg=="], "dom-serializer": ["dom-serializer@2.0.0", "", { "dependencies": { "domelementtype": "^2.3.0", "domhandler": "^5.0.2", "entities": "^4.2.0" } }, "sha512-wIkAryiqt/nV5EQKqQpo3SToSOV9J0DnbJqwK7Wv/Trc92zIAYZ4FlMu+JPFW1DfGFt81ZTCGgDEabffXeLyJg=="],
@@ -73,33 +161,33 @@
"domutils": ["domutils@3.2.2", "", { "dependencies": { "dom-serializer": "^2.0.0", "domelementtype": "^2.3.0", "domhandler": "^5.0.3" } }, "sha512-6kZKyUajlDuqlHKVX1w7gyslj9MPIXzIFiz/rGu35uC1wMi+kMhQwGhl4lt9unC9Vb9INnY9Z3/ZA3+FhASLaw=="], "domutils": ["domutils@3.2.2", "", { "dependencies": { "dom-serializer": "^2.0.0", "domelementtype": "^2.3.0", "domhandler": "^5.0.3" } }, "sha512-6kZKyUajlDuqlHKVX1w7gyslj9MPIXzIFiz/rGu35uC1wMi+kMhQwGhl4lt9unC9Vb9INnY9Z3/ZA3+FhASLaw=="],
"dunder-proto": ["dunder-proto@1.0.1", "", { "dependencies": { "call-bind-apply-helpers": "^1.0.1", "es-errors": "^1.3.0", "gopd": "^1.2.0" } }, "sha512-KIN/nDJBQRcXw0MLVhZE9iQHmG68qAVIBg9CqmUYjmQIhgij9U5MFvrqkUL5FbtyyzZuOeOt0zdeRe4UY7ct+A=="], "enquirer": ["enquirer@2.4.1", "", { "dependencies": { "ansi-colors": "^4.1.1", "strip-ansi": "^6.0.1" } }, "sha512-rRqJg/6gd538VHvR3PSrdRBb/1Vy2YfzHqzvbhGIQpDRKIa4FgV/54b5Q1xYSxOOwKvjXweS26E0Q+nAMwp2pQ=="],
"entities": ["entities@6.0.1", "", {}, "sha512-aN97NXWF6AWBTahfVOIrB/NShkzi5H7F9r1s9mD3cDj4Ko5f2qhhVoYMibXF7GlLveb/D2ioWay8lxI97Ven3g=="], "entities": ["entities@6.0.1", "", {}, "sha512-aN97NXWF6AWBTahfVOIrB/NShkzi5H7F9r1s9mD3cDj4Ko5f2qhhVoYMibXF7GlLveb/D2ioWay8lxI97Ven3g=="],
"es-define-property": ["es-define-property@1.0.1", "", {}, "sha512-e3nRfgfUZ4rNGL232gUgX06QNyyez04KdjFrF+LTRoOXmrOgFKDg4BCdsjW8EnT69eqdYGmRpJwiPVYNrCaW3g=="], "escape-string-regexp": ["escape-string-regexp@4.0.0", "", {}, "sha512-TtpcNJ3XAzx3Gq8sWRzJaVajRs0uVxA2YAkdb1jm2YkPz4G6egUFAyA3n5vtEIZefPk5Wa4UXbKuS5fKkJWdgA=="],
"es-errors": ["es-errors@1.3.0", "", {}, "sha512-Zf5H2Kxt2xjTvbJvP2ZWLEICxA6j+hAmMzIlypy4xcBg1vKVnx89Wy0GbS+kf5cwCVFFzdCFh2XSCFNULS6csw=="], "esprima": ["esprima@4.0.1", "", { "bin": { "esparse": "./bin/esparse.js", "esvalidate": "./bin/esvalidate.js" } }, "sha512-eGuFFw7Upda+g4p+QHvnW0RyTX/SVeJBDM/gCtMARO0cLuT2HcEKnTPvhjV6aGeqrCB/sbNop0Kszm0jsaWU4A=="],
"es-object-atoms": ["es-object-atoms@1.1.1", "", { "dependencies": { "es-errors": "^1.3.0" } }, "sha512-FGgH2h8zKNim9ljj7dankFPcICIK9Cp5bm+c2gQSYePhpaG5+esrLODihIorn+Pe6FGJzWhXQotPv73jTaldXA=="], "extend": ["extend@3.0.2", "", {}, "sha512-fjquC59cD7CyW6urNXK0FBufkZcoiGG80wTuPujX590cB5Ttln20E2UB4S/WARVqhXffZl2LNgS+gQdPIIim/g=="],
"es-set-tostringtag": ["es-set-tostringtag@2.1.0", "", { "dependencies": { "es-errors": "^1.3.0", "get-intrinsic": "^1.2.6", "has-tostringtag": "^1.0.2", "hasown": "^2.0.2" } }, "sha512-j6vWzfrGVfyXxge+O0x5sh6cvxAog0a/4Rdd2K36zCMV5eJ+/+tOAngRO8cODMNWbVRdVlmGZQL2YS3yR8bIUA=="], "extendable-error": ["extendable-error@0.1.7", "", {}, "sha512-UOiS2in6/Q0FK0R0q6UY9vYpQ21mr/Qn1KOnte7vsACuNJf514WvCCUHSRCPcgjPT2bAhNIJdlE6bVap1GKmeg=="],
"form-data": ["form-data@4.0.5", "", { "dependencies": { "asynckit": "^0.4.0", "combined-stream": "^1.0.8", "es-set-tostringtag": "^2.1.0", "hasown": "^2.0.2", "mime-types": "^2.1.12" } }, "sha512-8RipRLol37bNs2bhoV67fiTEvdTrbMUYcFTiy3+wuuOnUog2QBHCZWXDRijWQfAkhBj2Uf5UnVaiWwA5vdd82w=="], "fast-glob": ["fast-glob@3.3.3", "", { "dependencies": { "@nodelib/fs.stat": "^2.0.2", "@nodelib/fs.walk": "^1.2.3", "glob-parent": "^5.1.2", "merge2": "^1.3.0", "micromatch": "^4.0.8" } }, "sha512-7MptL8U0cqcFdzIzwOTHoilX9x5BrNqye7Z/LuC7kCMRio1EMSyqRK3BEAUD7sXRq4iT4AzTVuZdhgQ2TCvYLg=="],
"function-bind": ["function-bind@1.1.2", "", {}, "sha512-7XHNxH7qX9xG5mIwxkhumTox/MIRNcOgDrxWsMt2pAr23WHp6MrRlN7FBSFpCpr+oVO0F744iUgR82nJMfG2SA=="], "fastq": ["fastq@1.20.1", "", { "dependencies": { "reusify": "^1.0.4" } }, "sha512-GGToxJ/w1x32s/D2EKND7kTil4n8OVk/9mycTc4VDza13lOvpUZTGX3mFSCtV9ksdGBVzvsyAVLM6mHFThxXxw=="],
"get-intrinsic": ["get-intrinsic@1.3.0", "", { "dependencies": { "call-bind-apply-helpers": "^1.0.2", "es-define-property": "^1.0.1", "es-errors": "^1.3.0", "es-object-atoms": "^1.1.1", "function-bind": "^1.1.2", "get-proto": "^1.0.1", "gopd": "^1.2.0", "has-symbols": "^1.1.0", "hasown": "^2.0.2", "math-intrinsics": "^1.1.0" } }, "sha512-9fSjSaos/fRIVIp+xSJlE6lfwhES7LNtKaCBIamHsjr2na1BiABJPo0mOjjz8GJDURarmCPGqaiVg5mfjb98CQ=="], "fill-range": ["fill-range@7.1.1", "", { "dependencies": { "to-regex-range": "^5.0.1" } }, "sha512-YsGpe3WHLK8ZYi4tWDg2Jy3ebRz2rXowDxnld4bkQB00cc/1Zw9AWnC0i9ztDJitivtQvaI9KaLyKrc+hBW0yg=="],
"get-proto": ["get-proto@1.0.1", "", { "dependencies": { "dunder-proto": "^1.0.1", "es-object-atoms": "^1.0.0" } }, "sha512-sTSfBjoXBp89JvIKIefqw7U2CCebsc74kiY6awiGogKtoSGbgjYE/G/+l9sF3MWFPNc9IcoOC4ODfKHfxFmp0g=="], "find-up": ["find-up@4.1.0", "", { "dependencies": { "locate-path": "^5.0.0", "path-exists": "^4.0.0" } }, "sha512-PpOwAdQ/YlXQ2vj8a3h8IipDuYRi3wceVQQGYWxNINccq40Anw7BlsEXCMbt1Zt+OLA6Fq9suIpIWD0OsnISlw=="],
"gopd": ["gopd@1.2.0", "", {}, "sha512-ZUKRh6/kUFoAiTAtTYPZJ3hw9wNxx+BIBOijnlG9PnrJsCcSjs1wyyD6vJpaYtgnzDrKYRSqf3OO6Rfa93xsRg=="], "fs-extra": ["fs-extra@7.0.1", "", { "dependencies": { "graceful-fs": "^4.1.2", "jsonfile": "^4.0.0", "universalify": "^0.1.0" } }, "sha512-YJDaCJZEnBmcbw13fvdAM9AwNOJwOzrE4pqMqBq5nFiEqXUqHwlK4B+3pUw6JNvfSPtX05xFHtYy/1ni01eGCw=="],
"has-symbols": ["has-symbols@1.1.0", "", {}, "sha512-1cDNdwJ2Jaohmb3sg4OmKaMBwuC48sYni5HUw2DvsC8LjGTLK9h+eb1X6RyuOHe4hT0ULCW68iomhjUoKUqlPQ=="], "glob-parent": ["glob-parent@5.1.2", "", { "dependencies": { "is-glob": "^4.0.1" } }, "sha512-AOIgSQCepiJYwP3ARnGx+5VnTu2HBYdzbGP45eLw1vr3zB3vZLeyed1sC9hnbcOc9/SrMyM5RPQrkGz4aS9Zow=="],
"has-tostringtag": ["has-tostringtag@1.0.2", "", { "dependencies": { "has-symbols": "^1.0.3" } }, "sha512-NqADB8VjPFLM2V0VvHUewwwsw0ZWBaIdgo+ieHtK3hasLz4qeCRjYcqfB6AQrBggRKppKF8L52/VqdVsO47Dlw=="], "globby": ["globby@11.1.0", "", { "dependencies": { "array-union": "^2.1.0", "dir-glob": "^3.0.1", "fast-glob": "^3.2.9", "ignore": "^5.2.0", "merge2": "^1.4.1", "slash": "^3.0.0" } }, "sha512-jhIXaOzy1sb8IyocaruWSn1TjmnBVs8Ayhcy83rmxNJ8q2uWKCAj3CnJY+KpGSXCueAPc0i05kVvVKtP1t9S3g=="],
"hasown": ["hasown@2.0.2", "", { "dependencies": { "function-bind": "^1.1.2" } }, "sha512-0hJU9SCPvmMzIBdZFqNPXWa6dqh7WdH0cII9y+CyS8rG3nL48Bclra9HmKhVVUHyPWNH5Y7xDwAB7bfgSjkUMQ=="], "graceful-fs": ["graceful-fs@4.2.11", "", {}, "sha512-RbJ5/jmFcNNCcDV5o9eTnBLJ/HszWV0P73bc+Ff4nS/rJj+YaS6IGyiOL0VoBYX+l1Wrl3k63h/KrH+nhJ0XvQ=="],
"html-encoding-sniffer": ["html-encoding-sniffer@4.0.0", "", { "dependencies": { "whatwg-encoding": "^3.1.1" } }, "sha512-Y22oTqIU4uuPgEemfz7NDJz6OeKf12Lsu+QC+s3BVpda64lTiMYCyGwg5ki4vFxkMwQdeZDl2adZoqUgdFuTgQ=="], "html-encoding-sniffer": ["html-encoding-sniffer@4.0.0", "", { "dependencies": { "whatwg-encoding": "^3.1.1" } }, "sha512-Y22oTqIU4uuPgEemfz7NDJz6OeKf12Lsu+QC+s3BVpda64lTiMYCyGwg5ki4vFxkMwQdeZDl2adZoqUgdFuTgQ=="],
@@ -111,23 +199,139 @@
"https-proxy-agent": ["https-proxy-agent@7.0.6", "", { "dependencies": { "agent-base": "^7.1.2", "debug": "4" } }, "sha512-vK9P5/iUfdl95AI+JVyUuIcVtd4ofvtrOr3HNtM2yxC9bnMbEdp3x01OhQNnjb8IJYi38VlTE3mBXwcfvywuSw=="], "https-proxy-agent": ["https-proxy-agent@7.0.6", "", { "dependencies": { "agent-base": "^7.1.2", "debug": "4" } }, "sha512-vK9P5/iUfdl95AI+JVyUuIcVtd4ofvtrOr3HNtM2yxC9bnMbEdp3x01OhQNnjb8IJYi38VlTE3mBXwcfvywuSw=="],
"iconv-lite": ["iconv-lite@0.6.3", "", { "dependencies": { "safer-buffer": ">= 2.1.2 < 3.0.0" } }, "sha512-4fCk79wshMdzMp2rH06qWrJE4iolqLhCUH+OiuIgU++RB0+94NlDL81atO7GX55uUKueo0txHNtvEyI6D7WdMw=="], "human-id": ["human-id@4.1.3", "", { "bin": { "human-id": "dist/cli.js" } }, "sha512-tsYlhAYpjCKa//8rXZ9DqKEawhPoSytweBC2eNvcaDK+57RZLHGqNs3PZTQO6yekLFSuvA6AlnAfrw1uBvtb+Q=="],
"iconv-lite": ["iconv-lite@0.7.2", "", { "dependencies": { "safer-buffer": ">= 2.1.2 < 3.0.0" } }, "sha512-im9DjEDQ55s9fL4EYzOAv0yMqmMBSZp6G0VvFyTMPKWxiSBHUj9NW/qqLmXUwXrrM7AvqSlTCfvqRb0cM8yYqw=="],
"ignore": ["ignore@5.3.2", "", {}, "sha512-hsBTNUqQTDwkWtcdYI2i06Y/nUBEsNEDJKjWdigLvegy8kDuJAS8uRlpkkcQpyEXL0Z/pjDy5HBmMjRCJ2gq+g=="],
"is-docker": ["is-docker@2.2.1", "", { "bin": { "is-docker": "cli.js" } }, "sha512-F+i2BKsFrH66iaUFc0woD8sLy8getkwTwtOBjvs56Cx4CgJDeKQeqfz8wAYiSb8JOprWhHH5p77PbmYCvvUuXQ=="],
"is-extglob": ["is-extglob@2.1.1", "", {}, "sha512-SbKbANkN603Vi4jEZv49LeVJMn4yGwsbzZworEoyEiutsN3nJYdbO36zfhGJ6QEDpOZIFkDtnq5JRxmvl3jsoQ=="],
"is-glob": ["is-glob@4.0.3", "", { "dependencies": { "is-extglob": "^2.1.1" } }, "sha512-xelSayHH36ZgE7ZWhli7pW34hNbNl8Ojv5KVmkJD4hBdD3th8Tfk9vYasLM+mXWOZhFkgZfxhLSnrwRr4elSSg=="],
"is-number": ["is-number@7.0.0", "", {}, "sha512-41Cifkg6e8TylSpdtTpeLVMqvSBEVzTttHvERD741+pnZ8ANv0004MRL43QKPDlK9cGvNp6NZWZUBlbGXYxxng=="],
"is-plain-obj": ["is-plain-obj@4.1.0", "", {}, "sha512-+Pgi+vMuUNkJyExiMBt5IlFoMyKnr5zhJ4Uspz58WOhBF5QoIZkFyNHIbBAtHwzVAgk5RtndVNsDRN61/mmDqg=="],
"is-potential-custom-element-name": ["is-potential-custom-element-name@1.0.1", "", {}, "sha512-bCYeRA2rVibKZd+s2625gGnGF/t7DSqDs4dP7CrLA1m7jKWz6pps0LpYLJN8Q64HtmPKJ1hrN3nzPNKFEKOUiQ=="], "is-potential-custom-element-name": ["is-potential-custom-element-name@1.0.1", "", {}, "sha512-bCYeRA2rVibKZd+s2625gGnGF/t7DSqDs4dP7CrLA1m7jKWz6pps0LpYLJN8Q64HtmPKJ1hrN3nzPNKFEKOUiQ=="],
"jsdom": ["jsdom@24.1.3", "", { "dependencies": { "cssstyle": "^4.0.1", "data-urls": "^5.0.0", "decimal.js": "^10.4.3", "form-data": "^4.0.0", "html-encoding-sniffer": "^4.0.0", "http-proxy-agent": "^7.0.2", "https-proxy-agent": "^7.0.5", "is-potential-custom-element-name": "^1.0.1", "nwsapi": "^2.2.12", "parse5": "^7.1.2", "rrweb-cssom": "^0.7.1", "saxes": "^6.0.0", "symbol-tree": "^3.2.4", "tough-cookie": "^4.1.4", "w3c-xmlserializer": "^5.0.0", "webidl-conversions": "^7.0.0", "whatwg-encoding": "^3.1.1", "whatwg-mimetype": "^4.0.0", "whatwg-url": "^14.0.0", "ws": "^8.18.0", "xml-name-validator": "^5.0.0" }, "peerDependencies": { "canvas": "^2.11.2" }, "optionalPeers": ["canvas"] }, "sha512-MyL55p3Ut3cXbeBEG7Hcv0mVM8pp8PBNWxRqchZnSfAiES1v1mRnMeFfaHWIPULpwsYfvO+ZmMZz5tGCnjzDUQ=="], "is-subdir": ["is-subdir@1.2.0", "", { "dependencies": { "better-path-resolve": "1.0.0" } }, "sha512-2AT6j+gXe/1ueqbW6fLZJiIw3F8iXGJtt0yDrZaBhAZEG1raiTxKWU+IPqMCzQAXOUCKdA4UDMgacKH25XG2Cw=="],
"is-windows": ["is-windows@1.0.2", "", {}, "sha512-eXK1UInq2bPmjyX6e3VHIzMLobc4J94i4AWn+Hpq3OU5KkrRC96OAcR3PRJ/pGu6m8TRnBHP9dkXQVsT/COVIA=="],
"is-wsl": ["is-wsl@2.2.0", "", { "dependencies": { "is-docker": "^2.0.0" } }, "sha512-fKzAra0rGJUUBwGBgNkHZuToZcn+TtXHpeCgmkMJMMYx1sQDYaCSyjJBSCa2nH1DGm7s3n1oBnohoVTBaN7Lww=="],
"isexe": ["isexe@2.0.0", "", {}, "sha512-RHxMLp9lnKHGHRng9QFhRCMbYAcVpn69smSGcq3f36xjgVVWThj4qqLbTLlq7Ssj8B+fIQ1EuCEGI2lKsyQeIw=="],
"js-yaml": ["js-yaml@4.1.1", "", { "dependencies": { "argparse": "^2.0.1" }, "bin": { "js-yaml": "bin/js-yaml.js" } }, "sha512-qQKT4zQxXl8lLwBtHMWwaTcGfFOZviOJet3Oy/xmGk2gZH677CJM9EvtfdSkgWcATZhj/55JZ0rmy3myCT5lsA=="],
"jsdom": ["jsdom@26.1.0", "", { "dependencies": { "cssstyle": "^4.2.1", "data-urls": "^5.0.0", "decimal.js": "^10.5.0", "html-encoding-sniffer": "^4.0.0", "http-proxy-agent": "^7.0.2", "https-proxy-agent": "^7.0.6", "is-potential-custom-element-name": "^1.0.1", "nwsapi": "^2.2.16", "parse5": "^7.2.1", "rrweb-cssom": "^0.8.0", "saxes": "^6.0.0", "symbol-tree": "^3.2.4", "tough-cookie": "^5.1.1", "w3c-xmlserializer": "^5.0.0", "webidl-conversions": "^7.0.0", "whatwg-encoding": "^3.1.1", "whatwg-mimetype": "^4.0.0", "whatwg-url": "^14.1.1", "ws": "^8.18.0", "xml-name-validator": "^5.0.0" }, "peerDependencies": { "canvas": "^3.0.0" }, "optionalPeers": ["canvas"] }, "sha512-Cvc9WUhxSMEo4McES3P7oK3QaXldCfNWp7pl2NNeiIFlCoLr3kfq9kb1fxftiwk1FLV7CvpvDfonxtzUDeSOPg=="],
"jsonfile": ["jsonfile@4.0.0", "", { "optionalDependencies": { "graceful-fs": "^4.1.6" } }, "sha512-m6F1R3z8jjlf2imQHS2Qez5sjKWQzbuuhuJ/FKYFRZvPE3PuHcSMVZzfsLhGVOkfd20obL5SWEBew5ShlquNxg=="],
"lighthouse-logger": ["lighthouse-logger@2.0.2", "", { "dependencies": { "debug": "^4.4.1", "marky": "^1.2.2" } }, "sha512-vWl2+u5jgOQuZR55Z1WM0XDdrJT6mzMP8zHUct7xTlWhuQs+eV0g+QL0RQdFjT54zVmbhLCP8vIVpy1wGn/gCg=="],
"linkedom": ["linkedom@0.18.12", "", { "dependencies": { "css-select": "^5.1.0", "cssom": "^0.5.0", "html-escaper": "^3.0.3", "htmlparser2": "^10.0.0", "uhyphen": "^0.2.0" }, "peerDependencies": { "canvas": ">= 2" }, "optionalPeers": ["canvas"] }, "sha512-jalJsOwIKuQJSeTvsgzPe9iJzyfVaEJiEXl+25EkKevsULHvMJzpNqwvj1jOESWdmgKDiXObyjOYwlUqG7wo1Q=="], "linkedom": ["linkedom@0.18.12", "", { "dependencies": { "css-select": "^5.1.0", "cssom": "^0.5.0", "html-escaper": "^3.0.3", "htmlparser2": "^10.0.0", "uhyphen": "^0.2.0" }, "peerDependencies": { "canvas": ">= 2" }, "optionalPeers": ["canvas"] }, "sha512-jalJsOwIKuQJSeTvsgzPe9iJzyfVaEJiEXl+25EkKevsULHvMJzpNqwvj1jOESWdmgKDiXObyjOYwlUqG7wo1Q=="],
"locate-path": ["locate-path@5.0.0", "", { "dependencies": { "p-locate": "^4.1.0" } }, "sha512-t7hw9pI+WvuwNJXwk5zVHpyhIqzg2qTlklJOf0mVxGSbe3Fp2VieZcduNYjaLDoy6p9uGpQEGWG87WpMKlNq8g=="],
"lodash.startcase": ["lodash.startcase@4.4.0", "", {}, "sha512-+WKqsK294HMSc2jEbNgpHpd0JfIBhp7rEV4aqXWqFr6AlXov+SlcgB1Fv01y2kGe3Gc8nMW7VA0SrGuSkRfIEg=="],
"longest-streak": ["longest-streak@3.1.0", "", {}, "sha512-9Ri+o0JYgehTaVBBDoMqIl8GXtbWg711O3srftcHhZ0dqnETqLaoIK0x17fUw9rFSlK/0NlsKe0Ahhyl5pXE2g=="],
"lru-cache": ["lru-cache@10.4.3", "", {}, "sha512-JNAzZcXrCt42VGLuYz0zfAzDfAvJWW6AfYlDBQyDV5DClI2m5sAmK+OIO7s59XfsRsWHp02jAJrRadPRGTt6SQ=="], "lru-cache": ["lru-cache@10.4.3", "", {}, "sha512-JNAzZcXrCt42VGLuYz0zfAzDfAvJWW6AfYlDBQyDV5DClI2m5sAmK+OIO7s59XfsRsWHp02jAJrRadPRGTt6SQ=="],
"math-intrinsics": ["math-intrinsics@1.1.0", "", {}, "sha512-/IXtbwEk5HTPyEwyKX6hGkYXxM9nbj64B+ilVJnC/R6B0pH5G4V3b0pVbL7DBj4tkhBAppbQUlf6F6Xl9LHu1g=="], "markdown-table": ["markdown-table@3.0.4", "", {}, "sha512-wiYz4+JrLyb/DqW2hkFJxP7Vd7JuTDm77fvbM8VfEQdmSMqcImWeeRbHwZjBjIFki/VaMK2BhFi7oUUZeM5bqw=="],
"marky": ["marky@1.3.0", "", {}, "sha512-ocnPZQLNpvbedwTy9kNrQEsknEfgvcLMvOtz3sFeWApDq1MXH1TqkCIx58xlpESsfwQOnuBO9beyQuNGzVvuhQ=="],
"mathml-to-latex": ["mathml-to-latex@1.5.0", "", { "dependencies": { "@xmldom/xmldom": "^0.8.10" } }, "sha512-rrWn0eEvcEcdMM4xfHcSGIy+i01DX9byOdXTLWg+w1iJ6O6ohP5UXY1dVzNUZLhzfl3EGcRekWLhY7JT5Omaew=="], "mathml-to-latex": ["mathml-to-latex@1.5.0", "", { "dependencies": { "@xmldom/xmldom": "^0.8.10" } }, "sha512-rrWn0eEvcEcdMM4xfHcSGIy+i01DX9byOdXTLWg+w1iJ6O6ohP5UXY1dVzNUZLhzfl3EGcRekWLhY7JT5Omaew=="],
"mime-db": ["mime-db@1.52.0", "", {}, "sha512-sPU4uV7dYlvtWJxwwxHD0PuihVNiE7TyAbQ5SWxDCB9mUYvOgroQOwYQQOKPJ8CIbE+1ETVlOoK1UC2nU3gYvg=="], "mdast-util-find-and-replace": ["mdast-util-find-and-replace@3.0.2", "", { "dependencies": { "@types/mdast": "^4.0.0", "escape-string-regexp": "^5.0.0", "unist-util-is": "^6.0.0", "unist-util-visit-parents": "^6.0.0" } }, "sha512-Tmd1Vg/m3Xz43afeNxDIhWRtFZgM2VLyaf4vSTYwudTyeuTneoL3qtWMA5jeLyz/O1vDJmmV4QuScFCA2tBPwg=="],
"mime-types": ["mime-types@2.1.35", "", { "dependencies": { "mime-db": "1.52.0" } }, "sha512-ZDY+bPm5zTTF+YpCrAU9nK0UgICYPT0QtT1NZWFv4s++TNkcgVaT0g6+4R2uI4MjQjzysHB1zxuWL50hzaeXiw=="], "mdast-util-from-markdown": ["mdast-util-from-markdown@2.0.3", "", { "dependencies": { "@types/mdast": "^4.0.0", "@types/unist": "^3.0.0", "decode-named-character-reference": "^1.0.0", "devlop": "^1.0.0", "mdast-util-to-string": "^4.0.0", "micromark": "^4.0.0", "micromark-util-decode-numeric-character-reference": "^2.0.0", "micromark-util-decode-string": "^2.0.0", "micromark-util-normalize-identifier": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0", "unist-util-stringify-position": "^4.0.0" } }, "sha512-W4mAWTvSlKvf8L6J+VN9yLSqQ9AOAAvHuoDAmPkz4dHf553m5gVj2ejadHJhoJmcmxEnOv6Pa8XJhpxE93kb8Q=="],
"mdast-util-gfm": ["mdast-util-gfm@3.1.0", "", { "dependencies": { "mdast-util-from-markdown": "^2.0.0", "mdast-util-gfm-autolink-literal": "^2.0.0", "mdast-util-gfm-footnote": "^2.0.0", "mdast-util-gfm-strikethrough": "^2.0.0", "mdast-util-gfm-table": "^2.0.0", "mdast-util-gfm-task-list-item": "^2.0.0", "mdast-util-to-markdown": "^2.0.0" } }, "sha512-0ulfdQOM3ysHhCJ1p06l0b0VKlhU0wuQs3thxZQagjcjPrlFRqY215uZGHHJan9GEAXd9MbfPjFJz+qMkVR6zQ=="],
"mdast-util-gfm-autolink-literal": ["mdast-util-gfm-autolink-literal@2.0.1", "", { "dependencies": { "@types/mdast": "^4.0.0", "ccount": "^2.0.0", "devlop": "^1.0.0", "mdast-util-find-and-replace": "^3.0.0", "micromark-util-character": "^2.0.0" } }, "sha512-5HVP2MKaP6L+G6YaxPNjuL0BPrq9orG3TsrZ9YXbA3vDw/ACI4MEsnoDpn6ZNm7GnZgtAcONJyPhOP8tNJQavQ=="],
"mdast-util-gfm-footnote": ["mdast-util-gfm-footnote@2.1.0", "", { "dependencies": { "@types/mdast": "^4.0.0", "devlop": "^1.1.0", "mdast-util-from-markdown": "^2.0.0", "mdast-util-to-markdown": "^2.0.0", "micromark-util-normalize-identifier": "^2.0.0" } }, "sha512-sqpDWlsHn7Ac9GNZQMeUzPQSMzR6Wv0WKRNvQRg0KqHh02fpTz69Qc1QSseNX29bhz1ROIyNyxExfawVKTm1GQ=="],
"mdast-util-gfm-strikethrough": ["mdast-util-gfm-strikethrough@2.0.0", "", { "dependencies": { "@types/mdast": "^4.0.0", "mdast-util-from-markdown": "^2.0.0", "mdast-util-to-markdown": "^2.0.0" } }, "sha512-mKKb915TF+OC5ptj5bJ7WFRPdYtuHv0yTRxK2tJvi+BDqbkiG7h7u/9SI89nRAYcmap2xHQL9D+QG/6wSrTtXg=="],
"mdast-util-gfm-table": ["mdast-util-gfm-table@2.0.0", "", { "dependencies": { "@types/mdast": "^4.0.0", "devlop": "^1.0.0", "markdown-table": "^3.0.0", "mdast-util-from-markdown": "^2.0.0", "mdast-util-to-markdown": "^2.0.0" } }, "sha512-78UEvebzz/rJIxLvE7ZtDd/vIQ0RHv+3Mh5DR96p7cS7HsBhYIICDBCu8csTNWNO6tBWfqXPWekRuj2FNOGOZg=="],
"mdast-util-gfm-task-list-item": ["mdast-util-gfm-task-list-item@2.0.0", "", { "dependencies": { "@types/mdast": "^4.0.0", "devlop": "^1.0.0", "mdast-util-from-markdown": "^2.0.0", "mdast-util-to-markdown": "^2.0.0" } }, "sha512-IrtvNvjxC1o06taBAVJznEnkiHxLFTzgonUdy8hzFVeDun0uTjxxrRGVaNFqkU1wJR3RBPEfsxmU6jDWPofrTQ=="],
"mdast-util-phrasing": ["mdast-util-phrasing@4.1.0", "", { "dependencies": { "@types/mdast": "^4.0.0", "unist-util-is": "^6.0.0" } }, "sha512-TqICwyvJJpBwvGAMZjj4J2n0X8QWp21b9l0o7eXyVJ25YNWYbJDVIyD1bZXE6WtV6RmKJVYmQAKWa0zWOABz2w=="],
"mdast-util-to-markdown": ["mdast-util-to-markdown@2.1.2", "", { "dependencies": { "@types/mdast": "^4.0.0", "@types/unist": "^3.0.0", "longest-streak": "^3.0.0", "mdast-util-phrasing": "^4.0.0", "mdast-util-to-string": "^4.0.0", "micromark-util-classify-character": "^2.0.0", "micromark-util-decode-string": "^2.0.0", "unist-util-visit": "^5.0.0", "zwitch": "^2.0.0" } }, "sha512-xj68wMTvGXVOKonmog6LwyJKrYXZPvlwabaryTjLh9LuvovB/KAH+kvi8Gjj+7rJjsFi23nkUxRQv1KqSroMqA=="],
"mdast-util-to-string": ["mdast-util-to-string@4.0.0", "", { "dependencies": { "@types/mdast": "^4.0.0" } }, "sha512-0H44vDimn51F0YwvxSJSm0eCDOJTRlmN0R1yBh4HLj9wiV1Dn0QoXGbvFAWj2hSItVTlCmBF1hqKlIyUBVFLPg=="],
"merge2": ["merge2@1.4.1", "", {}, "sha512-8q7VEgMJW4J8tcfVPy8g09NcQwZdbwFEqhe/WZkoIzjn/3TGDwtOCYtXGxA3O8tPzpczCCDgv+P2P5y00ZJOOg=="],
"micromark": ["micromark@4.0.2", "", { "dependencies": { "@types/debug": "^4.0.0", "debug": "^4.0.0", "decode-named-character-reference": "^1.0.0", "devlop": "^1.0.0", "micromark-core-commonmark": "^2.0.0", "micromark-factory-space": "^2.0.0", "micromark-util-character": "^2.0.0", "micromark-util-chunked": "^2.0.0", "micromark-util-combine-extensions": "^2.0.0", "micromark-util-decode-numeric-character-reference": "^2.0.0", "micromark-util-encode": "^2.0.0", "micromark-util-normalize-identifier": "^2.0.0", "micromark-util-resolve-all": "^2.0.0", "micromark-util-sanitize-uri": "^2.0.0", "micromark-util-subtokenize": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-zpe98Q6kvavpCr1NPVSCMebCKfD7CA2NqZ+rykeNhONIJBpc1tFKt9hucLGwha3jNTNI8lHpctWJWoimVF4PfA=="],
"micromark-core-commonmark": ["micromark-core-commonmark@2.0.3", "", { "dependencies": { "decode-named-character-reference": "^1.0.0", "devlop": "^1.0.0", "micromark-factory-destination": "^2.0.0", "micromark-factory-label": "^2.0.0", "micromark-factory-space": "^2.0.0", "micromark-factory-title": "^2.0.0", "micromark-factory-whitespace": "^2.0.0", "micromark-util-character": "^2.0.0", "micromark-util-chunked": "^2.0.0", "micromark-util-classify-character": "^2.0.0", "micromark-util-html-tag-name": "^2.0.0", "micromark-util-normalize-identifier": "^2.0.0", "micromark-util-resolve-all": "^2.0.0", "micromark-util-subtokenize": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-RDBrHEMSxVFLg6xvnXmb1Ayr2WzLAWjeSATAoxwKYJV94TeNavgoIdA0a9ytzDSVzBy2YKFK+emCPOEibLeCrg=="],
"micromark-extension-gfm": ["micromark-extension-gfm@3.0.0", "", { "dependencies": { "micromark-extension-gfm-autolink-literal": "^2.0.0", "micromark-extension-gfm-footnote": "^2.0.0", "micromark-extension-gfm-strikethrough": "^2.0.0", "micromark-extension-gfm-table": "^2.0.0", "micromark-extension-gfm-tagfilter": "^2.0.0", "micromark-extension-gfm-task-list-item": "^2.0.0", "micromark-util-combine-extensions": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-vsKArQsicm7t0z2GugkCKtZehqUm31oeGBV/KVSorWSy8ZlNAv7ytjFhvaryUiCUJYqs+NoE6AFhpQvBTM6Q4w=="],
"micromark-extension-gfm-autolink-literal": ["micromark-extension-gfm-autolink-literal@2.1.0", "", { "dependencies": { "micromark-util-character": "^2.0.0", "micromark-util-sanitize-uri": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-oOg7knzhicgQ3t4QCjCWgTmfNhvQbDDnJeVu9v81r7NltNCVmhPy1fJRX27pISafdjL+SVc4d3l48Gb6pbRypw=="],
"micromark-extension-gfm-footnote": ["micromark-extension-gfm-footnote@2.1.0", "", { "dependencies": { "devlop": "^1.0.0", "micromark-core-commonmark": "^2.0.0", "micromark-factory-space": "^2.0.0", "micromark-util-character": "^2.0.0", "micromark-util-normalize-identifier": "^2.0.0", "micromark-util-sanitize-uri": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-/yPhxI1ntnDNsiHtzLKYnE3vf9JZ6cAisqVDauhp4CEHxlb4uoOTxOCJ+9s51bIB8U1N1FJ1RXOKTIlD5B/gqw=="],
"micromark-extension-gfm-strikethrough": ["micromark-extension-gfm-strikethrough@2.1.0", "", { "dependencies": { "devlop": "^1.0.0", "micromark-util-chunked": "^2.0.0", "micromark-util-classify-character": "^2.0.0", "micromark-util-resolve-all": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-ADVjpOOkjz1hhkZLlBiYA9cR2Anf8F4HqZUO6e5eDcPQd0Txw5fxLzzxnEkSkfnD0wziSGiv7sYhk/ktvbf1uw=="],
"micromark-extension-gfm-table": ["micromark-extension-gfm-table@2.1.1", "", { "dependencies": { "devlop": "^1.0.0", "micromark-factory-space": "^2.0.0", "micromark-util-character": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-t2OU/dXXioARrC6yWfJ4hqB7rct14e8f7m0cbI5hUmDyyIlwv5vEtooptH8INkbLzOatzKuVbQmAYcbWoyz6Dg=="],
"micromark-extension-gfm-tagfilter": ["micromark-extension-gfm-tagfilter@2.0.0", "", { "dependencies": { "micromark-util-types": "^2.0.0" } }, "sha512-xHlTOmuCSotIA8TW1mDIM6X2O1SiX5P9IuDtqGonFhEK0qgRI4yeC6vMxEV2dgyr2TiD+2PQ10o+cOhdVAcwfg=="],
"micromark-extension-gfm-task-list-item": ["micromark-extension-gfm-task-list-item@2.1.0", "", { "dependencies": { "devlop": "^1.0.0", "micromark-factory-space": "^2.0.0", "micromark-util-character": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-qIBZhqxqI6fjLDYFTBIa4eivDMnP+OZqsNwmQ3xNLE4Cxwc+zfQEfbs6tzAo2Hjq+bh6q5F+Z8/cksrLFYWQQw=="],
"micromark-factory-destination": ["micromark-factory-destination@2.0.1", "", { "dependencies": { "micromark-util-character": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-Xe6rDdJlkmbFRExpTOmRj9N3MaWmbAgdpSrBQvCFqhezUn4AHqJHbaEnfbVYYiexVSs//tqOdY/DxhjdCiJnIA=="],
"micromark-factory-label": ["micromark-factory-label@2.0.1", "", { "dependencies": { "devlop": "^1.0.0", "micromark-util-character": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-VFMekyQExqIW7xIChcXn4ok29YE3rnuyveW3wZQWWqF4Nv9Wk5rgJ99KzPvHjkmPXF93FXIbBp6YdW3t71/7Vg=="],
"micromark-factory-space": ["micromark-factory-space@2.0.1", "", { "dependencies": { "micromark-util-character": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-zRkxjtBxxLd2Sc0d+fbnEunsTj46SWXgXciZmHq0kDYGnck/ZSGj9/wULTV95uoeYiK5hRXP2mJ98Uo4cq/LQg=="],
"micromark-factory-title": ["micromark-factory-title@2.0.1", "", { "dependencies": { "micromark-factory-space": "^2.0.0", "micromark-util-character": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-5bZ+3CjhAd9eChYTHsjy6TGxpOFSKgKKJPJxr293jTbfry2KDoWkhBb6TcPVB4NmzaPhMs1Frm9AZH7OD4Cjzw=="],
"micromark-factory-whitespace": ["micromark-factory-whitespace@2.0.1", "", { "dependencies": { "micromark-factory-space": "^2.0.0", "micromark-util-character": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-Ob0nuZ3PKt/n0hORHyvoD9uZhr+Za8sFoP+OnMcnWK5lngSzALgQYKMr9RJVOWLqQYuyn6ulqGWSXdwf6F80lQ=="],
"micromark-util-character": ["micromark-util-character@2.1.1", "", { "dependencies": { "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-wv8tdUTJ3thSFFFJKtpYKOYiGP2+v96Hvk4Tu8KpCAsTMs6yi+nVmGh1syvSCsaxz45J6Jbw+9DD6g97+NV67Q=="],
"micromark-util-chunked": ["micromark-util-chunked@2.0.1", "", { "dependencies": { "micromark-util-symbol": "^2.0.0" } }, "sha512-QUNFEOPELfmvv+4xiNg2sRYeS/P84pTW0TCgP5zc9FpXetHY0ab7SxKyAQCNCc1eK0459uoLI1y5oO5Vc1dbhA=="],
"micromark-util-classify-character": ["micromark-util-classify-character@2.0.1", "", { "dependencies": { "micromark-util-character": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-K0kHzM6afW/MbeWYWLjoHQv1sgg2Q9EccHEDzSkxiP/EaagNzCm7T/WMKZ3rjMbvIpvBiZgwR3dKMygtA4mG1Q=="],
"micromark-util-combine-extensions": ["micromark-util-combine-extensions@2.0.1", "", { "dependencies": { "micromark-util-chunked": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-OnAnH8Ujmy59JcyZw8JSbK9cGpdVY44NKgSM7E9Eh7DiLS2E9RNQf0dONaGDzEG9yjEl5hcqeIsj4hfRkLH/Bg=="],
"micromark-util-decode-numeric-character-reference": ["micromark-util-decode-numeric-character-reference@2.0.2", "", { "dependencies": { "micromark-util-symbol": "^2.0.0" } }, "sha512-ccUbYk6CwVdkmCQMyr64dXz42EfHGkPQlBj5p7YVGzq8I7CtjXZJrubAYezf7Rp+bjPseiROqe7G6foFd+lEuw=="],
"micromark-util-decode-string": ["micromark-util-decode-string@2.0.1", "", { "dependencies": { "decode-named-character-reference": "^1.0.0", "micromark-util-character": "^2.0.0", "micromark-util-decode-numeric-character-reference": "^2.0.0", "micromark-util-symbol": "^2.0.0" } }, "sha512-nDV/77Fj6eH1ynwscYTOsbK7rR//Uj0bZXBwJZRfaLEJ1iGBR6kIfNmlNqaqJf649EP0F3NWNdeJi03elllNUQ=="],
"micromark-util-encode": ["micromark-util-encode@2.0.1", "", {}, "sha512-c3cVx2y4KqUnwopcO9b/SCdo2O67LwJJ/UyqGfbigahfegL9myoEFoDYZgkT7f36T0bLrM9hZTAaAyH+PCAXjw=="],
"micromark-util-html-tag-name": ["micromark-util-html-tag-name@2.0.1", "", {}, "sha512-2cNEiYDhCWKI+Gs9T0Tiysk136SnR13hhO8yW6BGNyhOC4qYFnwF1nKfD3HFAIXA5c45RrIG1ub11GiXeYd1xA=="],
"micromark-util-normalize-identifier": ["micromark-util-normalize-identifier@2.0.1", "", { "dependencies": { "micromark-util-symbol": "^2.0.0" } }, "sha512-sxPqmo70LyARJs0w2UclACPUUEqltCkJ6PhKdMIDuJ3gSf/Q+/GIe3WKl0Ijb/GyH9lOpUkRAO2wp0GVkLvS9Q=="],
"micromark-util-resolve-all": ["micromark-util-resolve-all@2.0.1", "", { "dependencies": { "micromark-util-types": "^2.0.0" } }, "sha512-VdQyxFWFT2/FGJgwQnJYbe1jjQoNTS4RjglmSjTUlpUMa95Htx9NHeYW4rGDJzbjvCsl9eLjMQwGeElsqmzcHg=="],
"micromark-util-sanitize-uri": ["micromark-util-sanitize-uri@2.0.1", "", { "dependencies": { "micromark-util-character": "^2.0.0", "micromark-util-encode": "^2.0.0", "micromark-util-symbol": "^2.0.0" } }, "sha512-9N9IomZ/YuGGZZmQec1MbgxtlgougxTodVwDzzEouPKo3qFWvymFHWcnDi2vzV1ff6kas9ucW+o3yzJK9YB1AQ=="],
"micromark-util-subtokenize": ["micromark-util-subtokenize@2.1.0", "", { "dependencies": { "devlop": "^1.0.0", "micromark-util-chunked": "^2.0.0", "micromark-util-symbol": "^2.0.0", "micromark-util-types": "^2.0.0" } }, "sha512-XQLu552iSctvnEcgXw6+Sx75GflAPNED1qx7eBJ+wydBb2KCbRZe+NwvIEEMM83uml1+2WSXpBAcp9IUCgCYWA=="],
"micromark-util-symbol": ["micromark-util-symbol@2.0.1", "", {}, "sha512-vs5t8Apaud9N28kgCrRUdEed4UJ+wWNvicHLPxCa9ENlYuAY31M0ETy5y1vA33YoNPDFTghEbnh6efaE8h4x0Q=="],
"micromark-util-types": ["micromark-util-types@2.0.2", "", {}, "sha512-Yw0ECSpJoViF1qTU4DC6NwtC4aWGt1EkzaQB8KPPyCRR8z9TWeV0HbEFGTO+ZY1wB22zmxnJqhPyTpOVCpeHTA=="],
"micromatch": ["micromatch@4.0.8", "", { "dependencies": { "braces": "^3.0.3", "picomatch": "^2.3.1" } }, "sha512-PXwfBhYu0hBCPw8Dn0E+WDYb7af3dSLVWKi3HGv84IdF4TyFoC0ysxFd0Goxw7nSv4T/PzEJQxsYsEiFCKo2BA=="],
"mri": ["mri@1.2.0", "", {}, "sha512-tzzskb3bG8LvYGFF/mDTpq3jpI6Q9wc3LEmBaghu+DdCssd1FakN7Bc0hVNmEyGq1bq3RgfkCb3cmQLpNPOroA=="],
"ms": ["ms@2.1.3", "", {}, "sha512-6FlzubTLZG3J2a/NVCAleEhjzq5oxgHyaCU9yYXvcLsvoVaHJq/s5xXI6/XXP6tz7R9xAOtHnSO/tXtF3WRTlA=="], "ms": ["ms@2.1.3", "", {}, "sha512-6FlzubTLZG3J2a/NVCAleEhjzq5oxgHyaCU9yYXvcLsvoVaHJq/s5xXI6/XXP6tz7R9xAOtHnSO/tXtF3WRTlA=="],
@@ -135,39 +339,123 @@
"nwsapi": ["nwsapi@2.2.23", "", {}, "sha512-7wfH4sLbt4M0gCDzGE6vzQBo0bfTKjU7Sfpqy/7gs1qBfYz2vEJH6vXcBKpO3+6Yu1telwd0t9HpyOoLEQQbIQ=="], "nwsapi": ["nwsapi@2.2.23", "", {}, "sha512-7wfH4sLbt4M0gCDzGE6vzQBo0bfTKjU7Sfpqy/7gs1qBfYz2vEJH6vXcBKpO3+6Yu1telwd0t9HpyOoLEQQbIQ=="],
"outdent": ["outdent@0.5.0", "", {}, "sha512-/jHxFIzoMXdqPzTaCpFzAAWhpkSjZPF4Vsn6jAfNpmbH/ymsmd7Qc6VE9BGn0L6YMj6uwpQLxCECpus4ukKS9Q=="],
"p-filter": ["p-filter@2.1.0", "", { "dependencies": { "p-map": "^2.0.0" } }, "sha512-ZBxxZ5sL2HghephhpGAQdoskxplTwr7ICaehZwLIlfL6acuVgZPm8yBNuRAFBGEqtD/hmUeq9eqLg2ys9Xr/yw=="],
"p-limit": ["p-limit@2.3.0", "", { "dependencies": { "p-try": "^2.0.0" } }, "sha512-//88mFWSJx8lxCzwdAABTJL2MyWB12+eIY7MDL2SqLmAkeKU9qxRvWuSyTjm3FUmpBEMuFfckAIqEaVGUDxb6w=="],
"p-locate": ["p-locate@4.1.0", "", { "dependencies": { "p-limit": "^2.2.0" } }, "sha512-R79ZZ/0wAxKGu3oYMlz8jy/kbhsNrS7SKZ7PxEHBgJ5+F2mtFW2fK2cOtBh1cHYkQsbzFV7I+EoRKe6Yt0oK7A=="],
"p-map": ["p-map@2.1.0", "", {}, "sha512-y3b8Kpd8OAN444hxfBbFfj1FY/RjtTd8tzYwhUqNYXx0fXx2iX4maP4Qr6qhIKbQXI02wTLAda4fYUbDagTUFw=="],
"p-try": ["p-try@2.2.0", "", {}, "sha512-R4nPAVTAU0B9D35/Gk3uJf/7XYbQcyohSKdvAxIRSNghFl4e71hVoGnBNQz9cWaXxO2I10KTC+3jMdvvoKw6dQ=="],
"package-manager-detector": ["package-manager-detector@0.2.11", "", { "dependencies": { "quansync": "^0.2.7" } }, "sha512-BEnLolu+yuz22S56CU1SUKq3XC3PkwD5wv4ikR4MfGvnRVcmzXR9DwSlW2fEamyTPyXHomBJRzgapeuBvRNzJQ=="],
"parse5": ["parse5@7.3.0", "", { "dependencies": { "entities": "^6.0.0" } }, "sha512-IInvU7fabl34qmi9gY8XOVxhYyMyuH2xUNpb2q8/Y+7552KlejkRvqvD19nMoUW/uQGGbqNpA6Tufu5FL5BZgw=="], "parse5": ["parse5@7.3.0", "", { "dependencies": { "entities": "^6.0.0" } }, "sha512-IInvU7fabl34qmi9gY8XOVxhYyMyuH2xUNpb2q8/Y+7552KlejkRvqvD19nMoUW/uQGGbqNpA6Tufu5FL5BZgw=="],
"psl": ["psl@1.15.0", "", { "dependencies": { "punycode": "^2.3.1" } }, "sha512-JZd3gMVBAVQkSs6HdNZo9Sdo0LNcQeMNP3CozBJb3JYC/QUYZTnKxP+f8oWRX4rHP5EurWxqAHTSwUCjlNKa1w=="], "path-exists": ["path-exists@4.0.0", "", {}, "sha512-ak9Qy5Q7jYb2Wwcey5Fpvg2KoAc/ZIhLSLOSBmRmygPsGwkVVt0fZa0qrtMz+m6tJTAHfZQ8FnmB4MG4LWy7/w=="],
"path-key": ["path-key@3.1.1", "", {}, "sha512-ojmeN0qd+y0jszEtoY48r0Peq5dwMEkIlCOu6Q5f41lfkswXuKtYrhgoTpLnyIcHm24Uhqx+5Tqm2InSwLhE6Q=="],
"path-type": ["path-type@4.0.0", "", {}, "sha512-gDKb8aZMDeD/tZWs9P6+q0J9Mwkdl6xMV8TjnGP3qJVJ06bdMgkbBlLU8IdfOsIsFz2BW1rNVT3XuNEl8zPAvw=="],
"picocolors": ["picocolors@1.1.1", "", {}, "sha512-xceH2snhtb5M9liqDsmEw56le376mTZkEX/jEb/RxNFyegNul7eNslCXP9FDj/Lcu0X8KEyMceP2ntpaHrDEVA=="],
"picomatch": ["picomatch@2.3.2", "", {}, "sha512-V7+vQEJ06Z+c5tSye8S+nHUfI51xoXIXjHQ99cQtKUkQqqO1kO/KCJUfZXuB47h/YBlDhah2H3hdUGXn8ie0oA=="],
"pify": ["pify@4.0.1", "", {}, "sha512-uB80kBFb/tfd68bVleG9T5GGsGPjJrLAUpR5PZIrhBnIaRTQRjqdJSsIKkOP6OAIFbj7GOrcudc5pNjZ+geV2g=="],
"prettier": ["prettier@2.8.8", "", { "bin": { "prettier": "bin-prettier.js" } }, "sha512-tdN8qQGvNjw4CHbY+XXk0JgCXn9QiF21a55rBe5LJAU+kDyC4WQn4+awm2Xfk2lQMk5fKup9XgzTZtGkjBdP9Q=="],
"punycode": ["punycode@2.3.1", "", {}, "sha512-vYt7UD1U9Wg6138shLtLOvdAu+8DsC/ilFtEVHcH+wydcSpNE20AfSOduf6MkRFahL5FY7X1oU7nKVZFtfq8Fg=="], "punycode": ["punycode@2.3.1", "", {}, "sha512-vYt7UD1U9Wg6138shLtLOvdAu+8DsC/ilFtEVHcH+wydcSpNE20AfSOduf6MkRFahL5FY7X1oU7nKVZFtfq8Fg=="],
"querystringify": ["querystringify@2.2.0", "", {}, "sha512-FIqgj2EUvTa7R50u0rGsyTftzjYmv/a3hO345bZNrqabNqjtgiDMgmo4mkUjd+nzU5oF3dClKqFIPUKybUyqoQ=="], "quansync": ["quansync@0.2.11", "", {}, "sha512-AifT7QEbW9Nri4tAwR5M/uzpBuqfZf+zwaEM/QkzEjj7NBuFD2rBuy0K3dE+8wltbezDV7JMA0WfnCPYRSYbXA=="],
"requires-port": ["requires-port@1.0.0", "", {}, "sha512-KigOCHcocU3XODJxsu8i/j8T9tzT4adHiecwORRQ0ZZFcp7ahwXuRU1m+yuO90C5ZUyGeGfocHDI14M3L3yDAQ=="], "queue-microtask": ["queue-microtask@1.2.3", "", {}, "sha512-NuaNSa6flKT5JaSYQzJok04JzTL1CA6aGhv5rfLW3PgqA+M2ChpZQnAC8h8i4ZFkBS8X5RqkDBHA7r4hej3K9A=="],
"rrweb-cssom": ["rrweb-cssom@0.7.1", "", {}, "sha512-TrEMa7JGdVm0UThDJSx7ddw5nVm3UJS9o9CCIZ72B1vSyEZoziDqBYP3XIoi/12lKrJR8rE3jeFHMok2F/Mnsg=="], "read-yaml-file": ["read-yaml-file@1.1.0", "", { "dependencies": { "graceful-fs": "^4.1.5", "js-yaml": "^3.6.1", "pify": "^4.0.1", "strip-bom": "^3.0.0" } }, "sha512-VIMnQi/Z4HT2Fxuwg5KrY174U1VdUIASQVWXXyqtNRtxSr9IYkn1rsI6Tb6HsrHCmB7gVpNwX6JxPTHcH6IoTA=="],
"remark-gfm": ["remark-gfm@4.0.1", "", { "dependencies": { "@types/mdast": "^4.0.0", "mdast-util-gfm": "^3.0.0", "micromark-extension-gfm": "^3.0.0", "remark-parse": "^11.0.0", "remark-stringify": "^11.0.0", "unified": "^11.0.0" } }, "sha512-1quofZ2RQ9EWdeN34S79+KExV1764+wCUGop5CPL1WGdD0ocPpu91lzPGbwWMECpEpd42kJGQwzRfyov9j4yNg=="],
"remark-parse": ["remark-parse@11.0.0", "", { "dependencies": { "@types/mdast": "^4.0.0", "mdast-util-from-markdown": "^2.0.0", "micromark-util-types": "^2.0.0", "unified": "^11.0.0" } }, "sha512-FCxlKLNGknS5ba/1lmpYijMUzX2esxW5xQqjWxw2eHFfS2MSdaHVINFmhjo+qN1WhZhNimq0dZATN9pH0IDrpA=="],
"remark-stringify": ["remark-stringify@11.0.0", "", { "dependencies": { "@types/mdast": "^4.0.0", "mdast-util-to-markdown": "^2.0.0", "unified": "^11.0.0" } }, "sha512-1OSmLd3awB/t8qdoEOMazZkNsfVTeY4fTsgzcQFdXNq8ToTN4ZGwrMnlda4K6smTFKD+GRV6O48i6Z4iKgPPpw=="],
"resolve-from": ["resolve-from@5.0.0", "", {}, "sha512-qYg9KP24dD5qka9J47d0aVky0N+b4fTU89LN9iDnjB5waksiC49rvMB0PrUJQGoTmH50XPiqOvAjDfaijGxYZw=="],
"reusify": ["reusify@1.1.0", "", {}, "sha512-g6QUff04oZpHs0eG5p83rFLhHeV00ug/Yf9nZM6fLeUrPguBTkTQOdpAWWspMh55TZfVQDPaN3NQJfbVRAxdIw=="],
"rrweb-cssom": ["rrweb-cssom@0.8.0", "", {}, "sha512-guoltQEx+9aMf2gDZ0s62EcV8lsXR+0w8915TC3ITdn2YueuNjdAYh/levpU9nFaoChh9RUS5ZdQMrKfVEN9tw=="],
"run-parallel": ["run-parallel@1.2.0", "", { "dependencies": { "queue-microtask": "^1.2.2" } }, "sha512-5l4VyZR86LZ/lDxZTR6jqL8AFE2S0IFLMP26AbjsLVADxHdhB/c0GUsH+y39UfCi3dzz8OlQuPmnaJOMoDHQBA=="],
"safer-buffer": ["safer-buffer@2.1.2", "", {}, "sha512-YZo3K82SD7Riyi0E1EQPojLz7kpepnSQI9IyPbHHg1XXXevb5dJI7tpyN2ADxGcQbHG7vcyRHk0cbwqcQriUtg=="], "safer-buffer": ["safer-buffer@2.1.2", "", {}, "sha512-YZo3K82SD7Riyi0E1EQPojLz7kpepnSQI9IyPbHHg1XXXevb5dJI7tpyN2ADxGcQbHG7vcyRHk0cbwqcQriUtg=="],
"saxes": ["saxes@6.0.0", "", { "dependencies": { "xmlchars": "^2.2.0" } }, "sha512-xAg7SOnEhrm5zI3puOOKyy1OMcMlIJZYNJY7xLBwSze0UjhPLnWfj2GF2EpT0jmzaJKIWKHLsaSSajf35bcYnA=="], "saxes": ["saxes@6.0.0", "", { "dependencies": { "xmlchars": "^2.2.0" } }, "sha512-xAg7SOnEhrm5zI3puOOKyy1OMcMlIJZYNJY7xLBwSze0UjhPLnWfj2GF2EpT0jmzaJKIWKHLsaSSajf35bcYnA=="],
"semver": ["semver@7.7.4", "", { "bin": { "semver": "bin/semver.js" } }, "sha512-vFKC2IEtQnVhpT78h1Yp8wzwrf8CM+MzKMHGJZfBtzhZNycRFnXsHk6E5TxIkkMsgNS7mdX3AGB7x2QM2di4lA=="],
"shebang-command": ["shebang-command@2.0.0", "", { "dependencies": { "shebang-regex": "^3.0.0" } }, "sha512-kHxr2zZpYtdmrN1qDjrrX/Z1rR1kG8Dx+gkpK1G4eXmvXswmcE1hTWBWYUzlraYw1/yZp6YuDY77YtvbN0dmDA=="],
"shebang-regex": ["shebang-regex@3.0.0", "", {}, "sha512-7++dFhtcx3353uBaq8DDR4NuxBetBzC7ZQOhmTQInHEd6bSrXdiEyzCvG07Z44UYdLShWUyXt5M/yhz8ekcb1A=="],
"signal-exit": ["signal-exit@4.1.0", "", {}, "sha512-bzyZ1e88w9O1iNJbKnOlvYTrWPDl46O1bG0D3XInv+9tkPrxrN8jUUTiFlDkkmKWgn1M6CfIA13SuGqOa9Korw=="],
"slash": ["slash@3.0.0", "", {}, "sha512-g9Q1haeby36OSStwb4ntCGGGaKsaVSjQ68fBxoQcutl5fS1vuY18H3wSt3jFyFtrkx+Kz0V1G85A4MyAdDMi2Q=="],
"spawndamnit": ["spawndamnit@3.0.1", "", { "dependencies": { "cross-spawn": "^7.0.5", "signal-exit": "^4.0.1" } }, "sha512-MmnduQUuHCoFckZoWnXsTg7JaiLBJrKFj9UI2MbRPGaJeVpsLcVBu6P/IGZovziM/YBsellCmsprgNA+w0CzVg=="],
"sprintf-js": ["sprintf-js@1.0.3", "", {}, "sha512-D9cPgkvLlV3t3IzL0D0YLvGA9Ahk4PcvVwUbN0dSGr1aP0Nrt4AEnTUbuGvquEC0mA64Gqt1fzirlRs5ibXx8g=="],
"strip-ansi": ["strip-ansi@6.0.1", "", { "dependencies": { "ansi-regex": "^5.0.1" } }, "sha512-Y38VPSHcqkFrCpFnQ9vuSXmquuv5oXOKpGeT6aGrr3o3Gc9AlVa6JBfUSOCnbxGGZF+/0ooI7KrPuUSztUdU5A=="],
"strip-bom": ["strip-bom@3.0.0", "", {}, "sha512-vavAMRXOgBVNF6nyEEmL3DBK19iRpDcoIwW+swQ+CbGiu7lju6t+JklA1MHweoWtadgt4ISVUsXLyDq34ddcwA=="],
"symbol-tree": ["symbol-tree@3.2.4", "", {}, "sha512-9QNk5KwDF+Bvz+PyObkmSYjI5ksVUYtjW7AU22r2NKcfLJcXp96hkDWU3+XndOsUb+AQ9QhfzfCT2O+CNWT5Tw=="], "symbol-tree": ["symbol-tree@3.2.4", "", {}, "sha512-9QNk5KwDF+Bvz+PyObkmSYjI5ksVUYtjW7AU22r2NKcfLJcXp96hkDWU3+XndOsUb+AQ9QhfzfCT2O+CNWT5Tw=="],
"temml": ["temml@0.13.1", "", {}, "sha512-/fL1utq8QUD9YpcLeZHPRnp9Cbzbexq5hZl5uSBhf8mNYiKkcS4eYbLidDB+/nF8C+RHAcBQbKw2bKoS83mz1Q=="], "temml": ["temml@0.13.2", "", {}, "sha512-n8fDRSsLscq9nh9j6z+FgkCvFMT0IJm6GCgwfzh+7AHT3Sfb4jFTQlsA6hVcF2dYYr3b66oDBVES95RfoukyrA=="],
"tough-cookie": ["tough-cookie@4.1.4", "", { "dependencies": { "psl": "^1.1.33", "punycode": "^2.1.1", "universalify": "^0.2.0", "url-parse": "^1.5.3" } }, "sha512-Loo5UUvLD9ScZ6jh8beX1T6sO1w2/MpCRpEP7V280GKMVUQ0Jzar2U3UJPsrdbziLEMMhu3Ujnq//rhiFuIeag=="], "term-size": ["term-size@2.2.1", "", {}, "sha512-wK0Ri4fOGjv/XPy8SBHZChl8CM7uMc5VML7SqiQ0zG7+J5Vr+RMQDoHa2CNT6KHUnTGIXH34UDMkPzAUyapBZg=="],
"tldts": ["tldts@6.1.86", "", { "dependencies": { "tldts-core": "^6.1.86" }, "bin": { "tldts": "bin/cli.js" } }, "sha512-WMi/OQ2axVTf/ykqCQgXiIct+mSQDFdH2fkwhPwgEwvJ1kSzZRiinb0zF2Xb8u4+OqPChmyI6MEu4EezNJz+FQ=="],
"tldts-core": ["tldts-core@6.1.86", "", {}, "sha512-Je6p7pkk+KMzMv2XXKmAE3McmolOQFdxkKw0R8EYNr7sELW46JqnNeTX8ybPiQgvg1ymCoF8LXs5fzFaZvJPTA=="],
"to-regex-range": ["to-regex-range@5.0.1", "", { "dependencies": { "is-number": "^7.0.0" } }, "sha512-65P7iz6X5yEr1cwcgvQxbbIw7Uk3gOy5dIdtZ4rDveLqhrdJP+Li/Hx6tyK0NEb+2GCyneCMJiGqrADCSNk8sQ=="],
"tough-cookie": ["tough-cookie@5.1.2", "", { "dependencies": { "tldts": "^6.1.32" } }, "sha512-FVDYdxtnj0G6Qm/DhNPSb8Ju59ULcup3tuJxkFb5K8Bv2pUXILbf0xZWU8PX8Ov19OXljbUyveOFwRMwkXzO+A=="],
"tr46": ["tr46@5.1.1", "", { "dependencies": { "punycode": "^2.3.1" } }, "sha512-hdF5ZgjTqgAntKkklYw0R03MG2x/bSzTtkxmIRw/sTNV8YXsCJ1tfLAX23lhxhHJlEf3CRCOCGGWw3vI3GaSPw=="], "tr46": ["tr46@5.1.1", "", { "dependencies": { "punycode": "^2.3.1" } }, "sha512-hdF5ZgjTqgAntKkklYw0R03MG2x/bSzTtkxmIRw/sTNV8YXsCJ1tfLAX23lhxhHJlEf3CRCOCGGWw3vI3GaSPw=="],
"trough": ["trough@2.2.0", "", {}, "sha512-tmMpK00BjZiUyVyvrBK7knerNgmgvcV/KLVyuma/SC+TQN167GrMRciANTz09+k3zW8L8t60jWO1GpfkZdjTaw=="],
"turndown": ["turndown@7.2.2", "", { "dependencies": { "@mixmark-io/domino": "^2.2.0" } }, "sha512-1F7db8BiExOKxjSMU2b7if62D/XOyQyZbPKq/nUwopfgnHlqXHqQ0lvfUTeUIr1lZJzOPFn43dODyMSIfvWRKQ=="], "turndown": ["turndown@7.2.2", "", { "dependencies": { "@mixmark-io/domino": "^2.2.0" } }, "sha512-1F7db8BiExOKxjSMU2b7if62D/XOyQyZbPKq/nUwopfgnHlqXHqQ0lvfUTeUIr1lZJzOPFn43dODyMSIfvWRKQ=="],
"turndown-plugin-gfm": ["turndown-plugin-gfm@1.0.2", "", {}, "sha512-vwz9tfvF7XN/jE0dGoBei3FXWuvll78ohzCZQuOb+ZjWrs3a0XhQVomJEb2Qh4VHTPNRO4GPZh0V7VRbiWwkRg=="], "turndown-plugin-gfm": ["turndown-plugin-gfm@1.0.2", "", {}, "sha512-vwz9tfvF7XN/jE0dGoBei3FXWuvll78ohzCZQuOb+ZjWrs3a0XhQVomJEb2Qh4VHTPNRO4GPZh0V7VRbiWwkRg=="],
"typescript": ["typescript@5.9.3", "", { "bin": { "tsc": "bin/tsc", "tsserver": "bin/tsserver" } }, "sha512-jl1vZzPDinLr9eUt3J/t7V6FgNEw9QjvBPdysz9KfQDD41fQrC2Y4vKQdiaUpFT4bXlb1RHhLpp8wtm6M5TgSw=="],
"uhyphen": ["uhyphen@0.2.0", "", {}, "sha512-qz3o9CHXmJJPGBdqzab7qAYuW8kQGKNEuoHFYrBwV6hWIMcpAmxDLXojcHfFr9US1Pe6zUswEIJIbLI610fuqA=="], "uhyphen": ["uhyphen@0.2.0", "", {}, "sha512-qz3o9CHXmJJPGBdqzab7qAYuW8kQGKNEuoHFYrBwV6hWIMcpAmxDLXojcHfFr9US1Pe6zUswEIJIbLI610fuqA=="],
"universalify": ["universalify@0.2.0", "", {}, "sha512-CJ1QgKmNg3CwvAv/kOFmtnEN05f0D/cn9QntgNOQlQF9dgvVTHj3t+8JPdjqawCHk7V/KA+fbUqzZ9XWhcqPUg=="], "undici-types": ["undici-types@7.18.2", "", {}, "sha512-AsuCzffGHJybSaRrmr5eHr81mwJU3kjw6M+uprWvCXiNeN9SOGwQ3Jn8jb8m3Z6izVgknn1R0FTCEAP2QrLY/w=="],
"url-parse": ["url-parse@1.5.10", "", { "dependencies": { "querystringify": "^2.1.1", "requires-port": "^1.0.0" } }, "sha512-WypcfiRhfeUP9vvF0j6rw0J3hrWrw6iZv3+22h6iRMJ/8z1Tj6XfLP4DsUix5MhMPnXpiHDoKyoZ/bdCkwBCiQ=="], "unified": ["unified@11.0.5", "", { "dependencies": { "@types/unist": "^3.0.0", "bail": "^2.0.0", "devlop": "^1.0.0", "extend": "^3.0.0", "is-plain-obj": "^4.0.0", "trough": "^2.0.0", "vfile": "^6.0.0" } }, "sha512-xKvGhPWw3k84Qjh8bI3ZeJjqnyadK+GEFtazSfZv/rKeTkTjOJho6mFqh2SM96iIcZokxiOpg78GazTSg8+KHA=="],
"unist-util-is": ["unist-util-is@6.0.1", "", { "dependencies": { "@types/unist": "^3.0.0" } }, "sha512-LsiILbtBETkDz8I9p1dQ0uyRUWuaQzd/cuEeS1hoRSyW5E5XGmTzlwY1OrNzzakGowI9Dr/I8HVaw4hTtnxy8g=="],
"unist-util-stringify-position": ["unist-util-stringify-position@4.0.0", "", { "dependencies": { "@types/unist": "^3.0.0" } }, "sha512-0ASV06AAoKCDkS2+xw5RXJywruurpbC4JZSm7nr7MOt1ojAzvyyaO+UxZf18j8FCF6kmzCZKcAgN/yu2gm2XgQ=="],
"unist-util-visit": ["unist-util-visit@5.1.0", "", { "dependencies": { "@types/unist": "^3.0.0", "unist-util-is": "^6.0.0", "unist-util-visit-parents": "^6.0.0" } }, "sha512-m+vIdyeCOpdr/QeQCu2EzxX/ohgS8KbnPDgFni4dQsfSCtpz8UqDyY5GjRru8PDKuYn7Fq19j1CQ+nJSsGKOzg=="],
"unist-util-visit-parents": ["unist-util-visit-parents@6.0.2", "", { "dependencies": { "@types/unist": "^3.0.0", "unist-util-is": "^6.0.0" } }, "sha512-goh1s1TBrqSqukSc8wrjwWhL0hiJxgA8m4kFxGlQ+8FYQ3C/m11FcTs4YYem7V664AhHVvgoQLk890Ssdsr2IQ=="],
"universalify": ["universalify@0.1.2", "", {}, "sha512-rBJeI5CXAlmy1pV+617WB9J63U6XcazHHF2f2dbJix4XzpUF0RS3Zbj0FGIOCAva5P/d/GBOYaACQ1w+0azUkg=="],
"vfile": ["vfile@6.0.3", "", { "dependencies": { "@types/unist": "^3.0.0", "vfile-message": "^4.0.0" } }, "sha512-KzIbH/9tXat2u30jf+smMwFCsno4wHVdNmzFyL+T/L3UGqqk6JKfVqOFOZEpZSHADH1k40ab6NUIXZq422ov3Q=="],
"vfile-message": ["vfile-message@4.0.3", "", { "dependencies": { "@types/unist": "^3.0.0", "unist-util-stringify-position": "^4.0.0" } }, "sha512-QTHzsGd1EhbZs4AsQ20JX1rC3cOlt/IWJruk893DfLRr57lcnOeMaWG4K0JrRta4mIJZKth2Au3mM3u03/JWKw=="],
"w3c-xmlserializer": ["w3c-xmlserializer@5.0.0", "", { "dependencies": { "xml-name-validator": "^5.0.0" } }, "sha512-o8qghlI8NZHU1lLPrpi2+Uq7abh4GGPpYANlalzWxyWteJOCsr/P+oPBA49TOLu5FTZO4d3F9MnWJfiMo4BkmA=="], "w3c-xmlserializer": ["w3c-xmlserializer@5.0.0", "", { "dependencies": { "xml-name-validator": "^5.0.0" } }, "sha512-o8qghlI8NZHU1lLPrpi2+Uq7abh4GGPpYANlalzWxyWteJOCsr/P+oPBA49TOLu5FTZO4d3F9MnWJfiMo4BkmA=="],
@@ -179,16 +467,34 @@
"whatwg-url": ["whatwg-url@14.2.0", "", { "dependencies": { "tr46": "^5.1.0", "webidl-conversions": "^7.0.0" } }, "sha512-De72GdQZzNTUBBChsXueQUnPKDkg/5A5zp7pFDuQAj5UFoENpiACU0wlCvzpAGnTkj++ihpKwKyYewn/XNUbKw=="], "whatwg-url": ["whatwg-url@14.2.0", "", { "dependencies": { "tr46": "^5.1.0", "webidl-conversions": "^7.0.0" } }, "sha512-De72GdQZzNTUBBChsXueQUnPKDkg/5A5zp7pFDuQAj5UFoENpiACU0wlCvzpAGnTkj++ihpKwKyYewn/XNUbKw=="],
"ws": ["ws@8.19.0", "", { "peerDependencies": { "bufferutil": "^4.0.1", "utf-8-validate": ">=5.0.2" }, "optionalPeers": ["bufferutil", "utf-8-validate"] }, "sha512-blAT2mjOEIi0ZzruJfIhb3nps74PRWTCz1IjglWEEpQl5XS/UNama6u2/rjFkDDouqr4L67ry+1aGIALViWjDg=="], "which": ["which@2.0.2", "", { "dependencies": { "isexe": "^2.0.0" }, "bin": { "node-which": "./bin/node-which" } }, "sha512-BLI3Tl1TW3Pvl70l3yq3Y64i+awpwXqsGBYWkkqMtnbXgrMD+yj7rhW0kuEDxzJaYXGjEW5ogapKNMEKNMjibA=="],
"ws": ["ws@8.20.0", "", { "peerDependencies": { "bufferutil": "^4.0.1", "utf-8-validate": ">=5.0.2" }, "optionalPeers": ["bufferutil", "utf-8-validate"] }, "sha512-sAt8BhgNbzCtgGbt2OxmpuryO63ZoDk/sqaB/znQm94T4fCEsy/yV+7CdC1kJhOU9lboAEU7R3kquuycDoibVA=="],
"xml-name-validator": ["xml-name-validator@5.0.0", "", {}, "sha512-EvGK8EJ3DhaHfbRlETOWAS5pO9MZITeauHKJyb8wyajUfQUenkIg2MvLDTZ4T/TgIcm3HU0TFBgWWboAZ30UHg=="], "xml-name-validator": ["xml-name-validator@5.0.0", "", {}, "sha512-EvGK8EJ3DhaHfbRlETOWAS5pO9MZITeauHKJyb8wyajUfQUenkIg2MvLDTZ4T/TgIcm3HU0TFBgWWboAZ30UHg=="],
"xmlchars": ["xmlchars@2.2.0", "", {}, "sha512-JZnDKK8B0RCDw84FNdDAIpZK+JuJw+s7Lz8nksI7SIuU3UXJJslUthsi+uWBUYOwPFwW7W7PRLRfUKpxjtjFCw=="], "xmlchars": ["xmlchars@2.2.0", "", {}, "sha512-JZnDKK8B0RCDw84FNdDAIpZK+JuJw+s7Lz8nksI7SIuU3UXJJslUthsi+uWBUYOwPFwW7W7PRLRfUKpxjtjFCw=="],
"cssstyle/rrweb-cssom": ["rrweb-cssom@0.8.0", "", {}, "sha512-guoltQEx+9aMf2gDZ0s62EcV8lsXR+0w8915TC3ITdn2YueuNjdAYh/levpU9nFaoChh9RUS5ZdQMrKfVEN9tw=="], "zwitch": ["zwitch@2.0.4", "", {}, "sha512-bXE4cR/kVZhKZX/RjPEflHaKVhUVl85noU3v6b8apfQEc1x4A+zBxjZ4lN8LqGd6WZ3dl98pY4o717VFmoPp+A=="],
"@manypkg/find-root/@types/node": ["@types/node@12.20.55", "", {}, "sha512-J8xLz7q2OFulZ2cyGTLE1TbbZcjpno7FaN6zdJNrgAdrJ+DZzh/uFR6YrTb4C+nXakvud8Q4+rbhoIWlYQbUFQ=="],
"@manypkg/find-root/fs-extra": ["fs-extra@8.1.0", "", { "dependencies": { "graceful-fs": "^4.2.0", "jsonfile": "^4.0.0", "universalify": "^0.1.0" } }, "sha512-yhlQgA6mnOJUKOsRUFsgJdQCvkKhcz8tlZG5HBQfReYZy46OwLcY+Zia0mtdHsOo9y/hP+CxMN0TU9QxoOtG4g=="],
"@manypkg/get-packages/@changesets/types": ["@changesets/types@4.1.0", "", {}, "sha512-LDQvVDv5Kb50ny2s25Fhm3d9QSZimsoUGBsUioj6MC3qbMUCuC8GPIvk/M6IvXx3lYhAs0lwWUQLb+VIEUCECw=="],
"@manypkg/get-packages/fs-extra": ["fs-extra@8.1.0", "", { "dependencies": { "graceful-fs": "^4.2.0", "jsonfile": "^4.0.0", "universalify": "^0.1.0" } }, "sha512-yhlQgA6mnOJUKOsRUFsgJdQCvkKhcz8tlZG5HBQfReYZy46OwLcY+Zia0mtdHsOo9y/hP+CxMN0TU9QxoOtG4g=="],
"dom-serializer/entities": ["entities@4.5.0", "", {}, "sha512-V0hjH4dGPh9Ao5p0MoRY6BVqtwCjhz6vI5LT8AJ55H+4g9/4vbHx1I54fS0XuclLhDHArPQCiMjDxjaL8fPxhw=="], "dom-serializer/entities": ["entities@4.5.0", "", {}, "sha512-V0hjH4dGPh9Ao5p0MoRY6BVqtwCjhz6vI5LT8AJ55H+4g9/4vbHx1I54fS0XuclLhDHArPQCiMjDxjaL8fPxhw=="],
"htmlparser2/entities": ["entities@7.0.1", "", {}, "sha512-TWrgLOFUQTH994YUyl1yT4uyavY5nNB5muff+RtWaqNVCAK408b5ZnnbNAUEWLTCpum9w6arT70i1XdQ4UeOPA=="], "htmlparser2/entities": ["entities@7.0.1", "", {}, "sha512-TWrgLOFUQTH994YUyl1yT4uyavY5nNB5muff+RtWaqNVCAK408b5ZnnbNAUEWLTCpum9w6arT70i1XdQ4UeOPA=="],
"mdast-util-find-and-replace/escape-string-regexp": ["escape-string-regexp@5.0.0", "", {}, "sha512-/veY75JbMK4j1yjvuUxuVsiS/hr/4iHs9FTT6cgTexxdE0Ly/glccBAkloH/DofkjRbZU3bnoj38mOmhkZ0lHw=="],
"read-yaml-file/js-yaml": ["js-yaml@3.14.2", "", { "dependencies": { "argparse": "^1.0.7", "esprima": "^4.0.0" }, "bin": { "js-yaml": "bin/js-yaml.js" } }, "sha512-PMSmkqxr106Xa156c2M265Z+FTrPl+oxd/rgOQy2tijQeK5TxQ43psO1ZCwhVOSdnn+RzkzlRz/eY4BgJBYVpg=="],
"whatwg-encoding/iconv-lite": ["iconv-lite@0.6.3", "", { "dependencies": { "safer-buffer": ">= 2.1.2 < 3.0.0" } }, "sha512-4fCk79wshMdzMp2rH06qWrJE4iolqLhCUH+OiuIgU++RB0+94NlDL81atO7GX55uUKueo0txHNtvEyI6D7WdMw=="],
"read-yaml-file/js-yaml/argparse": ["argparse@1.0.10", "", { "dependencies": { "sprintf-js": "~1.0.2" } }, "sha512-o5Roy6tNG4SL/FOkCAN6RzjiakZS25RLYFrcMttJqbdd8BWrnA+fGz57iN5Pb06pvBGvl5gQ0B48dJlslXvoTg=="],
} }
} }
-179
View File
@@ -1,179 +0,0 @@
import {
CdpConnection,
findChromeExecutable as findChromeExecutableBase,
findExistingChromeDebugPort,
getFreePort,
killChrome,
launchChrome as launchChromeBase,
sleep,
waitForChromeDebugPort,
type PlatformCandidates,
} from 'baoyu-chrome-cdp';
import { resolveUrlToMarkdownChromeProfileDir } from './paths.js';
import { NETWORK_IDLE_TIMEOUT_MS } from './constants.js';
const CHROME_CANDIDATES_FULL: PlatformCandidates = {
darwin: [
'/Applications/Google Chrome.app/Contents/MacOS/Google Chrome',
'/Applications/Google Chrome Canary.app/Contents/MacOS/Google Chrome Canary',
'/Applications/Google Chrome Beta.app/Contents/MacOS/Google Chrome Beta',
'/Applications/Chromium.app/Contents/MacOS/Chromium',
'/Applications/Microsoft Edge.app/Contents/MacOS/Microsoft Edge',
],
win32: [
'C:\\Program Files\\Google\\Chrome\\Application\\chrome.exe',
'C:\\Program Files (x86)\\Google\\Chrome\\Application\\chrome.exe',
'C:\\Program Files\\Microsoft\\Edge\\Application\\msedge.exe',
'C:\\Program Files (x86)\\Microsoft\\Edge\\Application\\msedge.exe',
],
default: [
'/usr/bin/google-chrome',
'/usr/bin/google-chrome-stable',
'/usr/bin/chromium',
'/usr/bin/chromium-browser',
'/snap/bin/chromium',
'/usr/bin/microsoft-edge',
],
};
export { CdpConnection, getFreePort, killChrome, sleep, waitForChromeDebugPort };
export async function findExistingChromePort(): Promise<number | null> {
return await findExistingChromeDebugPort({
profileDir: resolveUrlToMarkdownChromeProfileDir(),
});
}
export function findChromeExecutable(): string | null {
return findChromeExecutableBase({
candidates: CHROME_CANDIDATES_FULL,
envNames: ['URL_CHROME_PATH'],
}) ?? null;
}
export async function launchChrome(url: string, port: number, headless = false) {
const chromePath = findChromeExecutable();
if (!chromePath) throw new Error('Chrome executable not found. Install Chrome or set URL_CHROME_PATH env.');
return await launchChromeBase({
chromePath,
profileDir: resolveUrlToMarkdownChromeProfileDir(),
port,
url,
headless,
extraArgs: ['--disable-popup-blocking'],
});
}
export async function waitForNetworkIdle(
cdp: CdpConnection,
sessionId: string,
timeoutMs: number = NETWORK_IDLE_TIMEOUT_MS,
): Promise<void> {
return new Promise((resolve) => {
let timer: ReturnType<typeof setTimeout> | null = null;
let pending = 0;
const cleanup = () => {
if (timer) clearTimeout(timer);
cdp.off('Network.requestWillBeSent', onRequest);
cdp.off('Network.loadingFinished', onFinish);
cdp.off('Network.loadingFailed', onFinish);
};
const done = () => { cleanup(); resolve(); };
const resetTimer = () => {
if (timer) clearTimeout(timer);
timer = setTimeout(done, timeoutMs);
};
const onRequest = () => { pending++; resetTimer(); };
const onFinish = () => { pending = Math.max(0, pending - 1); if (pending <= 2) resetTimer(); };
cdp.on('Network.requestWillBeSent', onRequest);
cdp.on('Network.loadingFinished', onFinish);
cdp.on('Network.loadingFailed', onFinish);
resetTimer();
});
}
export async function waitForPageLoad(
cdp: CdpConnection,
sessionId: string,
timeoutMs: number = 30_000,
): Promise<void> {
void sessionId;
return new Promise((resolve) => {
const timer = setTimeout(() => {
cdp.off('Page.loadEventFired', handler);
resolve();
}, timeoutMs);
const handler = () => {
clearTimeout(timer);
cdp.off('Page.loadEventFired', handler);
resolve();
};
cdp.on('Page.loadEventFired', handler);
});
}
export async function createTargetAndAttach(
cdp: CdpConnection,
url: string,
): Promise<{ targetId: string; sessionId: string }> {
const { targetId } = await cdp.send<{ targetId: string }>('Target.createTarget', { url });
const { sessionId } = await cdp.send<{ sessionId: string }>('Target.attachToTarget', { targetId, flatten: true });
await cdp.send('Network.enable', {}, { sessionId });
await cdp.send('Page.enable', {}, { sessionId });
return { targetId, sessionId };
}
export async function navigateAndWait(
cdp: CdpConnection,
sessionId: string,
url: string,
timeoutMs: number,
): Promise<void> {
const loadPromise = new Promise<void>((resolve, reject) => {
const timer = setTimeout(() => reject(new Error('Page load timeout')), timeoutMs);
const handler = (params: unknown) => {
const event = params as { name?: string };
if (event.name === 'load' || event.name === 'DOMContentLoaded') {
clearTimeout(timer);
cdp.off('Page.lifecycleEvent', handler);
resolve();
}
};
cdp.on('Page.lifecycleEvent', handler);
});
await cdp.send('Page.navigate', { url }, { sessionId });
await loadPromise;
}
export async function evaluateScript<T>(
cdp: CdpConnection,
sessionId: string,
expression: string,
timeoutMs: number = 30_000,
): Promise<T> {
const result = await cdp.send<{ result: { value?: T } }>(
'Runtime.evaluate',
{ expression, returnByValue: true, awaitPromise: true },
{ sessionId, timeoutMs },
);
return result.result.value as T;
}
export async function autoScroll(
cdp: CdpConnection,
sessionId: string,
steps: number = 8,
waitMs: number = 600,
): Promise<void> {
let lastHeight = await evaluateScript<number>(cdp, sessionId, 'document.body.scrollHeight');
for (let i = 0; i < steps; i++) {
await evaluateScript<void>(cdp, sessionId, 'window.scrollTo(0, document.body.scrollHeight)');
await sleep(waitMs);
const newHeight = await evaluateScript<number>(cdp, sessionId, 'document.body.scrollHeight');
if (newHeight === lastHeight) break;
lastHeight = newHeight;
}
await evaluateScript<void>(cdp, sessionId, 'window.scrollTo(0, 0)');
}
@@ -1,13 +0,0 @@
import { resolveUrlToMarkdownChromeProfileDir } from "./paths.js";
export const DEFAULT_USER_AGENT =
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36";
export const USER_DATA_DIR = resolveUrlToMarkdownChromeProfileDir();
export const DEFAULT_TIMEOUT_MS = 30_000;
export const CDP_CONNECT_TIMEOUT_MS = 15_000;
export const NETWORK_IDLE_TIMEOUT_MS = 1_500;
export const POST_LOAD_DELAY_MS = 800;
export const SCROLL_STEP_WAIT_MS = 600;
export const SCROLL_MAX_STEPS = 8;
@@ -1,55 +0,0 @@
import assert from "node:assert/strict";
import test from "node:test";
import { cleanContent } from "./content-cleaner.js";
const SAMPLE_HTML = `<!doctype html>
<html>
<head>
<title>Example Story</title>
<style>.cookie-banner { position: fixed; }</style>
<script>window.__noise = true;</script>
</head>
<body>
<!-- comment that should be removed -->
<header>
<nav>
<a href="/home">Home</a>
<a href="/topics">Topics</a>
</nav>
</header>
<div class="cookie-banner">Accept cookies</div>
<aside>Sidebar links</aside>
<main>
<article class="content">
<h1>Actual Story Title</h1>
<p>
This is the first paragraph of the real story body, and it is intentionally long enough
to survive the cleaner's main-content heuristics without being mistaken for navigation.
</p>
<p>
This is the second paragraph with more useful detail, a
<a href="/read-more">supporting link</a>, and a normal image.
</p>
<img src="/images/cover.jpg" alt="Cover">
<img src="data:image/png;base64,AAAA" alt="Inline data">
</article>
</main>
<footer>Footer boilerplate</footer>
</body>
</html>`;
test("cleanContent keeps the article body and removes obvious boilerplate", () => {
const cleaned = cleanContent(SAMPLE_HTML, "https://example.com/posts/story");
assert.match(cleaned, /Actual Story Title/);
assert.match(cleaned, /https:\/\/example\.com\/read-more/);
assert.match(cleaned, /https:\/\/example\.com\/images\/cover\.jpg/);
assert.doesNotMatch(cleaned, /Accept cookies/);
assert.doesNotMatch(cleaned, /Sidebar links/);
assert.doesNotMatch(cleaned, /Footer boilerplate/);
assert.doesNotMatch(cleaned, /window\.__noise/);
assert.doesNotMatch(cleaned, /comment that should be removed/);
assert.doesNotMatch(cleaned, /data:image\/png;base64/);
});
@@ -1,58 +0,0 @@
import { JSDOM, VirtualConsole } from "jsdom";
import { Defuddle } from "defuddle/node";
import {
type ConversionResult,
type PageMetadata,
isMarkdownUsable,
normalizeMarkdown,
pickString,
} from "./markdown-conversion-shared.js";
export async function tryDefuddleConversion(
html: string,
url: string,
baseMetadata: PageMetadata
): Promise<{ ok: true; result: ConversionResult } | { ok: false; reason: string }> {
try {
const virtualConsole = new VirtualConsole();
virtualConsole.on("jsdomError", (error: Error & { type?: string }) => {
if (error.type === "css parsing" || /Could not parse CSS stylesheet/i.test(error.message)) {
return;
}
console.warn(`[url-to-markdown] jsdom: ${error.message}`);
});
const dom = new JSDOM(html, { url, virtualConsole });
const result = await Defuddle(dom, url, { markdown: true });
const markdown = normalizeMarkdown(result.content || "");
if (!isMarkdownUsable(markdown, html)) {
return { ok: false, reason: "Defuddle returned empty or incomplete markdown" };
}
return {
ok: true,
result: {
metadata: {
...baseMetadata,
title: pickString(result.title, baseMetadata.title) ?? "",
description: pickString(result.description, baseMetadata.description) ?? undefined,
author: pickString(result.author, baseMetadata.author) ?? undefined,
published: pickString(result.published, baseMetadata.published) ?? undefined,
coverImage: pickString(result.image, baseMetadata.coverImage) ?? undefined,
language: pickString(result.language, baseMetadata.language) ?? undefined,
},
markdown,
rawHtml: html,
conversionMethod: "defuddle",
variables: result.variables,
},
};
} catch (error) {
return {
ok: false,
reason: error instanceof Error ? error.message : String(error),
};
}
}
@@ -1,28 +0,0 @@
import assert from "node:assert/strict";
import test from "node:test";
import { extractContent } from "./html-to-markdown.js";
const EMBEDDED_IMAGE_HTML = `<!doctype html>
<html>
<body>
<main>
<article>
<h1>Embedded Image Story</h1>
<p>
This paragraph is intentionally long enough to satisfy the extractor thresholds so the
resulting markdown keeps the main article body and the embedded image reference.
</p>
<img src="data:image/png;base64,AAAA" alt="inline">
</article>
</main>
</body>
</html>`;
test("extractContent preserves base64 images when requested for media download", async () => {
const result = await extractContent(EMBEDDED_IMAGE_HTML, "https://example.com/embedded", {
preserveBase64Images: true,
});
assert.match(result.markdown, /!\[inline\]\(data:image\/png;base64,AAAA\)/);
});
@@ -1,162 +0,0 @@
import {
createMarkdownDocument,
extractMetadataFromHtml,
formatMetadataYaml,
type ConversionResult,
type PageMetadata,
isYouTubeUrl,
} from "./markdown-conversion-shared.js";
import { tryDefuddleConversion } from "./defuddle-converter.js";
import {
convertWithLegacyExtractor,
scoreMarkdownQuality,
shouldCompareWithLegacy,
} from "./legacy-converter.js";
import { tryUrlRuleParsers } from "./parsers/index.js";
import { cleanContent } from "./content-cleaner.js";
export type { ConversionResult, PageMetadata };
export { createMarkdownDocument, formatMetadataYaml };
export interface ExtractContentOptions {
preserveBase64Images?: boolean;
}
export const absolutizeUrlsScript = String.raw`
(function() {
const baseUrl = document.baseURI || location.href;
const htmlClone = document.documentElement.cloneNode(true);
function materializeShadowDom(sourceRoot, cloneRoot) {
const sourceElements = Array.from(sourceRoot.querySelectorAll("*"));
const cloneElements = Array.from(cloneRoot.querySelectorAll("*"));
for (let i = sourceElements.length - 1; i >= 0; i--) {
const sourceEl = sourceElements[i];
const cloneEl = cloneElements[i];
const shadowRoot = sourceEl && sourceEl.shadowRoot;
if (!shadowRoot || !cloneEl || !shadowRoot.innerHTML) continue;
if (cloneEl.tagName && cloneEl.tagName.includes("-")) {
const wrapper = document.createElement("div");
wrapper.setAttribute("data-shadow-host", cloneEl.tagName.toLowerCase());
wrapper.innerHTML = shadowRoot.innerHTML;
cloneEl.replaceWith(wrapper);
} else {
cloneEl.innerHTML = shadowRoot.innerHTML;
}
}
}
function toAbsolute(url) {
if (!url) return url;
try { return new URL(url, baseUrl).href; } catch { return url; }
}
function absAttr(root, sel, attr) {
root.querySelectorAll(sel).forEach(el => {
const v = el.getAttribute(attr);
if (v) {
const a = toAbsolute(v);
if (a) el.setAttribute(attr, a);
}
});
}
function absSrcset(root, sel) {
root.querySelectorAll(sel).forEach(el => {
const s = el.getAttribute("srcset");
if (!s) return;
el.setAttribute("srcset", s.split(",").map(p => {
const t = p.trim();
if (!t) return "";
const [url, ...d] = t.split(/\s+/);
return d.length ? toAbsolute(url) + " " + d.join(" ") : toAbsolute(url);
}).filter(Boolean).join(", "));
});
}
materializeShadowDom(document.documentElement, htmlClone);
htmlClone.querySelectorAll("img[data-src], video[data-src], audio[data-src], source[data-src]").forEach(el => {
const ds = el.getAttribute("data-src");
if (ds && (!el.getAttribute("src") || el.getAttribute("src") === "" || el.getAttribute("src")?.startsWith("data:"))) {
el.setAttribute("src", ds);
}
});
absAttr(htmlClone, "a[href]", "href");
absAttr(htmlClone, "img[src], video[src], audio[src], source[src], iframe[src]", "src");
absAttr(htmlClone, "video[poster]", "poster");
absSrcset(htmlClone, "img[srcset], source[srcset]");
return {
html: "<!doctype html>\n" + htmlClone.outerHTML,
finalUrl: location.href,
};
})()
`;
function shouldPreferDefuddle(result: ConversionResult): boolean {
if (isYouTubeUrl(result.metadata.url)) {
return true;
}
const transcript = result.variables?.transcript?.trim();
if (transcript) {
return true;
}
return /^##?\s+transcript\b/im.test(result.markdown);
}
export async function extractContent(
html: string,
url: string,
options: ExtractContentOptions = {}
): Promise<ConversionResult> {
const capturedAt = new Date().toISOString();
const baseMetadata = extractMetadataFromHtml(html, url, capturedAt);
const specializedResult = tryUrlRuleParsers(html, url, baseMetadata);
if (specializedResult) {
return specializedResult;
}
let cleanedHtml = html;
try {
cleanedHtml = cleanContent(html, url, {
removeBase64Images: !options.preserveBase64Images,
});
} catch {
cleanedHtml = html;
}
const defuddleResult = await tryDefuddleConversion(cleanedHtml, url, baseMetadata);
if (defuddleResult.ok) {
if (shouldPreferDefuddle(defuddleResult.result)) {
return { ...defuddleResult.result, rawHtml: html };
}
if (shouldCompareWithLegacy(defuddleResult.result.markdown)) {
const legacyResult = convertWithLegacyExtractor(html, baseMetadata, cleanedHtml);
const legacyScore = scoreMarkdownQuality(legacyResult.markdown);
const defuddleScore = scoreMarkdownQuality(defuddleResult.result.markdown);
if (legacyScore > defuddleScore + 120) {
return {
...legacyResult,
fallbackReason: "Legacy extractor produced higher-quality markdown than Defuddle",
};
}
}
return { ...defuddleResult.result, rawHtml: html };
}
const fallbackResult = convertWithLegacyExtractor(html, baseMetadata, cleanedHtml);
return {
...fallbackResult,
fallbackReason: defuddleResult.reason,
};
}
@@ -1,48 +0,0 @@
import assert from "node:assert/strict";
import test from "node:test";
import { cleanContent } from "./content-cleaner.js";
import { convertWithLegacyExtractor } from "./legacy-converter.js";
import { extractMetadataFromHtml } from "./markdown-conversion-shared.js";
const CAPTURED_AT = "2026-03-24T03:00:00.000Z";
const NEXT_DATA_HTML = `<!doctype html>
<html>
<head>
<title>Hydrated Story</title>
</head>
<body>
<div class="cookie-banner">Accept cookies</div>
<main>
<p>Short teaser text that should not win over the structured article payload.</p>
</main>
<script id="__NEXT_DATA__" type="application/json">
{
"props": {
"pageProps": {
"article": {
"title": "Hydrated Story",
"description": "A structured article payload from Next.js",
"body": "<p>The full article lives in __NEXT_DATA__ and should still be extracted even when the cleaned HTML removes scripts before the selector and readability passes run.</p><p>A second paragraph keeps the content comfortably above the minimum extraction threshold and proves the legacy extractor still has access to the original structured payload.</p>"
}
}
}
}
</script>
</body>
</html>`;
test("legacy extractor still uses original __NEXT_DATA__ after HTML cleaning", () => {
const url = "https://example.com/posts/hydrated-story";
const baseMetadata = extractMetadataFromHtml(NEXT_DATA_HTML, url, CAPTURED_AT);
const cleanedHtml = cleanContent(NEXT_DATA_HTML, url);
const result = convertWithLegacyExtractor(NEXT_DATA_HTML, baseMetadata, cleanedHtml);
assert.equal(result.conversionMethod, "legacy:next-data");
assert.match(result.markdown, /The full article lives in .*NEXT.*DATA/);
assert.match(result.markdown, /A second paragraph keeps the content comfortably above the minimum extraction threshold/);
assert.doesNotMatch(result.markdown, /Short teaser text that should not win/);
assert.equal(result.rawHtml, NEXT_DATA_HTML);
});
@@ -1,641 +0,0 @@
import { Readability } from "@mozilla/readability";
import TurndownService from "turndown";
import { gfm } from "turndown-plugin-gfm";
import {
type AnyRecord,
type ConversionResult,
type PageMetadata,
GOOD_CONTENT_LENGTH,
MIN_CONTENT_LENGTH,
extractPublishedTime,
extractTextFromHtml,
extractTitle,
normalizeMarkdown,
parseDocument,
pickString,
sanitizeHtml,
} from "./markdown-conversion-shared.js";
interface ExtractionCandidate {
title: string | null;
byline: string | null;
excerpt: string | null;
published: string | null;
html: string | null;
textContent: string;
method: string;
}
const CONTENT_SELECTORS = [
"article",
"main article",
"[role='main'] article",
"[itemprop='articleBody']",
".article-content",
".article-body",
".post-content",
".entry-content",
".story-body",
"main",
"[role='main']",
"#content",
".content",
];
const REMOVE_SELECTORS = [
"script",
"style",
"noscript",
"template",
"iframe",
"svg",
"path",
"nav",
"aside",
"footer",
"header",
"form",
".advertisement",
".ads",
".social-share",
".related-articles",
".comments",
".newsletter",
".cookie-banner",
".cookie-consent",
"[role='navigation']",
"[aria-label*='cookie' i]",
];
const NEXT_DATA_CONTENT_PATHS = [
"props.pageProps.content.body",
"props.pageProps.article.body",
"props.pageProps.article.content",
"props.pageProps.post.body",
"props.pageProps.post.content",
"props.pageProps.data.body",
"props.pageProps.story.body.content",
];
const LOW_QUALITY_MARKERS = [
/Join The Conversation/i,
/One Community\. Many Voices/i,
/Read our community guidelines/i,
/Create a free account to share your thoughts/i,
/Become a Forbes Member/i,
/Subscribe to trusted journalism/i,
/\bComments\b/i,
];
function generateExcerpt(excerpt: string | null, textContent: string | null): string | null {
if (excerpt) return excerpt;
if (!textContent) return null;
const trimmed = textContent.trim();
if (!trimmed) return null;
return trimmed.length > 200 ? `${trimmed.slice(0, 200)}...` : trimmed;
}
function parseJsonLdItem(item: AnyRecord): ExtractionCandidate | null {
const type = Array.isArray(item["@type"]) ? item["@type"][0] : item["@type"];
if (typeof type !== "string" || !["Article", "NewsArticle", "BlogPosting", "WebPage", "ReportageNewsArticle"].includes(type)) {
return null;
}
const rawContent =
(typeof item.articleBody === "string" && item.articleBody) ||
(typeof item.text === "string" && item.text) ||
(typeof item.description === "string" && item.description) ||
null;
if (!rawContent) return null;
const content = rawContent.trim();
const htmlLike = /<\/?[a-z][\s\S]*>/i.test(content);
const textContent = htmlLike ? extractTextFromHtml(content) : content;
if (textContent.length < MIN_CONTENT_LENGTH) return null;
return {
title: pickString(item.headline, item.name),
byline: extractAuthorFromJsonLd(item.author),
excerpt: pickString(item.description),
published: pickString(item.datePublished, item.dateCreated),
html: htmlLike ? content : null,
textContent,
method: "json-ld",
};
}
function extractAuthorFromJsonLd(authorData: unknown): string | null {
if (typeof authorData === "string") return authorData;
if (!authorData || typeof authorData !== "object") return null;
if (Array.isArray(authorData)) {
const names = authorData
.map((author) => extractAuthorFromJsonLd(author))
.filter((name): name is string => Boolean(name));
return names.length > 0 ? names.join(", ") : null;
}
const author = authorData as AnyRecord;
return typeof author.name === "string" ? author.name : null;
}
function flattenJsonLdItems(data: unknown): AnyRecord[] {
if (!data || typeof data !== "object") return [];
if (Array.isArray(data)) return data.flatMap(flattenJsonLdItems);
const item = data as AnyRecord;
if (Array.isArray(item["@graph"])) {
return (item["@graph"] as unknown[]).flatMap(flattenJsonLdItems);
}
return [item];
}
function tryJsonLdExtraction(document: Document): ExtractionCandidate | null {
const scripts = document.querySelectorAll("script[type='application/ld+json']");
for (const script of scripts) {
try {
const data = JSON.parse(script.textContent ?? "");
for (const item of flattenJsonLdItems(data)) {
const extracted = parseJsonLdItem(item);
if (extracted) return extracted;
}
} catch {
// Ignore malformed blocks.
}
}
return null;
}
function getByPath(value: unknown, path: string): unknown {
let current = value;
for (const part of path.split(".")) {
if (!current || typeof current !== "object") return undefined;
current = (current as AnyRecord)[part];
}
return current;
}
function isContentBlockArray(value: unknown): value is AnyRecord[] {
if (!Array.isArray(value) || value.length === 0) return false;
return value.slice(0, 5).some((item) => {
if (!item || typeof item !== "object") return false;
const obj = item as AnyRecord;
return "type" in obj || "text" in obj || "textHtml" in obj || "content" in obj;
});
}
function extractTextFromContentBlocks(blocks: AnyRecord[]): string {
const parts: string[] = [];
function pushParagraph(text: string): void {
const trimmed = text.trim();
if (!trimmed) return;
parts.push(trimmed, "\n\n");
}
function walk(node: unknown): void {
if (!node || typeof node !== "object") return;
const block = node as AnyRecord;
if (typeof block.text === "string") {
pushParagraph(block.text);
return;
}
if (typeof block.textHtml === "string") {
pushParagraph(extractTextFromHtml(block.textHtml));
return;
}
if (Array.isArray(block.items)) {
for (const item of block.items) {
if (item && typeof item === "object") {
const text = pickString((item as AnyRecord).text);
if (text) parts.push(`- ${text}\n`);
}
}
parts.push("\n");
}
if (Array.isArray(block.components)) {
for (const component of block.components) {
walk(component);
}
}
if (Array.isArray(block.content)) {
for (const child of block.content) {
walk(child);
}
}
}
for (const block of blocks) {
walk(block);
}
return parts.join("").replace(/\n{3,}/g, "\n\n").trim();
}
function tryStringBodyExtraction(
content: string,
meta: AnyRecord,
document: Document,
method: string
): ExtractionCandidate | null {
if (!content || content.length < MIN_CONTENT_LENGTH) return null;
const isHtml = /<\/?[a-z][\s\S]*>/i.test(content);
const html = isHtml ? sanitizeHtml(content) : null;
const textContent = isHtml ? extractTextFromHtml(html) : content.trim();
if (textContent.length < MIN_CONTENT_LENGTH) return null;
return {
title: pickString(meta.headline, meta.title, extractTitle(document)),
byline: pickString(meta.byline, meta.author),
excerpt: pickString(meta.description, meta.excerpt, generateExcerpt(null, textContent)),
published: pickString(meta.datePublished, meta.publishedAt, extractPublishedTime(document)),
html,
textContent,
method,
};
}
function tryNextDataExtraction(document: Document): ExtractionCandidate | null {
try {
const script = document.querySelector("script#__NEXT_DATA__");
if (!script?.textContent) return null;
const data = JSON.parse(script.textContent) as AnyRecord;
const pageProps = (getByPath(data, "props.pageProps") ?? {}) as AnyRecord;
for (const path of NEXT_DATA_CONTENT_PATHS) {
const value = getByPath(data, path);
if (typeof value === "string") {
const parentPath = path.split(".").slice(0, -1).join(".");
const parent = (getByPath(data, parentPath) ?? {}) as AnyRecord;
const meta = {
...pageProps,
...parent,
title: parent.title ?? (pageProps.title as string | undefined),
};
const candidate = tryStringBodyExtraction(value, meta, document, "next-data");
if (candidate) return candidate;
}
if (isContentBlockArray(value)) {
const textContent = extractTextFromContentBlocks(value);
if (textContent.length < MIN_CONTENT_LENGTH) continue;
return {
title: pickString(
getByPath(data, "props.pageProps.content.headline"),
getByPath(data, "props.pageProps.article.headline"),
getByPath(data, "props.pageProps.article.title"),
getByPath(data, "props.pageProps.post.title"),
pageProps.title,
extractTitle(document)
),
byline: pickString(
getByPath(data, "props.pageProps.author.name"),
getByPath(data, "props.pageProps.article.author.name")
),
excerpt: pickString(
getByPath(data, "props.pageProps.content.description"),
getByPath(data, "props.pageProps.article.description"),
pageProps.description,
generateExcerpt(null, textContent)
),
published: pickString(
getByPath(data, "props.pageProps.content.datePublished"),
getByPath(data, "props.pageProps.article.datePublished"),
getByPath(data, "props.pageProps.publishedAt"),
extractPublishedTime(document)
),
html: null,
textContent,
method: "next-data",
};
}
}
} catch {
return null;
}
return null;
}
function buildReadabilityCandidate(
article: ReturnType<Readability["parse"]>,
referenceDocument: Document,
method: string
): ExtractionCandidate | null {
const textContent = article?.textContent?.trim() ?? "";
if (textContent.length < MIN_CONTENT_LENGTH) return null;
return {
title: pickString(article?.title, extractTitle(referenceDocument)),
byline: pickString((article as { byline?: string } | null)?.byline),
excerpt: pickString(article?.excerpt, generateExcerpt(null, textContent)),
published: pickString(
(article as { publishedTime?: string } | null)?.publishedTime,
extractPublishedTime(referenceDocument)
),
html: article?.content ? sanitizeHtml(article.content) : null,
textContent,
method,
};
}
function tryReadability(document: Document, referenceDocument: Document = document): ExtractionCandidate | null {
try {
const strictClone = document.cloneNode(true) as Document;
const strictResult = buildReadabilityCandidate(
new Readability(strictClone).parse(),
referenceDocument,
"readability"
);
if (strictResult) return strictResult;
const relaxedClone = document.cloneNode(true) as Document;
return buildReadabilityCandidate(
new Readability(relaxedClone, { charThreshold: 120 }).parse(),
referenceDocument,
"readability-relaxed"
);
} catch {
return null;
}
}
function trySelectorExtraction(document: Document): ExtractionCandidate | null {
for (const selector of CONTENT_SELECTORS) {
const element = document.querySelector(selector);
if (!element) continue;
const clone = element.cloneNode(true) as Element;
for (const removeSelector of REMOVE_SELECTORS) {
for (const node of clone.querySelectorAll(removeSelector)) {
node.remove();
}
}
const html = sanitizeHtml(clone.innerHTML);
const textContent = extractTextFromHtml(html);
if (textContent.length < MIN_CONTENT_LENGTH) continue;
return {
title: extractTitle(document),
byline: null,
excerpt: generateExcerpt(null, textContent),
published: extractPublishedTime(document),
html,
textContent,
method: `selector:${selector}`,
};
}
return null;
}
function tryBodyExtraction(document: Document): ExtractionCandidate | null {
const body = document.body;
if (!body) return null;
const clone = body.cloneNode(true) as Element;
for (const removeSelector of REMOVE_SELECTORS) {
for (const node of clone.querySelectorAll(removeSelector)) {
node.remove();
}
}
const html = sanitizeHtml(clone.innerHTML);
const textContent = extractTextFromHtml(html);
if (!textContent) return null;
return {
title: extractTitle(document),
byline: null,
excerpt: generateExcerpt(null, textContent),
published: extractPublishedTime(document),
html,
textContent,
method: "body-fallback",
};
}
function pickBestCandidate(candidates: ExtractionCandidate[]): ExtractionCandidate | null {
if (candidates.length === 0) return null;
const methodOrder = [
"readability",
"readability-relaxed",
"next-data",
"json-ld",
"selector:",
"body-fallback",
];
function methodRank(method: string): number {
const idx = methodOrder.findIndex((entry) =>
entry.endsWith(":") ? method.startsWith(entry) : method === entry
);
return idx === -1 ? methodOrder.length : idx;
}
const ranked = [...candidates].sort((a, b) => {
const rankA = methodRank(a.method);
const rankB = methodRank(b.method);
if (rankA !== rankB) return rankA - rankB;
return (b.textContent.length ?? 0) - (a.textContent.length ?? 0);
});
for (const candidate of ranked) {
if (candidate.textContent.length >= GOOD_CONTENT_LENGTH) {
return candidate;
}
}
for (const candidate of ranked) {
if (candidate.textContent.length >= MIN_CONTENT_LENGTH) {
return candidate;
}
}
return ranked[0];
}
function extractFromHtml(html: string, cleanedHtml: string = html): ExtractionCandidate | null {
const originalDocument = parseDocument(html);
const cleanedDocument = parseDocument(cleanedHtml);
const readabilityCandidate = tryReadability(cleanedDocument, originalDocument);
const nextDataCandidate = tryNextDataExtraction(originalDocument);
const jsonLdCandidate = tryJsonLdExtraction(originalDocument);
const selectorCandidate = trySelectorExtraction(cleanedDocument);
const bodyCandidate = tryBodyExtraction(cleanedDocument);
const candidates = [
readabilityCandidate,
nextDataCandidate,
jsonLdCandidate,
selectorCandidate,
bodyCandidate,
].filter((candidate): candidate is ExtractionCandidate => Boolean(candidate));
const winner = pickBestCandidate(candidates);
if (!winner) return null;
return {
...winner,
title: winner.title ?? extractTitle(originalDocument),
published: winner.published ?? extractPublishedTime(originalDocument),
excerpt: winner.excerpt ?? generateExcerpt(null, winner.textContent),
};
}
const turndown = new TurndownService({
headingStyle: "atx",
hr: "---",
bulletListMarker: "-",
codeBlockStyle: "fenced",
emDelimiter: "*",
strongDelimiter: "**",
linkStyle: "inlined",
});
turndown.use(gfm);
turndown.remove(["script", "style", "iframe", "noscript", "template", "svg", "path"]);
turndown.addRule("collapseFigure", {
filter: "figure",
replacement(content) {
return `\n\n${content.trim()}\n\n`;
},
});
turndown.addRule("dropInvisibleAnchors", {
filter(node) {
return (
node.nodeName === "A" &&
!(node as Element).textContent?.trim() &&
!(node as Element).querySelector("img, video, picture, source")
);
},
replacement() {
return "";
},
});
export function convertHtmlFragmentToMarkdown(html: string): string {
if (!html || !html.trim()) return "";
try {
const sanitized = sanitizeHtml(html);
return turndown.turndown(sanitized);
} catch {
return "";
}
}
function fallbackPlainText(html: string): string {
const document = parseDocument(html);
for (const selector of ["script", "style", "noscript", "template", "iframe", "svg", "path"]) {
for (const el of document.querySelectorAll(selector)) {
el.remove();
}
}
const text = document.body?.textContent ?? document.documentElement?.textContent ?? "";
return normalizeMarkdown(text.replace(/\s+/g, " "));
}
function countBylines(markdown: string): number {
return (markdown.match(/(^|\n)By\s+/g) || []).length;
}
function countUsefulParagraphs(markdown: string): number {
const paragraphs = normalizeMarkdown(markdown).split(/\n{2,}/);
let count = 0;
for (const paragraph of paragraphs) {
const trimmed = paragraph.trim();
if (!trimmed) continue;
if (/^!?\[[^\]]*\]\([^)]+\)$/.test(trimmed)) continue;
if (/^#{1,6}\s+/.test(trimmed)) continue;
if ((trimmed.match(/\b[\p{L}\p{N}']+\b/gu) || []).length < 8) continue;
count++;
}
return count;
}
function countMarkerHits(markdown: string, markers: RegExp[]): number {
let hits = 0;
for (const marker of markers) {
if (marker.test(markdown)) hits++;
}
return hits;
}
export function scoreMarkdownQuality(markdown: string): number {
const normalized = normalizeMarkdown(markdown);
const wordCount = (normalized.match(/\b[\p{L}\p{N}']+\b/gu) || []).length;
const usefulParagraphs = countUsefulParagraphs(normalized);
const headingCount = (normalized.match(/^#{1,6}\s+/gm) || []).length;
const markerHits = countMarkerHits(normalized, LOW_QUALITY_MARKERS);
const bylineCount = countBylines(normalized);
const staffCount = (normalized.match(/\bForbes Staff\b/gi) || []).length;
return (
Math.min(wordCount, 4000) +
usefulParagraphs * 40 +
headingCount * 10 -
markerHits * 180 -
Math.max(0, bylineCount - 1) * 120 -
Math.max(0, staffCount - 1) * 80
);
}
export function shouldCompareWithLegacy(markdown: string): boolean {
const normalized = normalizeMarkdown(markdown);
return (
countMarkerHits(normalized, LOW_QUALITY_MARKERS) > 0 ||
countBylines(normalized) > 1 ||
countUsefulParagraphs(normalized) < 6
);
}
export function convertWithLegacyExtractor(
html: string,
baseMetadata: PageMetadata,
cleanedHtml: string = html
): ConversionResult {
const extracted = extractFromHtml(html, cleanedHtml);
let markdown = extracted?.html ? convertHtmlFragmentToMarkdown(extracted.html) : "";
if (!markdown.trim()) {
markdown = extracted?.textContent?.trim() || fallbackPlainText(cleanedHtml);
}
return {
metadata: {
...baseMetadata,
title: pickString(extracted?.title, baseMetadata.title) ?? "",
description: pickString(extracted?.excerpt, baseMetadata.description) ?? undefined,
author: pickString(extracted?.byline, baseMetadata.author) ?? undefined,
published: pickString(extracted?.published, baseMetadata.published) ?? undefined,
},
markdown: normalizeMarkdown(markdown),
rawHtml: html,
conversionMethod: extracted ? `legacy:${extracted.method}` : "legacy:plain-text",
};
}
@@ -1,473 +0,0 @@
import { createInterface } from "node:readline";
import { writeFile, mkdir, access } from "node:fs/promises";
import path from "node:path";
import process from "node:process";
import { CdpConnection, getFreePort, findExistingChromePort, launchChrome, waitForChromeDebugPort, waitForNetworkIdle, waitForPageLoad, autoScroll, evaluateScript, killChrome } from "./cdp.js";
import { absolutizeUrlsScript, extractContent, createMarkdownDocument, type ConversionResult } from "./html-to-markdown.js";
import { localizeMarkdownMedia, countRemoteMedia } from "./media-localizer.js";
import { resolveUrlToMarkdownDataDir } from "./paths.js";
import { DEFAULT_TIMEOUT_MS, CDP_CONNECT_TIMEOUT_MS, NETWORK_IDLE_TIMEOUT_MS, POST_LOAD_DELAY_MS, SCROLL_STEP_WAIT_MS, SCROLL_MAX_STEPS } from "./constants.js";
function sleep(ms: number): Promise<void> {
return new Promise((resolve) => setTimeout(resolve, ms));
}
async function fileExists(filePath: string): Promise<boolean> {
try {
await access(filePath);
return true;
} catch {
return false;
}
}
interface Args {
url: string;
output?: string;
outputDir?: string;
wait: boolean;
timeout: number;
downloadMedia: boolean;
browserMode: BrowserMode;
}
type BrowserMode = "auto" | "headless" | "headed";
interface CaptureAttemptOptions {
headless: boolean;
wait: boolean;
existingPort?: number;
waitPrompt?: string;
}
interface CaptureSnapshot {
html: string;
finalUrl: string;
}
const BROWSER_MODES = new Set<BrowserMode>(["auto", "headless", "headed"]);
function parseArgs(argv: string[]): Args {
const args: Args = {
url: "",
wait: false,
timeout: DEFAULT_TIMEOUT_MS,
downloadMedia: false,
browserMode: "auto",
};
for (let i = 2; i < argv.length; i++) {
const arg = argv[i];
if (arg === "--wait" || arg === "-w") {
args.wait = true;
} else if (arg === "-o" || arg === "--output") {
args.output = argv[++i];
} else if (arg === "--timeout" || arg === "-t") {
args.timeout = parseInt(argv[++i], 10) || DEFAULT_TIMEOUT_MS;
} else if (arg === "--output-dir") {
args.outputDir = argv[++i];
} else if (arg === "--download-media") {
args.downloadMedia = true;
} else if (arg === "--browser") {
args.browserMode = (argv[++i] as BrowserMode | undefined) ?? "auto";
} else if (arg === "--headless") {
args.browserMode = "headless";
} else if (arg === "--headed" || arg === "--noheadless" || arg === "--no-headless") {
args.browserMode = "headed";
} else if (!arg.startsWith("-") && !args.url) {
args.url = arg;
}
}
return args;
}
const SLUG_STOP_WORDS = new Set([
"the", "a", "an", "is", "are", "was", "were", "be", "been", "being",
"have", "has", "had", "do", "does", "did", "will", "would", "shall",
"should", "may", "might", "must", "can", "could", "to", "of", "in",
"for", "on", "with", "at", "by", "from", "as", "into", "through",
"during", "before", "after", "above", "below", "between", "out",
"off", "over", "under", "again", "further", "then", "once", "here",
"there", "when", "where", "why", "how", "all", "both", "each",
"few", "more", "most", "other", "some", "such", "no", "nor", "not",
"only", "own", "same", "so", "than", "too", "very", "just", "but",
"and", "or", "if", "this", "that", "these", "those", "it", "its",
"http", "https", "www", "com", "org", "net", "post", "article",
]);
function extractSlugFromContent(content: string): string | null {
const body = content.replace(/^---\n[\s\S]*?\n---\n?/, "").slice(0, 1000);
const words = body
.replace(/[^\w\s-]/g, " ")
.split(/\s+/)
.filter((w) => /^[a-zA-Z]/.test(w) && w.length >= 2 && !SLUG_STOP_WORDS.has(w.toLowerCase()))
.map((w) => w.toLowerCase());
const unique: string[] = [];
const seen = new Set<string>();
for (const w of words) {
if (!seen.has(w)) {
seen.add(w);
unique.push(w);
if (unique.length >= 6) break;
}
}
return unique.length >= 2 ? unique.join("-").slice(0, 50) : null;
}
function generateSlug(title: string, url: string, content?: string): string {
const asciiWords = title
.replace(/[^\w\s]/g, " ")
.split(/\s+/)
.filter((w) => /[a-zA-Z]/.test(w) && w.length >= 2 && !SLUG_STOP_WORDS.has(w.toLowerCase()))
.map((w) => w.toLowerCase());
if (asciiWords.length >= 2) {
return asciiWords.slice(0, 6).join("-").slice(0, 50);
}
if (content) {
const contentSlug = extractSlugFromContent(content);
if (contentSlug) return contentSlug;
}
const GENERIC_PATH_SEGMENTS = new Set(["status", "article", "post", "posts", "p", "blog", "news", "articles"]);
const parsed = new URL(url);
const pathSlug = parsed.pathname
.split("/")
.filter((s) => s.length > 0 && !/^\d{10,}$/.test(s) && !GENERIC_PATH_SEGMENTS.has(s.toLowerCase()))
.join("-")
.toLowerCase()
.replace(/[^\w-]/g, "-")
.replace(/-+/g, "-")
.replace(/^-|-$/g, "")
.slice(0, 40);
const prefix = asciiWords.slice(0, 2).join("-");
const combined = prefix ? `${prefix}-${pathSlug}` : pathSlug;
return combined.slice(0, 50) || "page";
}
function formatTimestamp(): string {
const now = new Date();
const pad = (n: number) => n.toString().padStart(2, "0");
return `${now.getFullYear()}${pad(now.getMonth() + 1)}${pad(now.getDate())}-${pad(now.getHours())}${pad(now.getMinutes())}${pad(now.getSeconds())}`;
}
function deriveHtmlSnapshotPath(markdownPath: string): string {
const parsed = path.parse(markdownPath);
const basename = parsed.ext ? parsed.name : parsed.base;
return path.join(parsed.dir, `${basename}-captured.html`);
}
function extractTitleFromMarkdownDocument(document: string): string {
const normalized = document.replace(/\r\n/g, "\n");
const frontmatterMatch = normalized.match(/^---\n([\s\S]*?)\n---\n?/);
if (frontmatterMatch) {
const titleLine = frontmatterMatch[1]
.split("\n")
.find((line) => /^title:\s*/i.test(line));
if (titleLine) {
const rawValue = titleLine.replace(/^title:\s*/i, "").trim();
const unquoted = rawValue
.replace(/^"(.*)"$/, "$1")
.replace(/^'(.*)'$/, "$1")
.replace(/\\"/g, '"');
if (unquoted) return unquoted;
}
}
const headingMatch = normalized.match(/^#\s+(.+)$/m);
return headingMatch?.[1]?.trim() ?? "";
}
function buildDefuddleApiUrl(targetUrl: string): string {
return `https://defuddle.md/${encodeURIComponent(targetUrl)}`;
}
async function fetchDefuddleApiMarkdown(targetUrl: string): Promise<{ markdown: string; title: string }> {
const apiUrl = buildDefuddleApiUrl(targetUrl);
const response = await fetch(apiUrl, {
headers: {
accept: "text/markdown,text/plain;q=0.9,*/*;q=0.1",
},
});
if (!response.ok) {
throw new Error(`defuddle.md returned ${response.status} ${response.statusText}`);
}
const markdown = (await response.text()).replace(/\r\n/g, "\n").trim();
if (!markdown) {
throw new Error("defuddle.md returned empty markdown");
}
return {
markdown,
title: extractTitleFromMarkdownDocument(markdown),
};
}
async function generateOutputPath(url: string, title: string, outputDir?: string, content?: string): Promise<string> {
const domain = new URL(url).hostname.replace(/^www\./, "");
const slug = generateSlug(title, url, content);
const dataDir = outputDir ? path.resolve(outputDir) : resolveUrlToMarkdownDataDir();
const basePath = path.join(dataDir, domain, slug, `${slug}.md`);
if (!(await fileExists(basePath))) {
return basePath;
}
const timestampSlug = `${slug}-${formatTimestamp()}`;
return path.join(dataDir, domain, timestampSlug, `${timestampSlug}.md`);
}
function defaultWaitPrompt(): string {
return "A browser window has been opened. If the page requires login or verification, complete it first, then press Enter to capture.";
}
async function waitForUserSignal(prompt: string): Promise<void> {
console.log(prompt);
const rl = createInterface({ input: process.stdin, output: process.stdout });
await new Promise<void>((resolve) => {
rl.once("line", () => { rl.close(); resolve(); });
});
}
async function captureUrlOnce(args: Args, options: CaptureAttemptOptions): Promise<ConversionResult> {
const reusing = options.existingPort !== undefined;
const port = options.existingPort ?? await getFreePort();
const chrome = reusing ? null : await launchChrome(args.url, port, options.headless);
if (reusing) {
console.log(`Reusing existing Chrome on port ${port}`);
} else {
console.log(`Launching Chrome (${options.headless ? "headless" : "headed"})...`);
}
let cdp: CdpConnection | null = null;
let targetId: string | null = null;
try {
const wsUrl = await waitForChromeDebugPort(port, 30_000);
cdp = await CdpConnection.connect(wsUrl, CDP_CONNECT_TIMEOUT_MS);
let sessionId: string;
if (reusing) {
const created = await cdp.send<{ targetId: string }>("Target.createTarget", { url: args.url });
targetId = created.targetId;
const attached = await cdp.send<{ sessionId: string }>("Target.attachToTarget", { targetId, flatten: true });
sessionId = attached.sessionId;
await cdp.send("Network.enable", {}, { sessionId });
await cdp.send("Page.enable", {}, { sessionId });
} else {
const targets = await cdp.send<{ targetInfos: Array<{ targetId: string; type: string; url: string }> }>("Target.getTargets");
const pageTarget = targets.targetInfos.find(t => t.type === "page" && t.url.startsWith("http"));
if (!pageTarget) throw new Error("No page target found");
targetId = pageTarget.targetId;
const attached = await cdp.send<{ sessionId: string }>("Target.attachToTarget", { targetId, flatten: true });
sessionId = attached.sessionId;
await cdp.send("Network.enable", {}, { sessionId });
await cdp.send("Page.enable", {}, { sessionId });
}
if (options.wait) {
await waitForUserSignal(options.waitPrompt ?? defaultWaitPrompt());
} else {
console.log("Waiting for page to load...");
await Promise.race([
waitForPageLoad(cdp, sessionId, 15_000),
sleep(8_000)
]);
await waitForNetworkIdle(cdp, sessionId, NETWORK_IDLE_TIMEOUT_MS);
await sleep(POST_LOAD_DELAY_MS);
console.log("Scrolling to trigger lazy load...");
await autoScroll(cdp, sessionId, SCROLL_MAX_STEPS, SCROLL_STEP_WAIT_MS);
await sleep(POST_LOAD_DELAY_MS);
}
console.log("Capturing page content...");
const snapshot = await evaluateScript<CaptureSnapshot>(
cdp, sessionId, absolutizeUrlsScript, args.timeout
);
return await extractContent(snapshot.html, snapshot.finalUrl || args.url, {
preserveBase64Images: args.downloadMedia,
});
} finally {
if (reusing) {
if (cdp && targetId) {
try { await cdp.send("Target.closeTarget", { targetId }, { timeoutMs: 5_000 }); } catch {}
}
if (cdp) cdp.close();
} else {
if (cdp) {
try { await cdp.send("Browser.close", {}, { timeoutMs: 5_000 }); } catch {}
cdp.close();
}
if (chrome) killChrome(chrome);
}
}
}
async function runHeadedFlow(
args: Args,
options: { existingPort?: number; wait: boolean; waitPrompt?: string }
): Promise<ConversionResult> {
return await captureUrlOnce(args, {
headless: false,
wait: options.wait,
existingPort: options.existingPort,
waitPrompt: options.waitPrompt,
});
}
async function captureUrl(args: Args): Promise<ConversionResult> {
const existingPort = await findExistingChromePort();
if (existingPort !== null) {
console.log("Found an existing Chrome session for this profile. Reusing it instead of launching a new browser.");
return await runHeadedFlow(args, {
existingPort,
wait: args.wait,
waitPrompt: args.wait ? defaultWaitPrompt() : undefined,
});
}
if (args.browserMode === "headless") {
return await captureUrlOnce(args, { headless: true, wait: false });
}
if (args.browserMode === "headed") {
return await runHeadedFlow(args, {
wait: args.wait,
waitPrompt: args.wait ? defaultWaitPrompt() : undefined,
});
}
if (args.wait) {
return await runHeadedFlow(args, {
wait: true,
waitPrompt: defaultWaitPrompt(),
});
}
try {
return await captureUrlOnce(args, { headless: true, wait: false });
} catch (error) {
const headlessMessage = error instanceof Error ? error.message : String(error);
console.warn(`Headless capture failed: ${headlessMessage}`);
console.log("Retrying with a visible browser window...");
try {
return await runHeadedFlow(args, { wait: false });
} catch (headedError) {
const headedMessage = headedError instanceof Error ? headedError.message : String(headedError);
throw new Error(`Headless capture failed (${headlessMessage}); headed retry failed (${headedMessage})`);
}
}
}
async function main(): Promise<void> {
const args = parseArgs(process.argv);
if (!args.url) {
console.error("Usage: bun main.ts <url> [-o output.md] [--output-dir dir] [--wait] [--browser auto|headless|headed] [--timeout ms] [--download-media]");
process.exit(1);
}
try {
new URL(args.url);
} catch {
console.error(`Invalid URL: ${args.url}`);
process.exit(1);
}
if (!BROWSER_MODES.has(args.browserMode)) {
console.error(`Invalid --browser mode: ${args.browserMode}. Expected auto, headless, or headed.`);
process.exit(1);
}
if (args.wait && args.browserMode === "headless") {
console.error("Error: --wait requires a visible browser. Use --browser auto or --browser headed.");
process.exit(1);
}
if (args.output) {
const stat = await import("node:fs").then(fs => fs.statSync(args.output!, { throwIfNoEntry: false }));
if (stat?.isDirectory()) {
console.error(`Error: -o path is a directory, not a file: ${args.output}`);
process.exit(1);
}
}
console.log(`Fetching: ${args.url}`);
console.log(`Mode: ${args.wait ? "wait" : "auto"}`);
console.log(`Browser: ${args.browserMode}`);
let outputPath: string;
let htmlSnapshotPath: string | null = null;
let document: string;
let conversionMethod: string;
let fallbackReason: string | undefined;
try {
const result = await captureUrl(args);
document = createMarkdownDocument(result);
outputPath = args.output || await generateOutputPath(result.metadata.url || args.url, result.metadata.title, args.outputDir, document);
const outputDir = path.dirname(outputPath);
htmlSnapshotPath = deriveHtmlSnapshotPath(outputPath);
await mkdir(outputDir, { recursive: true });
await writeFile(htmlSnapshotPath, result.rawHtml, "utf-8");
conversionMethod = result.conversionMethod;
fallbackReason = result.fallbackReason;
} catch (error) {
const primaryError = error instanceof Error ? error.message : String(error);
console.warn(`Primary capture failed: ${primaryError}`);
console.warn("Trying defuddle.md API fallback...");
try {
const remoteResult = await fetchDefuddleApiMarkdown(args.url);
document = remoteResult.markdown;
outputPath = args.output || await generateOutputPath(args.url, remoteResult.title, args.outputDir, document);
await mkdir(path.dirname(outputPath), { recursive: true });
conversionMethod = "defuddle-api";
fallbackReason = `Local browser capture failed: ${primaryError}`;
} catch (remoteError) {
const remoteMessage = remoteError instanceof Error ? remoteError.message : String(remoteError);
throw new Error(`Local browser capture failed (${primaryError}); defuddle.md fallback failed (${remoteMessage})`);
}
}
if (args.downloadMedia) {
const mediaResult = await localizeMarkdownMedia(document, {
markdownPath: outputPath,
log: console.log,
});
document = mediaResult.markdown;
if (mediaResult.downloadedImages > 0 || mediaResult.downloadedVideos > 0) {
console.log(`Downloaded: ${mediaResult.downloadedImages} images, ${mediaResult.downloadedVideos} videos`);
}
} else {
const { images, videos } = countRemoteMedia(document);
if (images > 0 || videos > 0) {
console.log(`Remote media found: ${images} images, ${videos} videos`);
}
}
await writeFile(outputPath, document, "utf-8");
console.log(`Saved: ${outputPath}`);
if (htmlSnapshotPath) {
console.log(`Saved HTML: ${htmlSnapshotPath}`);
} else {
console.log("Saved HTML: unavailable (defuddle.md fallback)");
}
console.log(`Title: ${extractTitleFromMarkdownDocument(document) || "(no title)"}`);
console.log(`Converter: ${conversionMethod}`);
if (fallbackReason) {
console.warn(`Fallback used: ${fallbackReason}`);
}
}
main().catch((err) => {
console.error("Error:", err instanceof Error ? err.message : String(err));
process.exit(1);
});
@@ -1,323 +0,0 @@
import { parseHTML } from "linkedom";
export interface PageMetadata {
url: string;
title: string;
description?: string;
author?: string;
published?: string;
coverImage?: string;
language?: string;
captured_at: string;
}
export interface ConversionResult {
metadata: PageMetadata;
markdown: string;
rawHtml: string;
conversionMethod: string;
fallbackReason?: string;
variables?: Record<string, string>;
}
export type AnyRecord = Record<string, unknown>;
export const MIN_CONTENT_LENGTH = 120;
export const GOOD_CONTENT_LENGTH = 900;
const PUBLISHED_TIME_SELECTORS = [
"meta[property='article:published_time']",
"meta[name='pubdate']",
"meta[name='publishdate']",
"meta[name='date']",
"time[datetime]",
];
const ARTICLE_TYPES = new Set([
"Article",
"NewsArticle",
"BlogPosting",
"WebPage",
"ReportageNewsArticle",
]);
export function pickString(...values: unknown[]): string | null {
for (const value of values) {
if (typeof value === "string") {
const trimmed = value.trim();
if (trimmed) return trimmed;
}
}
return null;
}
export function normalizeMarkdown(markdown: string): string {
return markdown
.replace(/\r\n/g, "\n")
.replace(/[ \t]+\n/g, "\n")
.replace(/\n{3,}/g, "\n\n")
.trim();
}
export function parseDocument(html: string): Document {
const normalized = /<\s*html[\s>]/i.test(html)
? html
: `<!doctype html><html><body>${html}</body></html>`;
return parseHTML(normalized).document as unknown as Document;
}
export function sanitizeHtml(html: string): string {
const { document } = parseHTML(`<div id="__root">${html}</div>`);
const root = document.querySelector("#__root");
if (!root) return html;
for (const selector of ["script", "style", "iframe", "noscript", "template", "svg", "path"]) {
for (const el of root.querySelectorAll(selector)) {
el.remove();
}
}
return root.innerHTML;
}
export function extractTextFromHtml(html: string): string {
const { document } = parseHTML(`<!doctype html><html><body>${html}</body></html>`);
for (const selector of ["script", "style", "noscript", "template", "iframe", "svg", "path"]) {
for (const el of document.querySelectorAll(selector)) {
el.remove();
}
}
return document.body?.textContent?.replace(/\s+/g, " ").trim() ?? "";
}
export function getMetaContent(document: Document, names: string[]): string | null {
for (const name of names) {
const element =
document.querySelector(`meta[name="${name}"]`) ??
document.querySelector(`meta[property="${name}"]`);
const content = element?.getAttribute("content");
if (content && content.trim()) return content.trim();
}
return null;
}
function normalizeLanguageTag(value: string | null): string | null {
if (!value) return null;
const trimmed = value.trim();
if (!trimmed) return null;
const primary = trimmed.split(/[,\s;]/, 1)[0]?.trim();
if (!primary) return null;
return primary.replace(/_/g, "-");
}
function flattenJsonLdItems(data: unknown): AnyRecord[] {
if (!data || typeof data !== "object") return [];
if (Array.isArray(data)) return data.flatMap(flattenJsonLdItems);
const item = data as AnyRecord;
if (Array.isArray(item["@graph"])) {
return (item["@graph"] as unknown[]).flatMap(flattenJsonLdItems);
}
return [item];
}
function parseJsonLdScripts(document: Document): AnyRecord[] {
const results: AnyRecord[] = [];
const scripts = document.querySelectorAll("script[type='application/ld+json']");
for (const script of scripts) {
try {
const data = JSON.parse(script.textContent ?? "");
results.push(...flattenJsonLdItems(data));
} catch {
// Ignore malformed blocks.
}
}
return results;
}
function isArticleType(item: AnyRecord): boolean {
const value = Array.isArray(item["@type"]) ? item["@type"][0] : item["@type"];
return typeof value === "string" && ARTICLE_TYPES.has(value);
}
function extractAuthorFromJsonLd(authorData: unknown): string | null {
if (typeof authorData === "string") return authorData;
if (!authorData || typeof authorData !== "object") return null;
if (Array.isArray(authorData)) {
const names = authorData
.map((author) => extractAuthorFromJsonLd(author))
.filter((name): name is string => Boolean(name));
return names.length > 0 ? names.join(", ") : null;
}
const author = authorData as AnyRecord;
return typeof author.name === "string" ? author.name : null;
}
function extractPrimaryJsonLdMeta(document: Document): Partial<PageMetadata> {
for (const item of parseJsonLdScripts(document)) {
if (!isArticleType(item)) continue;
return {
title: pickString(item.headline, item.name) ?? undefined,
description: pickString(item.description) ?? undefined,
author: extractAuthorFromJsonLd(item.author) ?? undefined,
published: pickString(item.datePublished, item.dateCreated) ?? undefined,
coverImage:
pickString(
item.image,
(item.image as AnyRecord | undefined)?.url,
(Array.isArray(item.image) ? item.image[0] : undefined) as unknown
) ?? undefined,
};
}
return {};
}
export function extractPublishedTime(document: Document): string | null {
for (const selector of PUBLISHED_TIME_SELECTORS) {
const el = document.querySelector(selector);
if (!el) continue;
const value = el.getAttribute("content") ?? el.getAttribute("datetime");
if (value && value.trim()) return value.trim();
}
return null;
}
export function extractTitle(document: Document): string | null {
const ogTitle = document.querySelector("meta[property='og:title']")?.getAttribute("content");
if (ogTitle && ogTitle.trim()) return ogTitle.trim();
const twitterTitle = document.querySelector("meta[name='twitter:title']")?.getAttribute("content");
if (twitterTitle && twitterTitle.trim()) return twitterTitle.trim();
const title = document.querySelector("title")?.textContent?.trim();
if (title) {
const cleaned = title.split(/\s*[-|–—]\s*/)[0]?.trim();
if (cleaned) return cleaned;
}
const h1 = document.querySelector("h1")?.textContent?.trim();
return h1 || null;
}
export function extractMetadataFromHtml(html: string, url: string, capturedAt: string): PageMetadata {
const document = parseDocument(html);
const jsonLd = extractPrimaryJsonLdMeta(document);
const timeEl = document.querySelector("time[datetime]");
const htmlLang = normalizeLanguageTag(document.documentElement?.getAttribute("lang"));
const metaLanguage = normalizeLanguageTag(
pickString(
getMetaContent(document, ["language", "content-language", "og:locale"]),
document.querySelector("meta[http-equiv='content-language']")?.getAttribute("content")
)
);
return {
url,
title:
pickString(
getMetaContent(document, ["og:title", "twitter:title"]),
jsonLd.title,
document.querySelector("h1")?.textContent,
document.title
) ?? "",
description:
pickString(
getMetaContent(document, ["description", "og:description", "twitter:description"]),
jsonLd.description
) ?? undefined,
author:
pickString(
getMetaContent(document, ["author", "article:author", "twitter:creator"]),
jsonLd.author
) ?? undefined,
published:
pickString(
timeEl?.getAttribute("datetime"),
getMetaContent(document, ["article:published_time", "datePublished", "publishdate", "date"]),
jsonLd.published,
extractPublishedTime(document)
) ?? undefined,
coverImage:
pickString(
getMetaContent(document, ["og:image", "twitter:image", "twitter:image:src"]),
jsonLd.coverImage
) ?? undefined,
language: pickString(htmlLang, metaLanguage) ?? undefined,
captured_at: capturedAt,
};
}
export function isMarkdownUsable(markdown: string, html: string): boolean {
const normalized = normalizeMarkdown(markdown);
if (!normalized) return false;
const htmlTextLength = extractTextFromHtml(html).length;
if (htmlTextLength < MIN_CONTENT_LENGTH) return true;
if (normalized.length >= 80) return true;
return normalized.length >= Math.min(200, Math.floor(htmlTextLength * 0.2));
}
export function isYouTubeUrl(url: string): boolean {
try {
const hostname = new URL(url).hostname.toLowerCase();
return hostname === "youtu.be" || hostname.endsWith(".youtube.com") || hostname === "youtube.com";
} catch {
return false;
}
}
function escapeYamlValue(value: string): string {
return value.replace(/\\/g, "\\\\").replace(/"/g, '\\"').replace(/\r?\n/g, "\\n");
}
export function formatMetadataYaml(meta: PageMetadata): string {
const lines = ["---"];
lines.push(`url: ${meta.url}`);
lines.push(`title: "${escapeYamlValue(meta.title)}"`);
if (meta.description) lines.push(`description: "${escapeYamlValue(meta.description)}"`);
if (meta.author) lines.push(`author: "${escapeYamlValue(meta.author)}"`);
if (meta.published) lines.push(`published: "${escapeYamlValue(meta.published)}"`);
if (meta.coverImage) lines.push(`coverImage: "${escapeYamlValue(meta.coverImage)}"`);
if (meta.language) lines.push(`language: "${escapeYamlValue(meta.language)}"`);
lines.push(`captured_at: "${escapeYamlValue(meta.captured_at)}"`);
lines.push("---");
return lines.join("\n");
}
export function createMarkdownDocument(result: ConversionResult): string {
const yaml = formatMetadataYaml(result.metadata);
const escapedTitle = result.metadata.title.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
const titleRegex = new RegExp(`^#\\s+${escapedTitle}\\s*(\\n|$)`, "i");
const hasTitle = titleRegex.test(result.markdown.trimStart());
const firstMeaningfulLine = result.markdown
.replace(/\r\n/g, "\n")
.split("\n")
.map((line) => line.trim())
.find((line) => line && !/^!?\[[^\]]*\]\([^)]+\)$/.test(line))
?.replace(/^>\s*/, "")
?.replace(/^#+\s+/, "")
?.trim();
const comparableTitle = result.metadata.title.toLowerCase().replace(/(?:\.{3}|…)\s*$/, "");
const comparableFirstLine = firstMeaningfulLine?.toLowerCase() ?? "";
const titleRepeatsContent =
comparableTitle !== "" &&
comparableFirstLine !== "" &&
(comparableFirstLine === comparableTitle ||
comparableFirstLine.startsWith(comparableTitle) ||
comparableTitle.startsWith(comparableFirstLine));
const title = result.metadata.title && !hasTitle && !titleRepeatsContent
? `\n\n# ${result.metadata.title}\n\n`
: "\n\n";
return yaml + title + result.markdown;
}
@@ -1,40 +0,0 @@
import assert from "node:assert/strict";
import { mkdtemp, readFile, readdir } from "node:fs/promises";
import os from "node:os";
import path from "node:path";
import test from "node:test";
import { localizeMarkdownMedia } from "./media-localizer.js";
const PNG_1X1_BASE64 =
"iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mP8/x8AAwMCAO7Z0ioAAAAASUVORK5CYII=";
test("localizeMarkdownMedia saves embedded base64 images into imgs directory", async () => {
const tempDir = await mkdtemp(path.join(os.tmpdir(), "url-to-markdown-media-"));
const dataUri = `data:image/png;base64,${PNG_1X1_BASE64}`;
const markdown = [
"---",
`coverImage: "${dataUri}"`,
"---",
"",
"# Embedded Image",
"",
`![inline](${dataUri})`,
"",
].join("\n");
const result = await localizeMarkdownMedia(markdown, {
markdownPath: path.join(tempDir, "post.md"),
});
assert.equal(result.downloadedImages, 1);
assert.equal(result.downloadedVideos, 0);
assert.match(result.markdown, /coverImage: "imgs\/img-001\.png"/);
assert.match(result.markdown, /!\[inline\]\(imgs\/img-001\.png\)/);
const files = await readdir(path.join(tempDir, "imgs"));
assert.deepEqual(files, ["img-001.png"]);
const bytes = await readFile(path.join(tempDir, "imgs", "img-001.png"));
assert.equal(bytes.length, Buffer.from(PNG_1X1_BASE64, "base64").length);
});
@@ -1,374 +0,0 @@
import path from "node:path";
import { mkdir, writeFile } from "node:fs/promises";
type MediaKind = "image" | "video";
type MediaHint = "image" | "unknown";
type MediaSource = "remote" | "data";
type MarkdownLinkCandidate = {
url: string;
hint: MediaHint;
source: MediaSource;
};
export type LocalizeMarkdownMediaOptions = {
markdownPath: string;
log?: (message: string) => void;
};
export type LocalizeMarkdownMediaResult = {
markdown: string;
downloadedImages: number;
downloadedVideos: number;
imageDir: string | null;
videoDir: string | null;
};
const MARKDOWN_LINK_RE =
/(!?\[[^\]\n]*\])\((<)?((?:https?:\/\/[^)\s>]+)|(?:data:[^)>\s]+))(>)?\)/g;
const FRONTMATTER_COVER_RE = /^(coverImage:\s*")((?:https?:\/\/[^"]+)|(?:data:[^"]+))(")/m;
const IMAGE_EXTENSIONS = new Set([
"jpg",
"jpeg",
"png",
"webp",
"gif",
"bmp",
"avif",
"heic",
"heif",
"svg",
]);
const VIDEO_EXTENSIONS = new Set(["mp4", "m4v", "mov", "webm", "mkv"]);
const MIME_EXTENSION_MAP: Record<string, string> = {
"image/jpeg": "jpg",
"image/jpg": "jpg",
"image/png": "png",
"image/webp": "webp",
"image/gif": "gif",
"image/bmp": "bmp",
"image/avif": "avif",
"image/heic": "heic",
"image/heif": "heif",
"image/svg+xml": "svg",
"video/mp4": "mp4",
"video/webm": "webm",
"video/quicktime": "mov",
"video/x-m4v": "m4v",
};
const DOWNLOAD_USER_AGENT =
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36";
function normalizeContentType(raw: string | null): string {
return raw?.split(";")[0]?.trim().toLowerCase() ?? "";
}
function normalizeExtension(raw: string | undefined | null): string | undefined {
if (!raw) return undefined;
const trimmed = raw.replace(/^\./, "").trim().toLowerCase();
if (!trimmed) return undefined;
if (trimmed === "jpeg") return "jpg";
if (trimmed === "jpg") return "jpg";
return trimmed;
}
function resolveExtensionFromUrl(rawUrl: string): string | undefined {
try {
const parsed = new URL(rawUrl);
const extFromPath = normalizeExtension(path.posix.extname(parsed.pathname));
if (extFromPath) return extFromPath;
const extFromFormat = normalizeExtension(parsed.searchParams.get("format"));
if (extFromFormat) return extFromFormat;
} catch {
return undefined;
}
return undefined;
}
function resolveExtensionFromContentType(contentType: string): string | undefined {
return normalizeExtension(MIME_EXTENSION_MAP[contentType]);
}
function resolveKindFromContentType(contentType: string): MediaKind | undefined {
if (!contentType) return undefined;
if (contentType.startsWith("image/")) return "image";
if (contentType.startsWith("video/")) return "video";
return undefined;
}
function resolveKindFromExtension(ext: string | undefined): MediaKind | undefined {
if (!ext) return undefined;
if (IMAGE_EXTENSIONS.has(ext)) return "image";
if (VIDEO_EXTENSIONS.has(ext)) return "video";
return undefined;
}
function resolveMediaKind(
rawUrl: string,
contentType: string,
extension: string | undefined,
hint: MediaHint
): MediaKind | undefined {
const kindFromType = resolveKindFromContentType(contentType);
if (kindFromType) return kindFromType;
const kindFromExtension = resolveKindFromExtension(extension);
if (kindFromExtension) return kindFromExtension;
if (contentType && contentType !== "application/octet-stream") {
return undefined;
}
return hint === "image" ? "image" : undefined;
}
function resolveOutputExtension(
contentType: string,
extension: string | undefined,
kind: MediaKind
): string {
const extFromMime = resolveExtensionFromContentType(contentType);
if (extFromMime) return extFromMime;
const normalizedExt = normalizeExtension(extension);
if (normalizedExt) return normalizedExt;
return kind === "video" ? "mp4" : "jpg";
}
function safeDecodeURIComponent(value: string): string {
try {
return decodeURIComponent(value);
} catch {
return value;
}
}
function sanitizeFileSegment(input: string): string {
return input
.replace(/[^a-zA-Z0-9_-]+/g, "-")
.replace(/-+/g, "-")
.replace(/^[-_]+|[-_]+$/g, "")
.slice(0, 48);
}
function resolveFileStem(rawUrl: string, extension: string): string {
if (isDataUri(rawUrl)) {
return "";
}
try {
const parsed = new URL(rawUrl);
const base = path.posix.basename(parsed.pathname);
if (!base) return "";
const decodedBase = safeDecodeURIComponent(base);
const normalizedExt = normalizeExtension(extension);
const stripExt = normalizedExt ? new RegExp(`\\.${normalizedExt}$`, "i") : null;
const rawStem = stripExt ? decodedBase.replace(stripExt, "") : decodedBase;
return sanitizeFileSegment(rawStem);
} catch {
return "";
}
}
function buildFileName(kind: MediaKind, index: number, sourceUrl: string, extension: string): string {
const stem = resolveFileStem(sourceUrl, extension);
const prefix = kind === "image" ? "img" : "video";
const serial = String(index).padStart(3, "0");
const suffix = stem ? `-${stem}` : "";
return `${prefix}-${serial}${suffix}.${extension}`;
}
function isDataUri(value: string): boolean {
return value.startsWith("data:");
}
function parseBase64DataUri(rawUrl: string): { contentType: string; bytes: Buffer } | null {
const match = rawUrl.match(/^data:([^;,]+);base64,([A-Za-z0-9+/=\s]+)$/i);
if (!match?.[1] || !match[2]) return null;
const contentType = normalizeContentType(match[1]);
if (!contentType) return null;
try {
const bytes = Buffer.from(match[2].replace(/\s+/g, ""), "base64");
if (bytes.length === 0) return null;
return { contentType, bytes };
} catch {
return null;
}
}
function collectMarkdownLinkCandidates(markdown: string): MarkdownLinkCandidate[] {
const candidates: MarkdownLinkCandidate[] = [];
const seen = new Set<string>();
const fmMatch = markdown.match(/^---\n([\s\S]*?)\n---/);
if (fmMatch) {
const coverMatch = fmMatch[1]?.match(FRONTMATTER_COVER_RE);
if (coverMatch?.[2] && !seen.has(coverMatch[2])) {
seen.add(coverMatch[2]);
candidates.push({
url: coverMatch[2],
hint: "image",
source: isDataUri(coverMatch[2]) ? "data" : "remote",
});
}
}
MARKDOWN_LINK_RE.lastIndex = 0;
let match: RegExpExecArray | null;
while ((match = MARKDOWN_LINK_RE.exec(markdown))) {
const label = match[1] ?? "";
const rawUrl = match[3] ?? "";
if (!rawUrl || seen.has(rawUrl)) continue;
seen.add(rawUrl);
candidates.push({
url: rawUrl,
hint: label.startsWith("![") ? "image" : "unknown",
source: isDataUri(rawUrl) ? "data" : "remote",
});
}
return candidates;
}
function rewriteMarkdownMediaLinks(markdown: string, replacements: Map<string, string>): string {
if (replacements.size === 0) return markdown;
MARKDOWN_LINK_RE.lastIndex = 0;
let result = markdown.replace(MARKDOWN_LINK_RE, (full, label, _openAngle, rawUrl) => {
const localPath = replacements.get(rawUrl);
if (!localPath) return full;
return `${label}(${localPath})`;
});
result = result.replace(FRONTMATTER_COVER_RE, (full, prefix, rawUrl, suffix) => {
const localPath = replacements.get(rawUrl);
if (!localPath) return full;
return `${prefix}${localPath}${suffix}`;
});
return result;
}
export async function localizeMarkdownMedia(
markdown: string,
options: LocalizeMarkdownMediaOptions
): Promise<LocalizeMarkdownMediaResult> {
const log = options.log ?? (() => {});
const markdownDir = path.dirname(options.markdownPath);
const candidates = collectMarkdownLinkCandidates(markdown);
if (candidates.length === 0) {
return {
markdown,
downloadedImages: 0,
downloadedVideos: 0,
imageDir: null,
videoDir: null,
};
}
const replacements = new Map<string, string>();
let downloadedImages = 0;
let downloadedVideos = 0;
for (const candidate of candidates) {
try {
let sourceUrl = candidate.url;
let contentType = "";
let extension: string | undefined;
let kind: MediaKind | undefined;
let bytes: Buffer | null = null;
if (candidate.source === "data") {
const parsed = parseBase64DataUri(candidate.url);
if (!parsed) {
log("[url-to-markdown] Skip embedded media: unsupported or invalid data URI");
continue;
}
contentType = parsed.contentType;
extension = resolveExtensionFromContentType(contentType);
kind = resolveMediaKind(sourceUrl, contentType, extension, candidate.hint);
bytes = parsed.bytes;
} else {
const response = await fetch(candidate.url, {
method: "GET",
redirect: "follow",
headers: {
"user-agent": DOWNLOAD_USER_AGENT,
},
});
if (!response.ok) {
log(`[url-to-markdown] Skip media (${response.status}): ${candidate.url}`);
continue;
}
sourceUrl = response.url || candidate.url;
contentType = normalizeContentType(response.headers.get("content-type"));
extension = resolveExtensionFromUrl(sourceUrl) ?? resolveExtensionFromUrl(candidate.url);
kind = resolveMediaKind(sourceUrl, contentType, extension, candidate.hint);
bytes = Buffer.from(await response.arrayBuffer());
}
if (!kind || !bytes) {
continue;
}
const outputExtension = resolveOutputExtension(contentType, extension, kind);
const nextIndex = kind === "image" ? downloadedImages + 1 : downloadedVideos + 1;
const dirName = kind === "image" ? "imgs" : "videos";
const targetDir = path.join(markdownDir, dirName);
await mkdir(targetDir, { recursive: true });
const fileName = buildFileName(kind, nextIndex, sourceUrl, outputExtension);
const absolutePath = path.join(targetDir, fileName);
const relativePath = path.posix.join(dirName, fileName);
await writeFile(absolutePath, bytes);
replacements.set(candidate.url, relativePath);
if (kind === "image") {
downloadedImages = nextIndex;
} else {
downloadedVideos = nextIndex;
}
} catch (error) {
const message = error instanceof Error ? error.message : String(error ?? "");
log(`[url-to-markdown] Failed to download media ${candidate.url}: ${message}`);
}
}
return {
markdown: rewriteMarkdownMediaLinks(markdown, replacements),
downloadedImages,
downloadedVideos,
imageDir: downloadedImages > 0 ? path.join(markdownDir, "imgs") : null,
videoDir: downloadedVideos > 0 ? path.join(markdownDir, "videos") : null,
};
}
export function countRemoteMedia(markdown: string): { images: number; videos: number; hasCoverImage: boolean } {
const fmMatch = markdown.match(/^---\n([\s\S]*?)\n---/);
const hasCoverImage = !!(fmMatch?.[1]?.match(FRONTMATTER_COVER_RE)?.[2]);
const candidates = collectMarkdownLinkCandidates(markdown);
let images = 0;
let videos = 0;
for (const c of candidates) {
if (c.source !== "remote") continue;
const ext = resolveExtensionFromUrl(c.url);
const kind = resolveKindFromExtension(ext);
if (kind === "video") {
videos++;
} else if (kind === "image" || c.hint === "image") {
images++;
}
}
return { images, videos, hasCoverImage };
}
@@ -3,12 +3,6 @@
"private": true, "private": true,
"type": "module", "type": "module",
"dependencies": { "dependencies": {
"@mozilla/readability": "^0.6.0", "baoyu-fetch": "file:./vendor/baoyu-fetch"
"baoyu-chrome-cdp": "file:./vendor/baoyu-chrome-cdp",
"defuddle": "^0.14.0",
"jsdom": "^24.1.3",
"linkedom": "^0.18.12",
"turndown": "^7.2.2",
"turndown-plugin-gfm": "^1.0.2"
} }
} }
@@ -1,206 +0,0 @@
import assert from "node:assert/strict";
import test from "node:test";
import {
createMarkdownDocument,
extractMetadataFromHtml,
} from "../markdown-conversion-shared.js";
import { tryUrlRuleParsers } from "./index.js";
const CAPTURED_AT = "2026-03-22T06:00:00.000Z";
const ARTICLE_HTML = `<!doctype html>
<html lang="zh-CN">
<body>
<div data-testid="twitterArticleReadView">
<a href="/dotey/article/2035141635713941927/media/1">
<div data-testid="tweetPhoto">
<img src="https://pbs.twimg.com/media/article-cover.jpg" alt="Image">
</div>
</a>
<div data-testid="twitter-article-title">Karpathy"写代码"已经不是对的动词了</div>
<div data-testid="User-Name">
<a href="/dotey">宝玉 Verified account</a>
<a href="/dotey">@dotey</a>
<time datetime="2026-03-20T23:49:11.000Z">Mar 20</time>
</div>
<div data-testid="twitterArticleRichTextView">
<p>Andrej Karpathy 说他从 2024 年 12 月起就基本没手写过一行代码。</p>
<a href="/dotey/article/2035141635713941927/media/2">
<div>
<div>
<div data-testid="tweetPhoto">
<img src="https://pbs.twimg.com/media/article-inline.jpg" alt="Image">
</div>
</div>
</div>
</a>
<h2>要点速览</h2>
<ul>
<li>核心焦虑从 GPU 利用率转向 Token 吞吐量</li>
</ul>
<blockquote>
<p>写代码已经不是对的动词了。</p>
</blockquote>
</div>
</div>
</body>
</html>`;
const STATUS_HTML = `<!doctype html>
<html lang="en">
<body>
<article data-testid="tweet">
<div data-testid="User-Name">
<a href="/dotey">宝玉 Verified account</a>
<a href="/dotey">@dotey</a>
<time datetime="2026-03-22T05:33:00.000Z">Mar 22</time>
</div>
<div data-testid="tweetText">
<span>转译:把下面这段加到你的 Codex 自定义指令里,体验会好太多:</span>
</div>
<div data-testid="tweetPhoto">
<img src="https://pbs.twimg.com/media/tweet-main.jpg" alt="Image">
</div>
<div data-testid="User-Name">
<a href="/mattshumer_">Matt Shumer Verified account</a>
<a href="/mattshumer_">@mattshumer_</a>
<time datetime="2026-03-17T00:00:00.000Z">Mar 17</time>
</div>
<div data-testid="tweetText">
<span>Add this to your Codex custom instructions for a way better experience.</span>
</div>
</article>
</body>
</html>`;
const ARCHIVE_HTML = `<!doctype html>
<html>
<head>
<title>archive.ph</title>
</head>
<body>
<form>
<input
type="text"
name="q"
value="https://www.newscientist.com/article/2520204-major-leap-towards-reanimation-after-death-as-mammals-brain-preserved/"
>
</form>
<div id="HEADER">
Archive shell text that should be ignored when CONTENT exists.
</div>
<div id="CONTENT">
<h1>Major leap towards reanimation after death as mammal brain preserved</h1>
<p>
Researchers say the preserved structure and activity markers suggest a significant step
forward in keeping delicate brain tissue viable after clinical death.
</p>
<p>
The archive wrapper should not take precedence over the actual article body when the
CONTENT container is available for parsing.
</p>
<img src="https://cdn.example.com/brain.jpg" alt="Brain tissue">
</div>
</body>
</html>`;
const ARCHIVE_FALLBACK_HTML = `<!doctype html>
<html>
<head>
<title>archive.ph</title>
</head>
<body>
<input type="text" name="q" value="https://example.com/fallback-story">
<main>
<h1>Fallback body parsing still works</h1>
<p>
When CONTENT is absent, the parser should fall back to the body content instead of
returning null or keeping the archive wrapper as the final URL.
</p>
<p>
This ensures archived pages with slightly different layouts still produce usable markdown.
</p>
</main>
</body>
</html>`;
function parse(html: string, url: string) {
const baseMetadata = extractMetadataFromHtml(html, url, CAPTURED_AT);
return tryUrlRuleParsers(html, url, baseMetadata);
}
test("parses archive.ph pages from CONTENT and restores the original URL", () => {
const result = parse(ARCHIVE_HTML, "https://archive.ph/SMcX5");
assert.ok(result);
assert.equal(result.conversionMethod, "parser:archive-ph");
assert.equal(
result.metadata.url,
"https://www.newscientist.com/article/2520204-major-leap-towards-reanimation-after-death-as-mammals-brain-preserved/"
);
assert.equal(
result.metadata.title,
"Major leap towards reanimation after death as mammal brain preserved"
);
assert.equal(result.metadata.coverImage, "https://cdn.example.com/brain.jpg");
assert.ok(result.markdown.includes("Researchers say the preserved structure"));
assert.ok(result.markdown.includes("![Brain tissue](https://cdn.example.com/brain.jpg)"));
assert.ok(!result.markdown.includes("Archive shell text that should be ignored"));
});
test("falls back to body when archive.ph CONTENT is missing", () => {
const result = parse(ARCHIVE_FALLBACK_HTML, "https://archive.ph/fallback");
assert.ok(result);
assert.equal(result.conversionMethod, "parser:archive-ph");
assert.equal(result.metadata.url, "https://example.com/fallback-story");
assert.equal(result.metadata.title, "Fallback body parsing still works");
assert.ok(result.markdown.includes("When CONTENT is absent"));
});
test("parses X article pages from HTML", () => {
const result = parse(
ARTICLE_HTML,
"https://x.com/dotey/article/2035141635713941927"
);
assert.ok(result);
assert.equal(result.conversionMethod, "parser:x-article");
assert.equal(result.metadata.title, "Karpathy\"写代码\"已经不是对的动词了");
assert.equal(result.metadata.author, "宝玉 (@dotey)");
assert.equal(result.metadata.coverImage, "https://pbs.twimg.com/media/article-cover.jpg");
assert.equal(result.metadata.published, "2026-03-20T23:49:11.000Z");
assert.equal(result.metadata.language, "zh");
assert.ok(result.markdown.includes("## 要点速览"));
assert.ok(
result.markdown.includes(
"[![](https://pbs.twimg.com/media/article-inline.jpg)](/dotey/article/2035141635713941927/media/2)"
)
);
assert.ok(result.markdown.includes("写代码已经不是对的动词了。"));
const document = createMarkdownDocument(result);
assert.ok(document.includes("# Karpathy\"写代码\"已经不是对的动词了"));
});
test("parses X status pages from HTML without duplicating the title heading", () => {
const result = parse(
STATUS_HTML,
"https://x.com/dotey/status/2035590649081196710"
);
assert.ok(result);
assert.equal(result.conversionMethod, "parser:x-status");
assert.equal(result.metadata.author, "宝玉 (@dotey)");
assert.equal(result.metadata.coverImage, "https://pbs.twimg.com/media/tweet-main.jpg");
assert.equal(result.metadata.language, "zh");
assert.ok(result.markdown.includes("转译:把下面这段加到你的 Codex 自定义指令里"));
assert.ok(result.markdown.includes("> Quote from Matt Shumer (@mattshumer_)"));
assert.ok(result.markdown.includes("!["));
const document = createMarkdownDocument(result);
assert.ok(
!document.includes("\n\n# 转译:把下面这段加到你的 Codex 自定义指令里,体验会好太多:\n\n")
);
});
@@ -1,47 +0,0 @@
import {
isMarkdownUsable,
normalizeMarkdown,
parseDocument,
type ConversionResult,
type PageMetadata,
} from "../markdown-conversion-shared.js";
import { URL_RULE_PARSERS } from "./rules/index.js";
import type { UrlRuleParserContext } from "./types.js";
export type { UrlRuleParser, UrlRuleParserContext } from "./types.js";
export function tryUrlRuleParsers(
html: string,
url: string,
baseMetadata: PageMetadata
): ConversionResult | null {
const document = parseDocument(html);
const context: UrlRuleParserContext = {
html,
url,
document,
baseMetadata,
};
for (const parser of URL_RULE_PARSERS) {
if (!parser.supports(context)) continue;
try {
const result = parser.parse(context);
if (!result) continue;
const markdown = normalizeMarkdown(result.markdown);
if (!isMarkdownUsable(markdown, html)) continue;
return {
...result,
markdown,
};
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
console.warn(`[url-to-markdown] parser ${parser.id} failed: ${message}`);
}
}
return null;
}
@@ -1,97 +0,0 @@
import { convertHtmlFragmentToMarkdown } from "../../legacy-converter.js";
import {
normalizeMarkdown,
pickString,
type ConversionResult,
} from "../../markdown-conversion-shared.js";
import type { UrlRuleParser, UrlRuleParserContext } from "../types.js";
const ARCHIVE_HOSTS = new Set([
"archive.ph",
"archive.is",
"archive.today",
"archive.md",
"archive.vn",
"archive.li",
"archive.fo",
]);
function isArchiveHost(url: string): boolean {
try {
return ARCHIVE_HOSTS.has(new URL(url).hostname.toLowerCase());
} catch {
return false;
}
}
function readOriginalUrl(document: Document): string | undefined {
const value = document.querySelector("input[name='q']")?.getAttribute("value")?.trim();
if (!value) return undefined;
try {
return new URL(value).href;
} catch {
return undefined;
}
}
function summarize(text: string, maxLength: number): string | undefined {
const normalized = text.replace(/\s+/g, " ").trim();
if (!normalized) return undefined;
if (normalized.length <= maxLength) return normalized;
return `${normalized.slice(0, Math.max(0, maxLength - 1)).trimEnd()}`;
}
function pickContentRoot(document: Document): Element | null {
return (
document.querySelector("#CONTENT") ??
document.querySelector("#content") ??
document.body
);
}
function pickContentTitle(root: Element, fallbackTitle: string): string {
const contentTitle = pickString(
root.querySelector("h1")?.textContent,
root.querySelector("[itemprop='headline']")?.textContent,
root.querySelector("article h2")?.textContent
);
if (contentTitle) return contentTitle;
if (fallbackTitle && !/^archive\./i.test(fallbackTitle.trim())) return fallbackTitle;
return "";
}
function parseArchivePage(context: UrlRuleParserContext): ConversionResult | null {
const root = pickContentRoot(context.document);
if (!root) return null;
const markdown = normalizeMarkdown(convertHtmlFragmentToMarkdown(root.innerHTML));
if (!markdown) return null;
const originalUrl = readOriginalUrl(context.document) ?? context.baseMetadata.url;
const bodyText = root.textContent?.replace(/\s+/g, " ").trim() ?? "";
const published = root.querySelector("time[datetime]")?.getAttribute("datetime") ?? undefined;
const coverImage = root.querySelector("img[src]")?.getAttribute("src") ?? undefined;
return {
metadata: {
...context.baseMetadata,
url: originalUrl,
title: pickContentTitle(root, context.baseMetadata.title),
description: summarize(bodyText, 220) ?? context.baseMetadata.description,
published: pickString(published, context.baseMetadata.published) ?? undefined,
coverImage: pickString(coverImage, context.baseMetadata.coverImage) ?? undefined,
},
markdown,
rawHtml: context.html,
conversionMethod: "parser:archive-ph",
};
}
export const archivePhRuleParser: UrlRuleParser = {
id: "archive-ph",
supports(context) {
return isArchiveHost(context.url);
},
parse: parseArchivePage,
};
@@ -1,10 +0,0 @@
import { archivePhRuleParser } from "./archive-ph.js";
import { xArticleRuleParser } from "./x-article.js";
import { xStatusRuleParser } from "./x-status.js";
import type { UrlRuleParser } from "../types.js";
export const URL_RULE_PARSERS: UrlRuleParser[] = [
archivePhRuleParser,
xArticleRuleParser,
xStatusRuleParser,
];
@@ -1,137 +0,0 @@
import {
normalizeMarkdown,
pickString,
type ConversionResult,
} from "../../markdown-conversion-shared.js";
import type { UrlRuleParser, UrlRuleParserContext } from "../types.js";
import {
cleanText,
collectMediaMarkdown,
convertXRichTextElementToMarkdown,
extractPublishedForCurrentUrl,
inferLanguage,
isXArticlePath,
isXHost,
normalizeXMarkdown,
parseUrl,
pickFirstValidLinkText,
sanitizeCoverImage,
summarizeText,
} from "./x-shared.js";
function collectArticleMarkdown(root: Element): { markdown: string; mediaUrls: string[] } {
const parts: string[] = [];
const seenMedia = new Set<string>();
const mediaUrls: string[] = [];
function pushPart(value: string): void {
const normalized = normalizeMarkdown(value);
if (!normalized) return;
parts.push(normalized);
}
function walk(node: Element): void {
const testId = node.getAttribute("data-testid");
if (testId === "twitterArticleRichTextView" || testId === "longformRichTextComponent") {
const bodyMedia = collectMediaMarkdown(node, seenMedia);
mediaUrls.push(...bodyMedia.urls.filter((url) => !mediaUrls.includes(url)));
pushPart(convertXRichTextElementToMarkdown(node));
return;
}
if (testId === "tweetPhoto") {
const media = collectMediaMarkdown(node, seenMedia);
mediaUrls.push(...media.urls.filter((url) => !mediaUrls.includes(url)));
for (const line of media.lines) pushPart(line);
return;
}
if (
testId === "twitter-article-title" ||
testId === "User-Name" ||
testId === "Tweet-User-Avatar" ||
testId === "reply" ||
testId === "retweet" ||
testId === "like" ||
testId === "bookmark" ||
testId === "caret" ||
testId === "app-text-transition-container"
) {
return;
}
if (node.tagName === "TIME" || node.tagName === "BUTTON") {
return;
}
for (const child of Array.from(node.children)) {
walk(child);
}
}
for (const child of Array.from(root.children)) {
walk(child);
}
return {
markdown: normalizeXMarkdown(parts.join("\n\n")),
mediaUrls,
};
}
function parseXArticle(context: UrlRuleParserContext): ConversionResult | null {
const articleRoot = context.document.querySelector("[data-testid='twitterArticleReadView']") as Element | null;
if (!articleRoot) return null;
const title = cleanText(
context.document.querySelector("[data-testid='twitter-article-title']")?.textContent
);
const identity = pickFirstValidLinkText(
context.document.querySelector("[data-testid='User-Name']")
);
const published = extractPublishedForCurrentUrl(articleRoot, context.url);
const { markdown, mediaUrls } = collectArticleMarkdown(articleRoot);
if (!markdown) return null;
const bodyText = cleanText(
context.document.querySelector("[data-testid='twitterArticleRichTextView']")?.textContent ??
context.document.querySelector("[data-testid='longformRichTextComponent']")?.textContent
);
return {
metadata: {
...context.baseMetadata,
title: pickString(title, context.baseMetadata.title) ?? "",
description: summarizeText(bodyText, 220) ?? context.baseMetadata.description,
author: pickString(identity.author, context.baseMetadata.author) ?? undefined,
published: pickString(published, context.baseMetadata.published) ?? undefined,
coverImage: sanitizeCoverImage(mediaUrls[0], context.baseMetadata.coverImage),
language: inferLanguage(bodyText, context.baseMetadata.language),
},
markdown,
rawHtml: context.html,
conversionMethod: "parser:x-article",
};
}
export const xArticleRuleParser: UrlRuleParser = {
id: "x-article",
supports(context) {
const parsed = parseUrl(context.url);
if (!parsed || !isXHost(parsed.hostname)) {
return false;
}
return (
isXArticlePath(parsed.pathname) ||
Boolean(
context.document.querySelector("[data-testid='twitterArticleReadView']") ||
context.document.querySelector("[data-testid='twitterArticleRichTextView']")
)
);
},
parse(context) {
return parseXArticle(context);
},
};
@@ -1,249 +0,0 @@
import { convertHtmlFragmentToMarkdown } from "../../legacy-converter.js";
import { normalizeMarkdown } from "../../markdown-conversion-shared.js";
export const DEFAULT_X_OG_IMAGE = "https://abs.twimg.com/rweb/ssr/default/v2/og/image.png";
export type MediaResult = {
lines: string[];
urls: string[];
};
export function isXHost(hostname: string): boolean {
const normalized = hostname.toLowerCase();
return (
normalized === "x.com" ||
normalized === "twitter.com" ||
normalized.endsWith(".x.com") ||
normalized.endsWith(".twitter.com")
);
}
export function parseUrl(input: string): URL | null {
try {
return new URL(input);
} catch {
return null;
}
}
export function isXStatusPath(pathname: string): boolean {
return /^\/[^/]+\/status(?:es)?\/\d+$/i.test(pathname) || /^\/i\/web\/status\/\d+$/i.test(pathname);
}
export function isXArticlePath(pathname: string): boolean {
return /^\/[^/]+\/article\/\d+$/i.test(pathname) || /^\/(?:i\/)?article\/\d+$/i.test(pathname);
}
export function cleanText(value: string | null | undefined): string {
return (value ?? "").replace(/\s+/g, " ").trim();
}
export function cleanUserLabel(value: string | null | undefined): string {
return cleanText(value).replace(/\bVerified account\b/gi, "").replace(/\s{2,}/g, " ").trim();
}
export function escapeMarkdownAlt(text: string): string {
return text.replace(/[\[\]]/g, "\\$&");
}
export function normalizeAlt(text: string | null | undefined): string {
const cleaned = cleanText(text);
if (!cleaned || /^(image|photo)$/i.test(cleaned)) return "";
return escapeMarkdownAlt(cleaned);
}
export function summarizeText(text: string, maxLength: number): string | undefined {
const normalized = cleanText(text);
if (!normalized) return undefined;
return normalized.length > maxLength
? `${normalized.slice(0, maxLength - 3)}...`
: normalized;
}
export function buildTweetTitle(text: string, fallback: string): string {
return summarizeText(text, 80) ?? fallback;
}
export function normalizeXMarkdown(markdown: string): string {
return normalizeMarkdown(markdown.replace(/^(#{1,6})\s*\n+([^\n])/gm, "$1 $2"));
}
export function inferLanguage(text: string, fallback?: string): string | undefined {
const normalized = cleanText(text);
if (!normalized) return fallback;
const han = (normalized.match(/\p{Script=Han}/gu) || []).length;
const hiragana = (normalized.match(/\p{Script=Hiragana}/gu) || []).length;
const katakana = (normalized.match(/\p{Script=Katakana}/gu) || []).length;
const hangul = (normalized.match(/\p{Script=Hangul}/gu) || []).length;
if (hangul >= 8) return "ko";
if (hiragana + katakana >= 8) return "ja";
if (han >= 16) return "zh";
return fallback;
}
export function buildQuoteMarkdown(markdown: string, author?: string): string {
const normalized = normalizeMarkdown(markdown);
if (!normalized) return "";
const lines = normalized.split("\n");
const prefixed = lines.map((line) => (line ? `> ${line}` : ">")).join("\n");
const header = author ? `> Quote from ${author}` : "> Quote";
return `${header}\n${prefixed}`;
}
export function pickFirstValidLinkText(userNameEl: Element | null | undefined): {
name?: string;
username?: string;
author?: string;
} {
if (!userNameEl) return {};
const linkTexts = Array.from(userNameEl.querySelectorAll("a[href]"))
.map((link) => cleanUserLabel(link.textContent))
.filter(Boolean);
let username = linkTexts.find((text) => text.startsWith("@"));
let name = linkTexts.find((text) => !text.startsWith("@") && !/^(promote|more)$/i.test(text));
if (!username || !name) {
const text = cleanUserLabel(userNameEl.textContent);
const fallbackMatch = text.match(/^(.*?)\s*(@[A-Za-z0-9_]+)(?:\s*·.*)?$/);
if (fallbackMatch) {
name = name ?? cleanText(fallbackMatch[1]);
username = username ?? cleanText(fallbackMatch[2]);
}
}
const author = name && username ? `${name} (${username})` : username ?? name;
return { name, username, author };
}
export function extractPublishedForCurrentUrl(root: ParentNode, url: string): string | undefined {
const parsed = parseUrl(url);
if (!parsed) return undefined;
const currentPath = parsed.pathname.toLowerCase();
for (const timeElement of root.querySelectorAll("a[href] time[datetime]")) {
const href = timeElement.closest("a")?.getAttribute("href");
const hrefUrl = href ? parseUrl(href.startsWith("http") ? href : `${parsed.origin}${href}`) : null;
if (hrefUrl?.pathname.toLowerCase() === currentPath) {
return timeElement.getAttribute("datetime") ?? undefined;
}
}
return root.querySelector("time[datetime]")?.getAttribute("datetime") ?? undefined;
}
export function collectMediaMarkdown(root: ParentNode, seen: Set<string>): MediaResult {
const lines: string[] = [];
const urls: string[] = [];
const rootElement = root as Element & {
getAttribute?: (name: string) => string | null;
};
const photoNodes = [
...(typeof rootElement.getAttribute === "function" &&
rootElement.getAttribute("data-testid") === "tweetPhoto"
? [rootElement]
: []),
...Array.from(root.querySelectorAll("[data-testid='tweetPhoto']")),
];
for (const node of photoNodes) {
const img = node.querySelector("img");
const imageUrl = img?.getAttribute("src");
if (imageUrl && !seen.has(imageUrl)) {
seen.add(imageUrl);
urls.push(imageUrl);
lines.push(`![${normalizeAlt(img?.getAttribute("alt"))}](${imageUrl})`);
}
const video = node.querySelector("video");
const posterUrl = video?.getAttribute("poster");
if (posterUrl && !seen.has(posterUrl)) {
seen.add(posterUrl);
urls.push(posterUrl);
lines.push(`![video](${posterUrl})`);
}
const videoUrl = video?.getAttribute("src") ?? video?.querySelector("source")?.getAttribute("src");
if (videoUrl && !seen.has(videoUrl)) {
seen.add(videoUrl);
urls.push(videoUrl);
lines.push(`[video](${videoUrl})`);
}
}
return { lines, urls };
}
export function materializeTweetPhotoNodes(root: Element): void {
for (const photo of Array.from(root.querySelectorAll("[data-testid='tweetPhoto']"))) {
const document = photo.ownerDocument;
const container = document.createElement("span");
const img = photo.querySelector("img");
const imageUrl = img?.getAttribute("src");
if (imageUrl) {
const image = document.createElement("img");
image.setAttribute("src", imageUrl);
const alt = normalizeAlt(img?.getAttribute("alt"));
if (alt) {
image.setAttribute("alt", alt);
}
container.appendChild(image);
}
const video = photo.querySelector("video");
const posterUrl = video?.getAttribute("poster");
if (posterUrl) {
const poster = document.createElement("img");
poster.setAttribute("src", posterUrl);
poster.setAttribute("alt", "video");
container.appendChild(poster);
}
const videoUrl = video?.getAttribute("src") ?? video?.querySelector("source")?.getAttribute("src");
if (videoUrl) {
if (container.childNodes.length > 0) {
container.appendChild(document.createTextNode(" "));
}
const link = document.createElement("a");
link.setAttribute("href", videoUrl);
link.textContent = "video";
container.appendChild(link);
}
if (container.childNodes.length === 0) {
photo.remove();
continue;
}
photo.replaceWith(container);
}
}
function collapseLinkedMediaContainers(root: Element): void {
for (const anchor of Array.from(root.querySelectorAll("a[href]"))) {
const images = Array.from(anchor.querySelectorAll("img"));
if (images.length !== 1) continue;
if (cleanText(anchor.textContent)) continue;
const image = images[0].cloneNode(true);
anchor.replaceChildren(image);
}
}
export function convertXRichTextElementToMarkdown(node: Element): string {
const clone = node.cloneNode(true) as Element;
materializeTweetPhotoNodes(clone);
collapseLinkedMediaContainers(clone);
return normalizeXMarkdown(convertHtmlFragmentToMarkdown(clone.innerHTML));
}
export function sanitizeCoverImage(primary?: string, fallback?: string): string | undefined {
if (primary) return primary;
if (!fallback || fallback === DEFAULT_X_OG_IMAGE) return undefined;
return fallback;
}
@@ -1,82 +0,0 @@
import type { ConversionResult } from "../../markdown-conversion-shared.js";
import type { UrlRuleParser, UrlRuleParserContext } from "../types.js";
import {
buildQuoteMarkdown,
buildTweetTitle,
cleanText,
collectMediaMarkdown,
convertXRichTextElementToMarkdown,
extractPublishedForCurrentUrl,
inferLanguage,
isXHost,
isXStatusPath,
normalizeXMarkdown,
parseUrl,
pickFirstValidLinkText,
sanitizeCoverImage,
summarizeText,
} from "./x-shared.js";
function parseXStatus(context: UrlRuleParserContext): ConversionResult | null {
const article = context.document.querySelector("article[data-testid='tweet'], article") as Element | null;
if (!article) return null;
const tweetTextElements = Array.from(article.querySelectorAll("[data-testid='tweetText']")) as Element[];
if (tweetTextElements.length === 0) return null;
const userNameElements = Array.from(article.querySelectorAll("[data-testid='User-Name']")) as Element[];
const mainTextElement = tweetTextElements[0];
const mainIdentity = pickFirstValidLinkText(userNameElements[0]);
const published = extractPublishedForCurrentUrl(article, context.url);
const mainMarkdown = normalizeXMarkdown(convertXRichTextElementToMarkdown(mainTextElement));
if (!mainMarkdown) return null;
const parts = [mainMarkdown];
const quotedTextElements = tweetTextElements.slice(1);
const quotedUserNameElements = userNameElements.slice(1);
quotedTextElements.forEach((element, index) => {
const quoteMarkdown = normalizeXMarkdown(convertXRichTextElementToMarkdown(element));
if (!quoteMarkdown) return;
const quoteIdentity = pickFirstValidLinkText(quotedUserNameElements[index]);
parts.push(buildQuoteMarkdown(quoteMarkdown, quoteIdentity.author));
});
const media = collectMediaMarkdown(article, new Set<string>());
if (media.lines.length > 0) {
parts.push(media.lines.join("\n\n"));
}
const mainText = cleanText(mainTextElement.textContent);
const markdown = normalizeXMarkdown(parts.join("\n\n"));
return {
metadata: {
...context.baseMetadata,
title: buildTweetTitle(mainText, context.baseMetadata.title),
description: summarizeText(mainText, 220) ?? context.baseMetadata.description,
author: mainIdentity.author ?? context.baseMetadata.author,
published: published ?? context.baseMetadata.published,
coverImage: sanitizeCoverImage(media.urls[0], context.baseMetadata.coverImage),
language: inferLanguage(mainText, context.baseMetadata.language),
},
markdown,
rawHtml: context.html,
conversionMethod: "parser:x-status",
};
}
export const xStatusRuleParser: UrlRuleParser = {
id: "x-status",
supports(context) {
const parsed = parseUrl(context.url);
if (!parsed || !isXHost(parsed.hostname)) {
return false;
}
return isXStatusPath(parsed.pathname) && Boolean(context.document.querySelector("[data-testid='tweetText']"));
},
parse(context): ConversionResult | null {
return parseXStatus(context);
},
};
@@ -1,14 +0,0 @@
import type { ConversionResult, PageMetadata } from "../markdown-conversion-shared.js";
export interface UrlRuleParserContext {
html: string;
url: string;
document: Document;
baseMetadata: PageMetadata;
}
export interface UrlRuleParser {
id: string;
supports(context: UrlRuleParserContext): boolean;
parse(context: UrlRuleParserContext): ConversionResult | null;
}
@@ -1,29 +0,0 @@
import os from "node:os";
import path from "node:path";
import process from "node:process";
const APP_DATA_DIR = "baoyu-skills";
const URL_TO_MARKDOWN_DATA_DIR = "url-to-markdown";
const PROFILE_DIR_NAME = "chrome-profile";
export function resolveUserDataRoot(): string {
if (process.platform === "win32") {
return process.env.APPDATA ?? path.join(os.homedir(), "AppData", "Roaming");
}
if (process.platform === "darwin") {
return path.join(os.homedir(), "Library", "Application Support");
}
return process.env.XDG_DATA_HOME ?? path.join(os.homedir(), ".local", "share");
}
export function resolveUrlToMarkdownDataDir(): string {
const override = process.env.URL_DATA_DIR?.trim();
if (override) return path.resolve(override);
return path.join(process.cwd(), URL_TO_MARKDOWN_DATA_DIR);
}
export function resolveUrlToMarkdownChromeProfileDir(): string {
const override = process.env.BAOYU_CHROME_PROFILE_DIR?.trim() || process.env.URL_CHROME_PROFILE_DIR?.trim();
if (override) return path.resolve(override);
return path.join(resolveUserDataRoot(), APP_DATA_DIR, PROFILE_DIR_NAME);
}
@@ -1,9 +0,0 @@
{
"name": "baoyu-chrome-cdp",
"private": true,
"version": "0.1.0",
"type": "module",
"exports": {
".": "./src/index.ts"
}
}
@@ -1,307 +0,0 @@
import assert from "node:assert/strict";
import { spawn, type ChildProcess } from "node:child_process";
import fs from "node:fs/promises";
import http from "node:http";
import os from "node:os";
import path from "node:path";
import process from "node:process";
import test, { type TestContext } from "node:test";
import {
discoverRunningChromeDebugPort,
findChromeExecutable,
findExistingChromeDebugPort,
getFreePort,
openPageSession,
resolveSharedChromeProfileDir,
waitForChromeDebugPort,
} from "./index.ts";
function useEnv(
t: TestContext,
values: Record<string, string | null>,
): void {
const previous = new Map<string, string | undefined>();
for (const [key, value] of Object.entries(values)) {
previous.set(key, process.env[key]);
if (value == null) {
delete process.env[key];
} else {
process.env[key] = value;
}
}
t.after(() => {
for (const [key, value] of previous.entries()) {
if (value == null) {
delete process.env[key];
} else {
process.env[key] = value;
}
}
});
}
async function makeTempDir(prefix: string): Promise<string> {
return fs.mkdtemp(path.join(os.tmpdir(), prefix));
}
async function startDebugServer(port: number): Promise<http.Server> {
const server = http.createServer((req, res) => {
if (req.url === "/json/version") {
res.writeHead(200, { "Content-Type": "application/json" });
res.end(JSON.stringify({
webSocketDebuggerUrl: `ws://127.0.0.1:${port}/devtools/browser/demo`,
}));
return;
}
res.writeHead(404);
res.end();
});
await new Promise<void>((resolve, reject) => {
server.once("error", reject);
server.listen(port, "127.0.0.1", () => resolve());
});
return server;
}
async function closeServer(server: http.Server): Promise<void> {
await new Promise<void>((resolve, reject) => {
server.close((error) => {
if (error) reject(error);
else resolve();
});
});
}
function shellPathForPlatform(): string | null {
if (process.platform === "win32") return null;
return "/bin/bash";
}
async function startFakeChromiumProcess(port: number): Promise<ChildProcess | null> {
const shell = shellPathForPlatform();
if (!shell) return null;
const child = spawn(
shell,
[
"-lc",
`exec -a chromium-mock ${JSON.stringify(process.execPath)} -e 'setInterval(() => {}, 1000)' -- --remote-debugging-port=${port}`,
],
{ stdio: "ignore" },
);
await new Promise((resolve) => setTimeout(resolve, 250));
return child;
}
async function stopProcess(child: ChildProcess | null): Promise<void> {
if (!child) return;
if (child.exitCode !== null || child.signalCode !== null) return;
child.kill("SIGTERM");
await new Promise((resolve) => setTimeout(resolve, 100));
if (child.exitCode === null && child.signalCode === null) child.kill("SIGKILL");
if (child.exitCode !== null || child.signalCode !== null) return;
await new Promise((resolve) => child.once("exit", resolve));
}
test("getFreePort honors a fixed environment override and otherwise allocates a TCP port", async (t) => {
useEnv(t, { TEST_FIXED_PORT: "45678" });
assert.equal(await getFreePort("TEST_FIXED_PORT"), 45678);
const dynamicPort = await getFreePort();
assert.ok(Number.isInteger(dynamicPort));
assert.ok(dynamicPort > 0);
});
test("findChromeExecutable prefers env overrides and falls back to candidate paths", async (t) => {
const root = await makeTempDir("baoyu-chrome-bin-");
t.after(() => fs.rm(root, { recursive: true, force: true }));
const envChrome = path.join(root, "env-chrome");
const fallbackChrome = path.join(root, "fallback-chrome");
await fs.writeFile(envChrome, "");
await fs.writeFile(fallbackChrome, "");
useEnv(t, { BAOYU_CHROME_PATH: envChrome });
assert.equal(
findChromeExecutable({
envNames: ["BAOYU_CHROME_PATH"],
candidates: { default: [fallbackChrome] },
}),
envChrome,
);
useEnv(t, { BAOYU_CHROME_PATH: null });
assert.equal(
findChromeExecutable({
envNames: ["BAOYU_CHROME_PATH"],
candidates: { default: [fallbackChrome] },
}),
fallbackChrome,
);
});
test("resolveSharedChromeProfileDir supports env overrides, WSL paths, and default suffixes", (t) => {
useEnv(t, { BAOYU_SHARED_PROFILE: "/tmp/custom-profile" });
assert.equal(
resolveSharedChromeProfileDir({
envNames: ["BAOYU_SHARED_PROFILE"],
appDataDirName: "demo-app",
profileDirName: "demo-profile",
}),
path.resolve("/tmp/custom-profile"),
);
useEnv(t, { BAOYU_SHARED_PROFILE: null });
assert.equal(
resolveSharedChromeProfileDir({
wslWindowsHome: "/mnt/c/Users/demo",
appDataDirName: "demo-app",
profileDirName: "demo-profile",
}),
path.join("/mnt/c/Users/demo", ".local", "share", "demo-app", "demo-profile"),
);
const fallback = resolveSharedChromeProfileDir({
appDataDirName: "demo-app",
profileDirName: "demo-profile",
});
assert.match(fallback, /demo-app[\\/]demo-profile$/);
});
test("findExistingChromeDebugPort reads DevToolsActivePort and validates it against a live endpoint", async (t) => {
const root = await makeTempDir("baoyu-cdp-profile-");
t.after(() => fs.rm(root, { recursive: true, force: true }));
const port = await getFreePort();
const server = await startDebugServer(port);
t.after(() => closeServer(server));
await fs.writeFile(path.join(root, "DevToolsActivePort"), `${port}\n/devtools/browser/demo\n`);
const found = await findExistingChromeDebugPort({ profileDir: root, timeoutMs: 1000 });
assert.equal(found, port);
});
test("discoverRunningChromeDebugPort reads DevToolsActivePort from the provided user-data dir", async (t) => {
const root = await makeTempDir("baoyu-cdp-user-data-");
t.after(() => fs.rm(root, { recursive: true, force: true }));
const port = await getFreePort();
const server = await startDebugServer(port);
t.after(() => closeServer(server));
await fs.writeFile(path.join(root, "DevToolsActivePort"), `${port}\n/devtools/browser/demo\n`);
const found = await discoverRunningChromeDebugPort({
userDataDirs: [root],
timeoutMs: 1000,
});
assert.deepEqual(found, {
port,
wsUrl: `ws://127.0.0.1:${port}/devtools/browser/demo`,
});
});
test("discoverRunningChromeDebugPort ignores unrelated debugging processes", async (t) => {
if (process.platform === "win32") {
t.skip("Process discovery fallback is not used on Windows.");
return;
}
const root = await makeTempDir("baoyu-cdp-user-data-");
t.after(() => fs.rm(root, { recursive: true, force: true }));
const port = await getFreePort();
const server = await startDebugServer(port);
t.after(() => closeServer(server));
const fakeChromium = await startFakeChromiumProcess(port);
t.after(async () => { await stopProcess(fakeChromium); });
const found = await discoverRunningChromeDebugPort({
userDataDirs: [root],
timeoutMs: 1000,
});
assert.equal(found, null);
});
test("openPageSession reports whether it created a new target", async () => {
const calls: string[] = [];
const cdpExisting = {
send: async <T>(method: string): Promise<T> => {
calls.push(method);
if (method === "Target.getTargets") {
return {
targetInfos: [{ targetId: "existing-target", type: "page", url: "https://gemini.google.com/app" }],
} as T;
}
if (method === "Target.attachToTarget") return { sessionId: "session-existing" } as T;
throw new Error(`Unexpected method: ${method}`);
},
};
const existing = await openPageSession({
cdp: cdpExisting as never,
reusing: false,
url: "https://gemini.google.com/app",
matchTarget: (target) => target.url.includes("gemini.google.com"),
activateTarget: false,
});
assert.deepEqual(existing, {
sessionId: "session-existing",
targetId: "existing-target",
createdTarget: false,
});
assert.deepEqual(calls, ["Target.getTargets", "Target.attachToTarget"]);
const createCalls: string[] = [];
const cdpCreated = {
send: async <T>(method: string): Promise<T> => {
createCalls.push(method);
if (method === "Target.getTargets") return { targetInfos: [] } as T;
if (method === "Target.createTarget") return { targetId: "created-target" } as T;
if (method === "Target.attachToTarget") return { sessionId: "session-created" } as T;
throw new Error(`Unexpected method: ${method}`);
},
};
const created = await openPageSession({
cdp: cdpCreated as never,
reusing: false,
url: "https://gemini.google.com/app",
matchTarget: (target) => target.url.includes("gemini.google.com"),
activateTarget: false,
});
assert.deepEqual(created, {
sessionId: "session-created",
targetId: "created-target",
createdTarget: true,
});
assert.deepEqual(createCalls, ["Target.getTargets", "Target.createTarget", "Target.attachToTarget"]);
});
test("waitForChromeDebugPort retries until the debug endpoint becomes available", async (t) => {
const port = await getFreePort();
const serverPromise = (async () => {
await new Promise((resolve) => setTimeout(resolve, 200));
const server = await startDebugServer(port);
t.after(() => closeServer(server));
})();
const websocketUrl = await waitForChromeDebugPort(port, 4000, {
includeLastError: true,
});
await serverPromise;
assert.equal(websocketUrl, `ws://127.0.0.1:${port}/devtools/browser/demo`);
});
@@ -1,523 +0,0 @@
import { spawn, spawnSync, type ChildProcess } from "node:child_process";
import fs from "node:fs";
import net from "node:net";
import os from "node:os";
import path from "node:path";
import process from "node:process";
export type PlatformCandidates = {
darwin?: string[];
win32?: string[];
default: string[];
};
type PendingRequest = {
resolve: (value: unknown) => void;
reject: (error: Error) => void;
timer: ReturnType<typeof setTimeout> | null;
};
type CdpSendOptions = {
sessionId?: string;
timeoutMs?: number;
};
type FetchJsonOptions = {
timeoutMs?: number;
};
type FindChromeExecutableOptions = {
candidates: PlatformCandidates;
envNames?: string[];
};
type ResolveSharedChromeProfileDirOptions = {
envNames?: string[];
appDataDirName?: string;
profileDirName?: string;
wslWindowsHome?: string | null;
};
type FindExistingChromeDebugPortOptions = {
profileDir: string;
timeoutMs?: number;
};
export type ChromeChannel = "stable" | "beta" | "canary" | "dev";
export type DiscoveredChrome = {
port: number;
wsUrl: string;
};
type DiscoverRunningChromeOptions = {
channels?: ChromeChannel[];
userDataDirs?: string[];
timeoutMs?: number;
};
type LaunchChromeOptions = {
chromePath: string;
profileDir: string;
port: number;
url?: string;
headless?: boolean;
extraArgs?: string[];
};
type ChromeTargetInfo = {
targetId: string;
url: string;
type: string;
};
type OpenPageSessionOptions = {
cdp: CdpConnection;
reusing: boolean;
url: string;
matchTarget: (target: ChromeTargetInfo) => boolean;
enablePage?: boolean;
enableRuntime?: boolean;
enableDom?: boolean;
enableNetwork?: boolean;
activateTarget?: boolean;
};
export type PageSession = {
sessionId: string;
targetId: string;
createdTarget: boolean;
};
export function sleep(ms: number): Promise<void> {
return new Promise((resolve) => setTimeout(resolve, ms));
}
export async function getFreePort(fixedEnvName?: string): Promise<number> {
const fixed = fixedEnvName ? Number.parseInt(process.env[fixedEnvName] ?? "", 10) : NaN;
if (Number.isInteger(fixed) && fixed > 0) return fixed;
return await new Promise((resolve, reject) => {
const server = net.createServer();
server.unref();
server.on("error", reject);
server.listen(0, "127.0.0.1", () => {
const address = server.address();
if (!address || typeof address === "string") {
server.close(() => reject(new Error("Unable to allocate a free TCP port.")));
return;
}
const port = address.port;
server.close((err) => {
if (err) reject(err);
else resolve(port);
});
});
});
}
export function findChromeExecutable(options: FindChromeExecutableOptions): string | undefined {
for (const envName of options.envNames ?? []) {
const override = process.env[envName]?.trim();
if (override && fs.existsSync(override)) return override;
}
const candidates = process.platform === "darwin"
? options.candidates.darwin ?? options.candidates.default
: process.platform === "win32"
? options.candidates.win32 ?? options.candidates.default
: options.candidates.default;
for (const candidate of candidates) {
if (fs.existsSync(candidate)) return candidate;
}
return undefined;
}
export function resolveSharedChromeProfileDir(options: ResolveSharedChromeProfileDirOptions = {}): string {
for (const envName of options.envNames ?? []) {
const override = process.env[envName]?.trim();
if (override) return path.resolve(override);
}
const appDataDirName = options.appDataDirName ?? "baoyu-skills";
const profileDirName = options.profileDirName ?? "chrome-profile";
if (options.wslWindowsHome) {
return path.join(options.wslWindowsHome, ".local", "share", appDataDirName, profileDirName);
}
const base = process.platform === "darwin"
? path.join(os.homedir(), "Library", "Application Support")
: process.platform === "win32"
? (process.env.APPDATA ?? path.join(os.homedir(), "AppData", "Roaming"))
: (process.env.XDG_DATA_HOME ?? path.join(os.homedir(), ".local", "share"));
return path.join(base, appDataDirName, profileDirName);
}
async function fetchWithTimeout(url: string, timeoutMs?: number): Promise<Response> {
if (!timeoutMs || timeoutMs <= 0) return await fetch(url, { redirect: "follow" });
const ctl = new AbortController();
const timer = setTimeout(() => ctl.abort(), timeoutMs);
try {
return await fetch(url, { redirect: "follow", signal: ctl.signal });
} finally {
clearTimeout(timer);
}
}
async function fetchJson<T = unknown>(url: string, options: FetchJsonOptions = {}): Promise<T> {
const response = await fetchWithTimeout(url, options.timeoutMs);
if (!response.ok) {
throw new Error(`Request failed: ${response.status} ${response.statusText}`);
}
return await response.json() as T;
}
async function isDebugPortReady(port: number, timeoutMs = 3_000): Promise<boolean> {
try {
const version = await fetchJson<{ webSocketDebuggerUrl?: string }>(
`http://127.0.0.1:${port}/json/version`,
{ timeoutMs }
);
return !!version.webSocketDebuggerUrl;
} catch {
return false;
}
}
function isPortListening(port: number, timeoutMs = 3_000): Promise<boolean> {
return new Promise((resolve) => {
const socket = new net.Socket();
const timer = setTimeout(() => { socket.destroy(); resolve(false); }, timeoutMs);
socket.once("connect", () => { clearTimeout(timer); socket.destroy(); resolve(true); });
socket.once("error", () => { clearTimeout(timer); resolve(false); });
socket.connect(port, "127.0.0.1");
});
}
function parseDevToolsActivePort(filePath: string): { port: number; wsPath: string } | null {
try {
const content = fs.readFileSync(filePath, "utf-8");
const lines = content.split(/\r?\n/);
const port = Number.parseInt(lines[0]?.trim() ?? "", 10);
const wsPath = lines[1]?.trim();
if (port > 0 && wsPath) return { port, wsPath };
} catch {}
return null;
}
export async function findExistingChromeDebugPort(options: FindExistingChromeDebugPortOptions): Promise<number | null> {
const timeoutMs = options.timeoutMs ?? 3_000;
const parsed = parseDevToolsActivePort(path.join(options.profileDir, "DevToolsActivePort"));
if (parsed && parsed.port > 0 && await isDebugPortReady(parsed.port, timeoutMs)) return parsed.port;
if (process.platform === "win32") return null;
try {
const result = spawnSync("ps", ["aux"], { encoding: "utf-8", timeout: 5_000 });
if (result.status !== 0 || !result.stdout) return null;
const lines = result.stdout
.split("\n")
.filter((line) => line.includes(options.profileDir) && line.includes("--remote-debugging-port="));
for (const line of lines) {
const portMatch = line.match(/--remote-debugging-port=(\d+)/);
const port = Number.parseInt(portMatch?.[1] ?? "", 10);
if (port > 0 && await isDebugPortReady(port, timeoutMs)) return port;
}
} catch {}
return null;
}
export function getDefaultChromeUserDataDirs(channels: ChromeChannel[] = ["stable"]): string[] {
const home = os.homedir();
const dirs: string[] = [];
const channelDirs: Record<string, { darwin: string; linux: string; win32: string }> = {
stable: {
darwin: path.join(home, "Library", "Application Support", "Google", "Chrome"),
linux: path.join(home, ".config", "google-chrome"),
win32: path.join(process.env.LOCALAPPDATA ?? path.join(home, "AppData", "Local"), "Google", "Chrome", "User Data"),
},
beta: {
darwin: path.join(home, "Library", "Application Support", "Google", "Chrome Beta"),
linux: path.join(home, ".config", "google-chrome-beta"),
win32: path.join(process.env.LOCALAPPDATA ?? path.join(home, "AppData", "Local"), "Google", "Chrome Beta", "User Data"),
},
canary: {
darwin: path.join(home, "Library", "Application Support", "Google", "Chrome Canary"),
linux: path.join(home, ".config", "google-chrome-canary"),
win32: path.join(process.env.LOCALAPPDATA ?? path.join(home, "AppData", "Local"), "Google", "Chrome SxS", "User Data"),
},
dev: {
darwin: path.join(home, "Library", "Application Support", "Google", "Chrome Dev"),
linux: path.join(home, ".config", "google-chrome-dev"),
win32: path.join(process.env.LOCALAPPDATA ?? path.join(home, "AppData", "Local"), "Google", "Chrome Dev", "User Data"),
},
};
const platform = process.platform === "darwin" ? "darwin" : process.platform === "win32" ? "win32" : "linux";
for (const ch of channels) {
const entry = channelDirs[ch];
if (entry) dirs.push(entry[platform]);
}
return dirs;
}
// Best-effort reuse of an already-running local CDP session discovered from
// known Chrome user-data dirs. This is distinct from Chrome DevTools MCP's
// prompt-based --autoConnect flow.
export async function discoverRunningChromeDebugPort(options: DiscoverRunningChromeOptions = {}): Promise<DiscoveredChrome | null> {
const channels = options.channels ?? ["stable", "beta", "canary", "dev"];
const timeoutMs = options.timeoutMs ?? 3_000;
const userDataDirs = (options.userDataDirs ?? getDefaultChromeUserDataDirs(channels))
.map((dir) => path.resolve(dir));
for (const dir of userDataDirs) {
const parsed = parseDevToolsActivePort(path.join(dir, "DevToolsActivePort"));
if (!parsed) continue;
if (await isPortListening(parsed.port, timeoutMs)) {
return { port: parsed.port, wsUrl: `ws://127.0.0.1:${parsed.port}${parsed.wsPath}` };
}
}
if (process.platform !== "win32") {
try {
const result = spawnSync("ps", ["aux"], { encoding: "utf-8", timeout: 5_000 });
if (result.status === 0 && result.stdout) {
const lines = result.stdout
.split("\n")
.filter((line) =>
line.includes("--remote-debugging-port=") &&
userDataDirs.some((dir) => line.includes(dir))
);
for (const line of lines) {
const portMatch = line.match(/--remote-debugging-port=(\d+)/);
const port = Number.parseInt(portMatch?.[1] ?? "", 10);
if (port > 0 && await isDebugPortReady(port, timeoutMs)) {
try {
const version = await fetchJson<{ webSocketDebuggerUrl?: string }>(`http://127.0.0.1:${port}/json/version`, { timeoutMs });
if (version.webSocketDebuggerUrl) return { port, wsUrl: version.webSocketDebuggerUrl };
} catch {}
}
}
}
} catch {}
}
return null;
}
export async function waitForChromeDebugPort(
port: number,
timeoutMs: number,
options?: { includeLastError?: boolean }
): Promise<string> {
const start = Date.now();
let lastError: unknown = null;
while (Date.now() - start < timeoutMs) {
try {
const version = await fetchJson<{ webSocketDebuggerUrl?: string }>(
`http://127.0.0.1:${port}/json/version`,
{ timeoutMs: 5_000 }
);
if (version.webSocketDebuggerUrl) return version.webSocketDebuggerUrl;
lastError = new Error("Missing webSocketDebuggerUrl");
} catch (error) {
lastError = error;
}
await sleep(200);
}
if (options?.includeLastError && lastError) {
throw new Error(
`Chrome debug port not ready: ${lastError instanceof Error ? lastError.message : String(lastError)}`
);
}
throw new Error("Chrome debug port not ready");
}
export class CdpConnection {
private ws: WebSocket;
private nextId = 0;
private pending = new Map<number, PendingRequest>();
private eventHandlers = new Map<string, Set<(params: unknown) => void>>();
private defaultTimeoutMs: number;
private constructor(ws: WebSocket, defaultTimeoutMs = 15_000) {
this.ws = ws;
this.defaultTimeoutMs = defaultTimeoutMs;
this.ws.addEventListener("message", (event) => {
try {
const data = typeof event.data === "string"
? event.data
: new TextDecoder().decode(event.data as ArrayBuffer);
const msg = JSON.parse(data) as {
id?: number;
method?: string;
params?: unknown;
result?: unknown;
error?: { message?: string };
};
if (msg.method) {
const handlers = this.eventHandlers.get(msg.method);
if (handlers) {
handlers.forEach((handler) => handler(msg.params));
}
}
if (msg.id) {
const pending = this.pending.get(msg.id);
if (pending) {
this.pending.delete(msg.id);
if (pending.timer) clearTimeout(pending.timer);
if (msg.error?.message) pending.reject(new Error(msg.error.message));
else pending.resolve(msg.result);
}
}
} catch {}
});
this.ws.addEventListener("close", () => {
for (const [id, pending] of this.pending.entries()) {
this.pending.delete(id);
if (pending.timer) clearTimeout(pending.timer);
pending.reject(new Error("CDP connection closed."));
}
});
}
static async connect(
url: string,
timeoutMs: number,
options?: { defaultTimeoutMs?: number }
): Promise<CdpConnection> {
const ws = new WebSocket(url);
await new Promise<void>((resolve, reject) => {
const timer = setTimeout(() => reject(new Error("CDP connection timeout.")), timeoutMs);
ws.addEventListener("open", () => {
clearTimeout(timer);
resolve();
});
ws.addEventListener("error", () => {
clearTimeout(timer);
reject(new Error("CDP connection failed."));
});
});
return new CdpConnection(ws, options?.defaultTimeoutMs ?? 15_000);
}
on(method: string, handler: (params: unknown) => void): void {
if (!this.eventHandlers.has(method)) {
this.eventHandlers.set(method, new Set());
}
this.eventHandlers.get(method)?.add(handler);
}
off(method: string, handler: (params: unknown) => void): void {
this.eventHandlers.get(method)?.delete(handler);
}
async send<T = unknown>(method: string, params?: Record<string, unknown>, options?: CdpSendOptions): Promise<T> {
const id = ++this.nextId;
const message: Record<string, unknown> = { id, method };
if (params) message.params = params;
if (options?.sessionId) message.sessionId = options.sessionId;
const timeoutMs = options?.timeoutMs ?? this.defaultTimeoutMs;
const result = await new Promise<unknown>((resolve, reject) => {
const timer = timeoutMs > 0
? setTimeout(() => {
this.pending.delete(id);
reject(new Error(`CDP timeout: ${method}`));
}, timeoutMs)
: null;
this.pending.set(id, { resolve, reject, timer });
this.ws.send(JSON.stringify(message));
});
return result as T;
}
close(): void {
try {
this.ws.close();
} catch {}
}
}
export async function launchChrome(options: LaunchChromeOptions): Promise<ChildProcess> {
await fs.promises.mkdir(options.profileDir, { recursive: true });
const args = [
`--remote-debugging-port=${options.port}`,
`--user-data-dir=${options.profileDir}`,
"--no-first-run",
"--no-default-browser-check",
...(options.extraArgs ?? []),
];
if (options.headless) args.push("--headless=new");
if (options.url) args.push(options.url);
return spawn(options.chromePath, args, { stdio: "ignore" });
}
export function killChrome(chrome: ChildProcess): void {
try {
chrome.kill("SIGTERM");
} catch {}
setTimeout(() => {
if (!chrome.killed) {
try {
chrome.kill("SIGKILL");
} catch {}
}
}, 2_000).unref?.();
}
export async function openPageSession(options: OpenPageSessionOptions): Promise<PageSession> {
let targetId: string;
let createdTarget = false;
if (options.reusing) {
const created = await options.cdp.send<{ targetId: string }>("Target.createTarget", { url: options.url });
targetId = created.targetId;
createdTarget = true;
} else {
const targets = await options.cdp.send<{ targetInfos: ChromeTargetInfo[] }>("Target.getTargets");
const existing = targets.targetInfos.find(options.matchTarget);
if (existing) {
targetId = existing.targetId;
} else {
const created = await options.cdp.send<{ targetId: string }>("Target.createTarget", { url: options.url });
targetId = created.targetId;
createdTarget = true;
}
}
const { sessionId } = await options.cdp.send<{ sessionId: string }>(
"Target.attachToTarget",
{ targetId, flatten: true }
);
if (options.activateTarget ?? true) {
await options.cdp.send("Target.activateTarget", { targetId });
}
if (options.enablePage) await options.cdp.send("Page.enable", {}, { sessionId });
if (options.enableRuntime) await options.cdp.send("Runtime.enable", {}, { sessionId });
if (options.enableDom) await options.cdp.send("DOM.enable", {}, { sessionId });
if (options.enableNetwork) await options.cdp.send("Network.enable", {}, { sessionId });
return { sessionId, targetId, createdTarget };
}
@@ -0,0 +1,124 @@
# baoyu-fetch
English | [简体中文](./README.zh-CN.md) | [Changelog](./CHANGELOG.md) | [中文更新日志](./CHANGELOG.zh-CN.md)
`baoyu-fetch` is a Bun CLI built on Chrome CDP. Give it a URL and it returns
high-quality `markdown` or `json`. When a site adapter matches, it prefers API
responses or structured page data; otherwise it falls back to generic HTML
extraction.
## Features
- Capture rendered page content through Chrome CDP
- Observe network requests and responses, and fetch bodies when needed
- Adapter registry that auto-selects a handler from the URL
- Built-in adapters for `x`, `youtube`, and `hn`
- Generic fallback: Defuddle first, then Readability + HTML-to-Markdown; when `--format markdown` is requested, it can also fall back to `defuddle.md`
- Print `markdown` / `json` to stdout or save with `--output`
- Optionally download extracted images or videos and rewrite Markdown links
- Optional wait modes for login and verification flows
- Chrome profile defaults to `baoyu-skills/chrome-profile`
## Installation
```bash
bun install
```
For package usage, the quickest option is:
```bash
bunx baoyu-fetch https://example.com
```
You can also install it globally:
```bash
npm install -g baoyu-fetch
```
The npm package ships TypeScript source entrypoints instead of a prebuilt
`dist`, so Bun is required at runtime.
## Usage
```bash
bun run src/cli.ts https://example.com
bunx baoyu-fetch https://example.com
baoyu-fetch https://example.com
baoyu-fetch https://example.com --format markdown --output article.md
baoyu-fetch https://example.com --format markdown --output article.md --download-media
baoyu-fetch https://x.com/jack/status/20 --format json --output article.json
baoyu-fetch https://x.com/jack/status/20 --json
baoyu-fetch https://x.com/jack/status/20 --wait-for interaction
baoyu-fetch https://x.com/jack/status/20 --wait-for force
baoyu-fetch https://x.com/jack/status/20 --chrome-profile-dir ~/Library/Application\\ Support/baoyu-skills/chrome-profile
```
## Options
```bash
baoyu-fetch <url> [options]
Options:
--output <file> Save output to file
--format <type> Output format: markdown | json
--json Alias for --format json
--adapter <name> Force an adapter (for example x / hn / generic)
--download-media Download adapter-reported media into ./imgs and ./videos, then rewrite markdown links
--media-dir <dir> Base directory for downloaded media. Defaults to the output directory
--debug-dir <dir> Write debug artifacts (html, document.json, network.json)
--cdp-url <url> Reuse an existing Chrome DevTools endpoint
--browser-path <path> Explicit Chrome binary path
--chrome-profile-dir <path>
Chrome user data dir. Defaults to BAOYU_CHROME_PROFILE_DIR
or baoyu-skills/chrome-profile
--headless Launch a temporary headless Chrome if needed
--wait-for <mode> Wait mode: interaction | force
--wait-for-interaction
Alias for --wait-for interaction
--wait-for-login Alias for --wait-for interaction
--interaction-timeout <ms>
Manual interaction timeout. Default: 600000
--interaction-poll-interval <ms>
Poll interval while waiting. Default: 1500
--login-timeout <ms> Alias for --interaction-timeout
--login-poll-interval <ms>
Alias for --interaction-poll-interval
--timeout <ms> Page load timeout. Default: 30000
--help Show help
```
## How It Works
1. The CLI parses the target URL and options.
2. It opens or connects to a Chrome CDP session and creates a controlled tab.
3. `NetworkJournal` records requests and responses.
4. The adapter registry resolves a site-specific adapter when possible.
5. The adapter returns a structured `ExtractedDocument`.
6. If nothing matches, generic HTML extraction runs instead.
7. The result is rendered as Markdown, or returned as JSON with both
`document` and `markdown`.
## Development
```bash
bun run check
bun run test
bun run build
```
## Release
When you make a user-visible change, add a changeset first:
```bash
bunx changeset
```
After the generated `.changeset/*.md` file lands on `main`, GitHub Actions will
open or update the release PR. Merging that release PR publishes the package to
npm.
The publish flow does not build `dist`; it publishes `src/*.ts` for Bun
execution directly.
@@ -0,0 +1,122 @@
# baoyu-fetch
[English](./README.md) | 简体中文 | [更新日志](./CHANGELOG.zh-CN.md) | [English Changelog](./CHANGELOG.md)
`baoyu-fetch` 是一个基于 Chrome CDP 的 Bun CLI。输入 URL,它会输出高质量
`markdown``json`;命中站点 adapter 时优先消费 API 返回或页面内结构化
数据,未命中时回退到通用 HTML 提取。
## 当前能力
- 通过 Chrome CDP 抓取渲染后的页面内容
- 监听网络请求与响应,按需拉取响应体
- adapter registry,支持按 URL 自动命中站点处理器
- 内置 `x``youtube``hn` adapters
- 通用 fallbackDefuddle 优先,Readability + HTML to Markdown 回退;`--format markdown` 时会再尝试 `defuddle.md` 兜底
- `stdout``--output` 输出 `markdown` / `json`
- 可选下载提取出的图片/视频并重写 Markdown 链接
- 提供登录/验证场景下的交互等待模式
- Chrome profile 默认对齐 `baoyu-skills/chrome-profile`
## 安装
```bash
bun install
```
作为包使用时,推荐直接这样运行:
```bash
bunx baoyu-fetch https://example.com
```
也可以全局安装:
```bash
npm install -g baoyu-fetch
```
npm 包发布的是 TypeScript 源码入口,不包含预编译的 `dist`,所以运行时需要
Bun。
## 用法
```bash
bun run src/cli.ts https://example.com
bunx baoyu-fetch https://example.com
baoyu-fetch https://example.com
baoyu-fetch https://example.com --format markdown --output article.md
baoyu-fetch https://example.com --format markdown --output article.md --download-media
baoyu-fetch https://x.com/jack/status/20 --format json --output article.json
baoyu-fetch https://x.com/jack/status/20 --json
baoyu-fetch https://x.com/jack/status/20 --wait-for interaction
baoyu-fetch https://x.com/jack/status/20 --wait-for force
baoyu-fetch https://x.com/jack/status/20 --chrome-profile-dir ~/Library/Application\\ Support/baoyu-skills/chrome-profile
```
## 主要参数
```bash
baoyu-fetch <url> [options]
Options:
--output <file> 保存输出内容到文件
--format <type> 输出格式:markdown | json
--json `--format json` 的兼容别名
--adapter <name> 强制使用指定 adapter(如 x / hn / generic
--download-media 下载 adapter 返回的媒体到 ./imgs 和 ./videos,并重写 markdown 链接
--media-dir <dir> 指定媒体下载根目录;默认使用输出文件所在目录
--debug-dir <dir> 导出调试信息(html、document.json、network.json
--cdp-url <url> 连接现有 Chrome 调试地址
--browser-path <path> 指定 Chrome 可执行文件
--chrome-profile-dir <path>
指定 Chrome profile 目录。默认使用 BAOYU_CHROME_PROFILE_DIR
否则回退到 baoyu-skills/chrome-profile
--headless 启动临时 headless Chrome(未连现有实例时)
--wait-for <mode> 等待模式:interaction | force
--wait-for-interaction
`--wait-for interaction` 的别名
--wait-for-login `--wait-for interaction` 的别名
--interaction-timeout <ms>
手动交互等待超时,默认 600000
--interaction-poll-interval <ms>
等待期间的轮询间隔,默认 1500
--login-timeout <ms> `--interaction-timeout` 的别名
--login-poll-interval <ms>
`--interaction-poll-interval` 的别名
--timeout <ms> 页面加载超时,默认 30000
--help 显示帮助
```
## 设计
核心链路:
1. CLI 解析 URL 和选项
2. 建立 CDP 会话并创建受控 tab
3. 启动 `NetworkJournal` 收集所有请求/响应
4. 由 adapter registry 匹配站点 adapter
5. adapter 返回结构化 `ExtractedDocument`
6. 没命中则走通用 HTML 提取
7. 按请求输出 Markdown,或输出包含 `document``markdown` 的 JSON
## 开发
```bash
bun run check
bun run test
bun run build
```
## 发版
新增用户可见改动后,先添加一个 changeset:
```bash
bunx changeset
```
把生成的 `.changeset/*.md` 一起合并到 `main` 后,GitHub Actions 会自动创建或
更新 release PR;合并 release PR 之后,会自动发布到 npm。
发布流程不会编译 `dist`,而是直接把 `src/*.ts` 发布出去供 Bun 执行。
@@ -0,0 +1,63 @@
{
"name": "baoyu-fetch",
"version": "0.1.1",
"description": "Read URLs into high-quality Markdown or JSON with Chrome CDP and site adapters.",
"type": "module",
"bin": {
"baoyu-fetch": "./src/cli.ts"
},
"files": [
"README.zh-CN.md",
"src/adapters",
"src/browser",
"src/cli.ts",
"src/commands",
"src/extract",
"src/media",
"src/types",
"src/utils",
"README.md"
],
"repository": {
"type": "git",
"url": "git+https://github.com/JimLiu/baoyu-skills.git",
"directory": "packages/baoyu-fetch"
},
"bugs": {
"url": "https://github.com/JimLiu/baoyu-skills/issues"
},
"homepage": "https://github.com/JimLiu/baoyu-skills/tree/main/packages/baoyu-fetch#readme",
"publishConfig": {
"access": "public"
},
"scripts": {
"build": "rm -rf dist && bun build ./src/cli.ts --target bun --outfile ./dist/cli.js && chmod +x ./dist/cli.js",
"check": "tsc --noEmit",
"dev": "bun run ./src/cli.ts",
"release": "changeset publish",
"test": "bun test",
"version-packages": "changeset version"
},
"engines": {
"bun": ">=1.2.0"
},
"dependencies": {
"@mozilla/readability": "^0.6.0",
"chrome-launcher": "^1.2.1",
"defuddle": "^0.14.0",
"jsdom": "^26.0.0",
"remark-gfm": "^4.0.1",
"remark-parse": "^11.0.0",
"turndown": "^7.2.0",
"turndown-plugin-gfm": "^1.0.2",
"unified": "^11.0.5",
"ws": "^8.18.3"
},
"devDependencies": {
"@changesets/cli": "^2.30.0",
"@types/bun": "^1.2.23",
"@types/jsdom": "^21.1.7",
"@types/ws": "^8.18.1",
"typescript": "^5.9.2"
}
}
@@ -0,0 +1,74 @@
import type { Adapter } from "../types";
import { detectInteractionGate } from "../../browser/interaction-gates";
import { captureNormalizedPageSnapshot } from "../../browser/page-snapshot";
import { convertHtmlToMarkdown } from "../../extract/html-to-markdown";
export const genericAdapter: Adapter = {
name: "generic",
match() {
return true;
},
async process(context) {
context.log.info(`Loading ${context.input.url.toString()} with generic adapter`);
await context.browser.goto(context.input.url.toString(), context.timeoutMs);
try {
await context.network.waitForIdle({
idleMs: 1_200,
timeoutMs: Math.min(context.timeoutMs, 15_000),
});
} catch {
context.log.debug("Network idle timed out on initial load; continuing.");
}
await context.browser.scrollToEnd({ maxSteps: 4, delayMs: 300 });
try {
await context.network.waitForIdle({
idleMs: 900,
timeoutMs: Math.min(context.timeoutMs, 10_000),
});
} catch {
context.log.debug("Network idle timed out after scrolling; continuing.");
}
const interaction = await detectInteractionGate(context.browser);
if (interaction) {
return {
status: "needs_interaction",
interaction,
};
}
const snapshot = await captureNormalizedPageSnapshot(context.browser);
const converted = await convertHtmlToMarkdown(snapshot.html, snapshot.finalUrl, {
enableRemoteMarkdownFallback: context.outputFormat === "markdown",
preserveBase64Images: context.downloadMedia,
});
const document = {
url: snapshot.finalUrl,
canonicalUrl: converted.metadata.canonicalUrl,
title: converted.metadata.title,
author: converted.metadata.author,
siteName: converted.metadata.siteName,
publishedAt: converted.metadata.publishedAt,
summary: converted.metadata.summary,
adapter: "generic",
metadata: {
coverImage: converted.metadata.coverImage,
language: converted.metadata.language,
capturedAt: converted.metadata.capturedAt,
conversionMethod: converted.conversionMethod,
fallbackReason: converted.fallbackReason,
kind: "generic/article",
},
content: converted.markdown ? [{ type: "markdown" as const, markdown: converted.markdown }] : [],
};
return {
status: "ok",
document,
media: converted.media,
};
},
};
@@ -0,0 +1,391 @@
import { JSDOM } from "jsdom";
import TurndownService from "turndown";
import { gfm } from "turndown-plugin-gfm";
import type { Adapter } from "../types";
import type { ExtractedDocument } from "../../extract/document";
import { collectMediaFromDocument } from "../../media/markdown-media";
const HN_BASE_URL = "https://news.ycombinator.com";
const turndown = new TurndownService({
headingStyle: "atx",
bulletListMarker: "-",
codeBlockStyle: "fenced",
});
turndown.use(gfm);
export interface HnItem {
id: number;
type: "story" | "comment" | "job" | "poll" | "pollopt" | string;
by?: string;
time?: number;
text?: string;
title?: string;
url?: string;
score?: number;
descendants?: number;
kids?: number[];
parent?: number;
deleted?: boolean;
dead?: boolean;
}
export interface HnCommentNode {
item: HnItem;
children: HnCommentNode[];
}
interface ParsedHnThread {
story: HnItem;
comments: HnCommentNode[];
}
function decodeHtmlText(value: string | undefined): string | undefined {
if (!value) {
return undefined;
}
const dom = new JSDOM(`<!doctype html><html><body>${value}</body></html>`);
return dom.window.document.body.textContent?.trim() || undefined;
}
function normalizeMarkdown(markdown: string): string {
return markdown
.replace(/\r\n/g, "\n")
.replace(/[ \t]+\n/g, "\n")
.replace(/\n{3,}/g, "\n\n")
.trim();
}
function convertHnHtmlToMarkdown(html: string | undefined, baseUrl: string): string {
if (!html?.trim()) {
return "";
}
const dom = new JSDOM(`<div id="__root">${html}</div>`, { url: baseUrl });
const root = dom.window.document.querySelector("#__root");
if (!root) {
return "";
}
root.querySelectorAll("a[href]").forEach((element) => {
const href = element.getAttribute("href");
if (!href) {
return;
}
try {
element.setAttribute("href", new URL(href, baseUrl).toString());
} catch {
// Ignore malformed URLs and keep the original href.
}
});
return normalizeMarkdown(turndown.turndown(root.innerHTML));
}
function formatIsoTimestamp(unixSeconds: number | undefined): string | undefined {
if (!unixSeconds || !Number.isFinite(unixSeconds)) {
return undefined;
}
return new Date(unixSeconds * 1_000).toISOString();
}
function formatDisplayTimestamp(unixSeconds: number | undefined): string {
const iso = formatIsoTimestamp(unixSeconds);
if (!iso) {
return "unknown time";
}
return iso.replace("T", " ").replace(".000Z", " UTC");
}
function indentMarkdown(markdown: string, spaces: number): string {
const prefix = " ".repeat(spaces);
return markdown
.split("\n")
.map((line) => (line ? `${prefix}${line}` : prefix))
.join("\n");
}
function renderCommentHeader(item: HnItem, pageUrl: string): string {
const author = item.by ?? "[deleted]";
const time = item.id
? `[${formatDisplayTimestamp(item.time)}](${pageUrl}#${item.id})`
: formatDisplayTimestamp(item.time);
return `${author} · ${time}`;
}
function renderCommentNode(node: HnCommentNode, pageUrl: string, depth = 0): string {
const baseIndent = " ".repeat(depth * 4);
const lines = [`${baseIndent}- ${renderCommentHeader(node.item, pageUrl)}`];
const body = convertHnHtmlToMarkdown(node.item.text, pageUrl);
if (body) {
lines.push("");
lines.push(indentMarkdown(body, depth * 4 + 4));
} else if (node.item.deleted || node.item.dead) {
lines.push("");
lines.push(`${baseIndent} [comment unavailable]`);
}
for (const child of node.children) {
lines.push("");
lines.push(renderCommentNode(child, pageUrl, depth + 1));
}
return lines.join("\n");
}
export function buildHnThreadMarkdown(
story: HnItem,
comments: HnCommentNode[],
pageUrl: string,
): string {
const lines: string[] = [];
const storyUrl = story.url ? new URL(story.url, pageUrl).toString() : undefined;
const storyText = convertHnHtmlToMarkdown(story.text, pageUrl);
if (storyUrl && storyUrl !== pageUrl) {
lines.push(`Source: [${storyUrl}](${storyUrl})`);
}
lines.push(`HN Item: [${story.id}](${pageUrl})`);
const submittedBy = story.by ? ` by ${story.by}` : "";
const submittedAt = formatDisplayTimestamp(story.time);
lines.push(`Submitted${submittedBy} at ${submittedAt}`);
const stats: string[] = [];
if (typeof story.score === "number") {
stats.push(`${story.score} points`);
}
if (typeof story.descendants === "number") {
stats.push(`${story.descendants} comments`);
}
if (stats.length > 0) {
lines.push(stats.join(" | "));
}
if (storyText) {
lines.push("");
lines.push("## Post");
lines.push("");
lines.push(storyText);
}
lines.push("");
lines.push("## Comments");
lines.push("");
if (comments.length === 0) {
lines.push("No comments.");
} else {
lines.push(comments.map((comment) => renderCommentNode(comment, pageUrl)).join("\n\n"));
}
return normalizeMarkdown(lines.join("\n"));
}
export function buildHnDocument(
story: HnItem,
comments: HnCommentNode[],
pageUrl: string,
): ExtractedDocument {
const decodedTitle = decodeHtmlText(story.title) ?? `HN Item ${story.id}`;
return {
url: pageUrl,
canonicalUrl: pageUrl,
title: decodedTitle,
author: story.by,
siteName: "Hacker News",
publishedAt: formatIsoTimestamp(story.time),
adapter: "hn",
metadata: {
kind: "hn/story",
storyId: story.id,
storyUrl: story.url ? new URL(story.url, pageUrl).toString() : undefined,
points: story.score,
commentCount: story.descendants,
},
content: [
{
type: "markdown",
markdown: buildHnThreadMarkdown(story, comments, pageUrl),
},
],
};
}
export function parseHnItemId(url: URL): number | null {
if (url.hostname !== "news.ycombinator.com") {
return null;
}
if (url.pathname !== "/item") {
return null;
}
const value = url.searchParams.get("id");
if (!value || !/^\d+$/.test(value)) {
return null;
}
return Number(value);
}
function extractUnixSecondsFromAge(element: Element | null): number | undefined {
const title = element?.getAttribute("title")?.trim();
if (!title) {
return undefined;
}
const match = title.match(/(\d{9,})$/);
return match ? Number(match[1]) : undefined;
}
function extractScore(text: string | null | undefined): number | undefined {
if (!text) {
return undefined;
}
const match = text.match(/(\d+)/);
return match ? Number(match[1]) : undefined;
}
function extractCommentCount(container: ParentNode): number | undefined {
const anchors = Array.from(container.querySelectorAll("a"));
for (const anchor of anchors) {
const match = anchor.textContent?.trim().match(/(\d+)\s+comments?/i);
if (match) {
return Number(match[1]);
}
}
return undefined;
}
function normalizeStoryUrl(storyId: number, href: string | null | undefined, pageUrl: string): string | undefined {
if (!href) {
return undefined;
}
try {
const resolved = new URL(href, pageUrl).toString();
if (resolved === pageUrl || resolved === `${HN_BASE_URL}/item?id=${storyId}`) {
return undefined;
}
return resolved;
} catch {
return undefined;
}
}
export function extractHnThreadFromHtml(html: string, pageUrl: string): ParsedHnThread | null {
const dom = new JSDOM(html, { url: pageUrl });
const { document } = dom.window;
const storyRow = document.querySelector("table.fatitem tr.athing.submission");
if (!storyRow) {
return null;
}
const storyId = Number(storyRow.getAttribute("id"));
if (!Number.isFinite(storyId)) {
return null;
}
const titleLink = storyRow.querySelector(".titleline > a");
const subline = document.querySelector("table.fatitem .subline");
const topText = document.querySelector("table.fatitem .toptext");
const story: HnItem = {
id: storyId,
type: "story",
by: subline?.querySelector(".hnuser")?.textContent?.trim() || undefined,
time: extractUnixSecondsFromAge(subline?.querySelector(".age") ?? null),
title: titleLink?.innerHTML?.trim() || undefined,
url: normalizeStoryUrl(storyId, titleLink?.getAttribute("href"), pageUrl),
text: topText?.innerHTML?.trim() || undefined,
score: extractScore(subline?.querySelector(".score")?.textContent),
descendants: extractCommentCount(subline ?? document),
};
const roots: HnCommentNode[] = [];
const stack: HnCommentNode[] = [];
document.querySelectorAll("tr.athing.comtr").forEach((row) => {
const commentId = Number(row.getAttribute("id"));
if (!Number.isFinite(commentId)) {
return;
}
const indentRaw = row.querySelector("td.ind")?.getAttribute("indent");
const depth = indentRaw && /^\d+$/.test(indentRaw) ? Number(indentRaw) : 0;
const comhead = row.querySelector(".comhead");
const item: HnItem = {
id: commentId,
type: "comment",
by: comhead?.querySelector(".hnuser")?.textContent?.trim() || undefined,
time: extractUnixSecondsFromAge(comhead?.querySelector(".age") ?? null),
text: row.querySelector(".comment > .commtext")?.innerHTML?.trim() || undefined,
deleted: row.querySelector(".comment > .commtext") === null,
};
const node: HnCommentNode = {
item,
children: [],
};
while (stack.length > depth) {
stack.pop();
}
const parent = stack[stack.length - 1];
if (parent) {
parent.children.push(node);
} else {
roots.push(node);
}
stack.push(node);
});
return {
story,
comments: roots,
};
}
export const hnAdapter: Adapter = {
name: "hn",
match(input) {
return parseHnItemId(input.url) !== null;
},
async process(context) {
const itemId = parseHnItemId(context.input.url);
if (!itemId) {
return {
status: "no_document",
};
}
const pageUrl = context.input.url.toString();
context.log.info(`Loading ${pageUrl} with hn adapter`);
await context.browser.goto(pageUrl, context.timeoutMs);
const html = await context.browser.getHTML();
const thread = extractHnThreadFromHtml(html, pageUrl);
if (!thread) {
return {
status: "no_document",
};
}
const document = buildHnDocument(thread.story, thread.comments, pageUrl);
return {
status: "ok",
document,
media: collectMediaFromDocument(document),
};
},
};
@@ -0,0 +1,29 @@
import type { Adapter, AdapterInput } from "./types";
import { genericAdapter } from "./generic";
import { hnAdapter } from "./hn";
import { xAdapter } from "./x";
import { youtubeAdapter } from "./youtube";
const adapters: Adapter[] = [xAdapter, youtubeAdapter, hnAdapter, genericAdapter];
export function listAdapters(): Adapter[] {
return adapters;
}
export function resolveAdapter(input: AdapterInput, forcedName?: string): Adapter {
if (forcedName) {
const forced = adapters.find((adapter) => adapter.name === forcedName);
if (!forced) {
throw new Error(`Unknown adapter: ${forcedName}`);
}
return forced;
}
const matched = adapters.find((adapter) => adapter.match(input));
if (!matched) {
throw new Error("No adapter matched the URL");
}
return matched;
}
export { genericAdapter };
@@ -0,0 +1,71 @@
import type { BrowserSession } from "../browser/session";
import type { CdpClient } from "../browser/cdp-client";
import type { NetworkJournal } from "../browser/network-journal";
import type { ExtractedDocument } from "../extract/document";
import type { MediaDownloadRequest, MediaDownloadResult, MediaAsset } from "../media/types";
import type { Logger } from "../utils/logger";
export interface AdapterInput {
url: URL;
}
export type LoginState = "logged_in" | "logged_out" | "unknown";
export type InteractionKind = "login" | "cloudflare" | "recaptcha" | "hcaptcha" | "captcha" | "challenge";
export interface AdapterLoginInfo {
provider: string;
state: LoginState;
required?: boolean;
username?: string;
reason?: string;
}
export interface WaitForInteractionRequest {
type: "wait_for_interaction";
kind: InteractionKind;
provider: string;
prompt: string;
reason?: string;
timeoutMs?: number;
pollIntervalMs?: number;
requiresVisibleBrowser?: boolean;
}
export type AdapterProcessResult =
| {
status: "ok";
document: ExtractedDocument;
media?: MediaAsset[];
login?: AdapterLoginInfo;
}
| {
status: "needs_interaction";
interaction: WaitForInteractionRequest;
login?: AdapterLoginInfo;
}
| {
status: "no_document";
login?: AdapterLoginInfo;
};
export interface AdapterContext {
input: AdapterInput;
browser: BrowserSession;
network: NetworkJournal;
cdp: CdpClient;
log: Logger;
outputFormat: "markdown" | "json";
timeoutMs: number;
interactive: boolean;
downloadMedia: boolean;
}
export interface Adapter {
name: string;
match(input: AdapterInput): boolean;
checkLogin?(context: AdapterContext): Promise<AdapterLoginInfo>;
downloadMedia?(request: MediaDownloadRequest): Promise<MediaDownloadResult>;
process(context: AdapterContext): Promise<AdapterProcessResult>;
}
export type { MediaAsset };
@@ -0,0 +1,433 @@
import type { ExtractedDocument } from "../../extract/document";
import {
findTweetNode,
findTweetNodeById,
formatMediaList,
formatTweetAuthor,
getTweetAuthorMetadata,
getTweetText,
getUser,
isRecord,
normalizeTitle,
toHighResXImageUrl,
toXTweet,
} from "./shared";
import type { JsonObject } from "./types";
function resolveArticleMediaUrl(mediaInfo: JsonObject): string {
const rawUrl =
(typeof mediaInfo.original_img_url === "string" && mediaInfo.original_img_url) ||
(typeof mediaInfo.url === "string" && mediaInfo.url) ||
"";
return rawUrl ? toHighResXImageUrl(rawUrl) : "";
}
function normalizeEntityMap(entityMap: unknown): Map<string, JsonObject> {
const normalized = new Map<string, JsonObject>();
if (Array.isArray(entityMap)) {
for (const entry of entityMap) {
if (!isRecord(entry)) {
continue;
}
const key =
typeof entry.key === "string" || typeof entry.key === "number"
? String(entry.key)
: undefined;
const value = isRecord(entry.value) ? entry.value : undefined;
if (!key || !value) {
continue;
}
normalized.set(key, value);
}
return normalized;
}
if (!isRecord(entityMap)) {
return normalized;
}
for (const [key, value] of Object.entries(entityMap)) {
if (!isRecord(value)) {
continue;
}
normalized.set(key, value);
}
return normalized;
}
function getEntityMarkdown(entityMap: Map<string, JsonObject>, entityKey: unknown): string | null {
const key =
typeof entityKey === "string" || typeof entityKey === "number"
? String(entityKey)
: undefined;
if (!key) {
return null;
}
const entity = entityMap.get(key);
if (!entity || entity.type !== "MARKDOWN") {
return null;
}
const data = isRecord(entity.data) ? entity.data : {};
if (typeof data.markdown !== "string") {
return null;
}
const markdown = data.markdown.trim();
return markdown || null;
}
function getLinkUrl(entityMap: Map<string, JsonObject>, entityKey: unknown): string | null {
const key =
typeof entityKey === "string" || typeof entityKey === "number"
? String(entityKey)
: undefined;
if (!key) {
return null;
}
const entity = entityMap.get(key);
if (!entity || entity.type !== "LINK") {
return null;
}
const data = isRecord(entity.data) ? entity.data : {};
const candidates = [
data.expanded_url,
data.expandedUrl,
data.original_url,
data.originalUrl,
data.url,
data.display_url,
data.displayUrl,
];
for (const candidate of candidates) {
if (typeof candidate === "string" && candidate.trim()) {
return candidate.trim();
}
}
return null;
}
function getTweetId(entityMap: Map<string, JsonObject>, entityKey: unknown): string | null {
const key =
typeof entityKey === "string" || typeof entityKey === "number"
? String(entityKey)
: undefined;
if (!key) {
return null;
}
const entity = entityMap.get(key);
if (!entity || entity.type !== "TWEET") {
return null;
}
const data = isRecord(entity.data) ? entity.data : {};
if (typeof data.tweetId !== "string") {
return null;
}
return data.tweetId;
}
function buildMediaUrlMap(articleResult: JsonObject): Map<string, string> {
const mediaMap = new Map<string, string>();
const mediaEntities = Array.isArray(articleResult.media_entities) ? articleResult.media_entities : [];
for (const entity of mediaEntities) {
if (!isRecord(entity) || typeof entity.media_id !== "string" || !isRecord(entity.media_info)) {
continue;
}
const mediaInfo = entity.media_info;
const url = resolveArticleMediaUrl(mediaInfo);
if (url) {
mediaMap.set(entity.media_id, url);
}
}
const coverMedia = isRecord(articleResult.cover_media) ? articleResult.cover_media : null;
if (coverMedia && typeof coverMedia.media_id === "string" && isRecord(coverMedia.media_info)) {
const url = resolveArticleMediaUrl(coverMedia.media_info);
if (url) {
mediaMap.set(coverMedia.media_id, url);
}
}
return mediaMap;
}
function getMediaMarkdown(entityMap: Map<string, JsonObject>, entityKey: unknown, mediaMap: Map<string, string>): string[] {
const key =
typeof entityKey === "string" || typeof entityKey === "number"
? String(entityKey)
: undefined;
if (!key) {
return [];
}
const entity = entityMap.get(key);
if (!entity || entity.type !== "MEDIA") {
return [];
}
const data = isRecord(entity.data) ? entity.data : {};
const mediaItems = Array.isArray(data.mediaItems) ? data.mediaItems : [];
const urls: string[] = [];
for (const item of mediaItems) {
if (!isRecord(item) || typeof item.mediaId !== "string") {
continue;
}
const url = mediaMap.get(item.mediaId);
if (url && !urls.includes(url)) {
urls.push(url);
}
}
return urls.map((url) => `![](${url})`);
}
function resolveTweetMarkdown(payloads: unknown[], tweetId: string, pageUrl: string): string | null {
for (const payload of payloads) {
const tweet = findTweetNodeById(payload, tweetId);
if (!tweet) {
continue;
}
const xTweet = toXTweet(tweet, pageUrl);
const author = formatTweetAuthor(xTweet) ?? xTweet.url;
const lines = [`> ${author}`, ...xTweet.text.split("\n").map((line) => `> ${line}`)];
const media = formatMediaList(xTweet.media).map((line) =>
line.startsWith("photo: ") ? `> ![](${line.slice("photo: ".length)})` : `> - ${line}`,
);
const parts = [lines.join("\n")];
if (media.length > 0) {
parts.push([">", ...media].join("\n"));
}
parts.push(`> ${xTweet.url}`);
return parts.join("\n").trim();
}
return `> Embedded tweet: https://x.com/i/status/${tweetId}`;
}
function replaceLinkEntities(text: string, block: JsonObject, entityMap: Map<string, JsonObject>): string {
const entityRanges = Array.isArray(block.entityRanges) ? block.entityRanges : [];
const replacements = entityRanges
.filter((range): range is JsonObject => isRecord(range))
.map((range) => {
const offset = typeof range.offset === "number" ? range.offset : -1;
const length = typeof range.length === "number" ? range.length : -1;
const url = getLinkUrl(entityMap, range.key);
return { offset, length, url };
})
.filter((range) => range.offset >= 0 && range.length > 0 && range.url)
.sort((left, right) => right.offset - left.offset);
let next = text;
for (const replacement of replacements) {
next =
next.slice(0, replacement.offset) +
replacement.url +
next.slice(replacement.offset + replacement.length);
}
return next;
}
function renderAtomicBlock(
block: JsonObject,
entityMap: Map<string, JsonObject>,
mediaMap: Map<string, string>,
payloads: unknown[],
pageUrl: string,
): string | null {
const entityRanges = Array.isArray(block.entityRanges) ? block.entityRanges : [];
const parts: string[] = [];
for (const range of entityRanges) {
if (!isRecord(range)) {
continue;
}
const markdown = getEntityMarkdown(entityMap, range.key);
if (markdown) {
parts.push(markdown);
continue;
}
const mediaMarkdown = getMediaMarkdown(entityMap, range.key, mediaMap);
if (mediaMarkdown.length > 0) {
parts.push(mediaMarkdown.join("\n\n"));
continue;
}
const tweetId = getTweetId(entityMap, range.key);
if (tweetId) {
const tweetMarkdown = resolveTweetMarkdown(payloads, tweetId, pageUrl);
if (tweetMarkdown) {
parts.push(tweetMarkdown);
}
}
}
if (parts.length === 0) {
return null;
}
return parts.join("\n\n");
}
function renderArticleBlocks(
blocks: unknown[],
entityMap: Map<string, JsonObject>,
mediaMap: Map<string, string>,
payloads: unknown[],
pageUrl: string,
): string {
const parts: string[] = [];
let orderedCounter = 0;
for (const block of blocks) {
if (!isRecord(block)) {
continue;
}
const blockType = typeof block.type === "string" ? block.type : "unstyled";
const rawText = typeof block.text === "string" ? block.text : "";
const text = replaceLinkEntities(rawText, block, entityMap).trim();
if (!text && blockType !== "atomic") {
continue;
}
if (blockType !== "ordered-list-item") {
orderedCounter = 0;
}
switch (blockType) {
case "header-one":
parts.push(`# ${text}`);
break;
case "header-two":
parts.push(`## ${text}`);
break;
case "header-three":
parts.push(`### ${text}`);
break;
case "blockquote":
parts.push(`> ${text}`);
break;
case "unordered-list-item":
parts.push(`- ${text}`);
break;
case "ordered-list-item":
orderedCounter += 1;
parts.push(`${orderedCounter}. ${text}`);
break;
case "code-block":
parts.push(`\`\`\`\n${text}\n\`\`\``);
break;
case "atomic": {
const markdown = renderAtomicBlock(block, entityMap, mediaMap, payloads, pageUrl);
if (markdown) {
parts.push(markdown);
}
break;
}
default:
parts.push(text);
break;
}
}
return parts.join("\n\n").trim();
}
function getArticleResult(tweet: JsonObject): JsonObject | null {
if (
isRecord(tweet.article) &&
isRecord(tweet.article.article_results) &&
isRecord(tweet.article.article_results.result)
) {
return tweet.article.article_results.result as JsonObject;
}
return null;
}
function extractSummary(markdown: string): string | undefined {
const segments = markdown
.split(/\n\n+/)
.map((segment) => segment.trim())
.filter(Boolean);
const preferred = segments.find((segment) => !/^(#|>|- |\d+\. |\`\`\`)/.test(segment));
return preferred?.slice(0, 220);
}
export function extractArticleDocumentFromPayload(
payload: unknown,
statusId: string,
pageUrl: string,
payloads: unknown[] = [payload],
): ExtractedDocument | null {
const tweet = findTweetNode(payload, statusId);
if (!tweet) {
return null;
}
const articleResult = getArticleResult(tweet);
if (!articleResult) {
return null;
}
const title = typeof articleResult.title === "string" ? articleResult.title.trim() : undefined;
const contentState = isRecord(articleResult.content_state) ? articleResult.content_state : {};
const blocks = Array.isArray(contentState.blocks) ? contentState.blocks : [];
const entityMap = normalizeEntityMap(contentState.entityMap);
const mediaMap = buildMediaUrlMap(articleResult);
const richMarkdown = renderArticleBlocks(blocks, entityMap, mediaMap, payloads, pageUrl);
const plainText = typeof articleResult.plain_text === "string" ? articleResult.plain_text.trim() : "";
const markdown = richMarkdown || plainText || getTweetText(tweet);
if (!markdown) {
return null;
}
const xTweet = toXTweet(tweet, pageUrl);
const user = getUser(tweet);
const coverMedia = isRecord(articleResult.cover_media) ? articleResult.cover_media : null;
const coverMediaInfo = coverMedia && isRecord(coverMedia.media_info) ? coverMedia.media_info : null;
const coverImage = coverMediaInfo ? resolveArticleMediaUrl(coverMediaInfo) || undefined : undefined;
return {
url: pageUrl,
canonicalUrl: xTweet.url,
title: title || normalizeTitle(xTweet.text, "X Article"),
author: formatTweetAuthor(xTweet),
siteName: "X",
publishedAt: xTweet.createdAt,
summary: extractSummary(markdown) || xTweet.text.slice(0, 200) || undefined,
adapter: "x",
metadata: {
kind: "x/article",
tweetId: xTweet.id,
coverImage,
authorName: xTweet.authorName ?? user.name,
authorUsername: xTweet.author ?? user.screenName,
authorUrl: (xTweet.author ?? user.screenName) ? `https://x.com/${xTweet.author ?? user.screenName}` : undefined,
...getTweetAuthorMetadata(xTweet),
},
content: [{ type: "markdown", markdown }],
};
}
@@ -0,0 +1,117 @@
import type { Adapter, AdapterLoginInfo } from "../types";
import { detectInteractionGate } from "../../browser/interaction-gates";
import type { ExtractedDocument } from "../../extract/document";
import { collectMediaFromDocument } from "../../media/markdown-media";
import { extractArticleDocumentFromPayload } from "./article";
import { buildNeedsLoginResult, detectXLogin } from "./login";
import { extractStatusId, isXHost } from "./match";
import { collectXJsonPayloads, waitForInitialXPayload } from "./payloads";
import { extractSingleTweetDocumentFromPayload } from "./single";
import { extractThreadDocumentFromPayloads } from "./thread";
import { loadFullXThread } from "./thread-loader";
function extractDocumentFromPayloads(
payloads: unknown[],
statusId: string,
pageUrl: string,
): ExtractedDocument | null {
for (const payload of payloads) {
const articleDocument = extractArticleDocumentFromPayload(payload, statusId, pageUrl, payloads);
if (articleDocument) {
return articleDocument;
}
}
const threadDocument = extractThreadDocumentFromPayloads(payloads, statusId, pageUrl);
if (threadDocument) {
return threadDocument;
}
for (const payload of payloads) {
const singleDocument = extractSingleTweetDocumentFromPayload(payload, statusId, pageUrl);
if (singleDocument) {
return singleDocument;
}
}
return null;
}
async function ensureXLoginState(context: Parameters<Adapter["process"]>[0]): Promise<AdapterLoginInfo> {
return detectXLogin(context);
}
export const xAdapter: Adapter = {
name: "x",
match(input) {
return isXHost(input.url.hostname);
},
async checkLogin(context) {
return detectXLogin(context);
},
async process(context) {
const statusId = extractStatusId(context.input.url);
if (!statusId) {
return {
status: "no_document",
};
}
context.log.info(`Loading ${context.input.url.toString()} with x adapter`);
await context.browser.goto(context.input.url.toString(), context.timeoutMs);
const interaction = await detectInteractionGate(context.browser);
if (interaction) {
return {
status: "needs_interaction",
interaction,
};
}
let login = await ensureXLoginState(context);
if (login.state === "logged_out") {
return buildNeedsLoginResult(login);
}
await waitForInitialXPayload(context);
await loadFullXThread(context, statusId);
const pageUrl = await context.browser.getURL();
const postLoadInteraction = await detectInteractionGate(context.browser);
if (postLoadInteraction) {
return {
status: "needs_interaction",
interaction: postLoadInteraction,
login,
};
}
login = await ensureXLoginState(context).catch(() => login);
if (login.state === "logged_out") {
return buildNeedsLoginResult(login);
}
const payloads = await collectXJsonPayloads(context);
if (payloads.length === 0) {
return {
status: "no_document",
login,
};
}
const document = extractDocumentFromPayloads(payloads, statusId, pageUrl);
if (document) {
return {
status: "ok",
document,
media: collectMediaFromDocument(document),
login,
};
}
return {
status: "no_document",
login,
};
},
};
@@ -0,0 +1,80 @@
import type { AdapterContext, AdapterLoginInfo, AdapterProcessResult } from "../types";
interface XLoginSnapshot {
currentUrl: string;
hasAccountMenu: boolean;
hasLoginInputs: boolean;
bodyText: string;
}
export async function detectXLogin(context: AdapterContext): Promise<AdapterLoginInfo> {
const snapshot = await context.browser.evaluate<XLoginSnapshot>(`
(() => {
const bodyText = (document.body?.innerText ?? "").slice(0, 2500);
return {
currentUrl: window.location.href,
hasAccountMenu: Boolean(
document.querySelector(
'[data-testid="SideNav_AccountSwitcher_Button"], [data-testid="AppTabBar_Profile_Link"], [aria-label="Account menu"]'
)
),
hasLoginInputs: Boolean(
document.querySelector(
'input[name="text"], input[name="password"], input[autocomplete="username"], input[autocomplete="current-password"]'
)
),
bodyText,
};
})()
`).catch(async () => ({
currentUrl: await context.browser.getURL().catch(() => context.input.url.toString()),
hasAccountMenu: false,
hasLoginInputs: false,
bodyText: "",
}));
if (
/\/i\/flow\/login|\/login/i.test(snapshot.currentUrl) ||
snapshot.hasLoginInputs ||
/sign in to x|join x today|登录 x|注册 x|登录到 x/i.test(snapshot.bodyText)
) {
return {
provider: "x",
state: "logged_out",
required: true,
reason: "X login page detected",
};
}
if (snapshot.hasAccountMenu) {
return {
provider: "x",
state: "logged_in",
};
}
return {
provider: "x",
state: "unknown",
};
}
export function buildNeedsLoginResult(login: AdapterLoginInfo): AdapterProcessResult {
return {
status: "needs_interaction",
login: {
...login,
provider: "x",
state: login.state === "logged_in" ? "unknown" : login.state,
required: true,
},
interaction: {
type: "wait_for_interaction",
kind: "login",
provider: "x",
reason: login.reason,
prompt: "Please sign in to X in the opened Chrome window. Extraction will continue automatically once login is detected.",
requiresVisibleBrowser: true,
},
};
}
@@ -0,0 +1,9 @@
export function isXHost(hostname: string): boolean {
return ["x.com", "www.x.com", "twitter.com", "www.twitter.com"].includes(hostname);
}
export function extractStatusId(url: URL): string | undefined {
const match = url.pathname.match(/\/(?:status|article)\/(\d+)/);
return match?.[1];
}
@@ -0,0 +1,50 @@
import type { AdapterContext } from "../types";
import { filterXGraphQlEntries } from "./shared";
export function getRelevantXThreadEntries(context: AdapterContext) {
return filterXGraphQlEntries(context.network.getEntries()).filter(
(entry) =>
entry.method === "GET" &&
entry.finished &&
(
entry.url.includes("TweetDetail") ||
entry.url.includes("TweetResultByRestId") ||
entry.url.includes("TweetResultsByRestIds")
),
);
}
export async function prefetchRelevantXThreadBodies(context: AdapterContext): Promise<void> {
const entries = getRelevantXThreadEntries(context).filter((entry) => entry.body === undefined && !entry.bodyError);
for (const entry of entries) {
await context.network.ensureBody(entry);
}
}
export async function collectXJsonPayloads(context: AdapterContext): Promise<unknown[]> {
await prefetchRelevantXThreadBodies(context);
const entries = getRelevantXThreadEntries(context);
const payloads: unknown[] = [];
for (const entry of entries) {
const payload = await context.network.getJsonBody(entry);
if (payload) {
payloads.push(payload);
}
}
return payloads;
}
export async function waitForInitialXPayload(context: AdapterContext): Promise<void> {
try {
await context.network.waitForResponse(
(entry) =>
entry.url.includes("/graphql/") &&
(entry.url.includes("TweetDetail") || entry.url.includes("TweetResultByRestId")),
{ timeoutMs: Math.min(context.timeoutMs, 15_000) },
);
await prefetchRelevantXThreadBodies(context);
} catch {
context.log.debug("No tweet GraphQL response observed before timeout.");
}
}
@@ -0,0 +1,386 @@
import path from "node:path";
import type { NetworkEntry } from "../../browser/network-journal";
import type { XMedia, XQuotedTweet, XTweet, XUser, JsonObject } from "./types";
const X_IMAGE_EXTENSIONS = new Set(["jpg", "jpeg", "png", "webp", "gif", "bmp", "avif"]);
function emptyObject(): JsonObject {
return {};
}
export function isRecord(value: unknown): value is JsonObject {
return Boolean(value) && typeof value === "object" && !Array.isArray(value);
}
export function walk(value: unknown, visitor: (node: unknown) => boolean | void): boolean {
if (visitor(value)) {
return true;
}
if (Array.isArray(value)) {
for (const item of value) {
if (walk(item, visitor)) {
return true;
}
}
return false;
}
if (isRecord(value)) {
for (const child of Object.values(value)) {
if (walk(child, visitor)) {
return true;
}
}
}
return false;
}
function hasTweetText(node: JsonObject): boolean {
const legacy = isRecord(node.legacy) ? node.legacy : emptyObject();
return (
typeof legacy.full_text === "string" ||
typeof getNoteTweetText(node) === "string"
);
}
export function findTweetNodeById(payload: unknown, tweetId: string): JsonObject | null {
let match: JsonObject | null = null;
walk(payload, (node) => {
if (!isRecord(node) || typeof node.rest_id !== "string" || !isRecord(node.legacy)) {
return false;
}
if (!hasTweetText(node)) {
return false;
}
if (node.rest_id === tweetId) {
match = node;
return true;
}
return false;
});
return match;
}
export function findTweetNode(payload: unknown, statusId: string): JsonObject | null {
let firstMatch: JsonObject | null = null;
const exactMatch = findTweetNodeById(payload, statusId);
if (exactMatch) {
return exactMatch;
}
walk(payload, (node) => {
if (!isRecord(node) || typeof node.rest_id !== "string" || !isRecord(node.legacy)) {
return false;
}
if (!hasTweetText(node)) {
return false;
}
if (!firstMatch) {
firstMatch = node;
}
return false;
});
return firstMatch;
}
export function getLegacy(tweet: JsonObject): JsonObject {
return isRecord(tweet.legacy) ? tweet.legacy : emptyObject();
}
export function unwrapTweetResult(node: unknown): JsonObject | null {
if (!isRecord(node)) {
return null;
}
if (node.__typename === "TweetWithVisibilityResults" && isRecord(node.tweet)) {
return unwrapTweetResult(node.tweet);
}
const tweet = isRecord(node.tweet) ? (node.tweet as JsonObject) : node;
if (typeof tweet.rest_id !== "string" || !isRecord(tweet.legacy)) {
return null;
}
return tweet;
}
export function getUser(tweet: JsonObject): XUser {
const result =
isRecord(tweet.core) &&
isRecord(tweet.core.user_results) &&
isRecord(tweet.core.user_results.result)
? (tweet.core.user_results.result as JsonObject)
: emptyObject();
const legacy = isRecord(result.legacy) ? result.legacy : emptyObject();
const core = isRecord(result.core) ? result.core : emptyObject();
return {
name:
(typeof legacy.name === "string" ? legacy.name : undefined) ??
(typeof core.name === "string" ? core.name : undefined),
screenName:
(typeof legacy.screen_name === "string" ? legacy.screen_name : undefined) ??
(typeof core.screen_name === "string" ? core.screen_name : undefined),
};
}
function getNoteTweetResult(tweet: JsonObject): JsonObject | null {
if (
!isRecord(tweet.note_tweet) ||
!isRecord(tweet.note_tweet.note_tweet_results) ||
!isRecord(tweet.note_tweet.note_tweet_results.result)
) {
return null;
}
return tweet.note_tweet.note_tweet_results.result as JsonObject;
}
function getNoteTweetText(tweet: JsonObject): string | undefined {
const noteTweet = getNoteTweetResult(tweet);
return typeof noteTweet?.text === "string" ? noteTweet.text : undefined;
}
interface TweetUrlEntity {
url: string;
expandedUrl?: string;
displayUrl?: string;
}
function collectTweetUrlEntities(values: unknown[]): TweetUrlEntity[] {
return values.reduce<TweetUrlEntity[]>((entities, value) => {
if (!isRecord(value) || typeof value.url !== "string" || !value.url) {
return entities;
}
entities.push({
url: value.url,
expandedUrl: typeof value.expanded_url === "string" ? value.expanded_url : undefined,
displayUrl: typeof value.display_url === "string" ? value.display_url : undefined,
});
return entities;
}, []);
}
function getTweetUrlEntities(tweet: JsonObject): TweetUrlEntity[] {
const noteTweet = getNoteTweetResult(tweet);
const noteTweetEntitySet = noteTweet && isRecord(noteTweet.entity_set) ? noteTweet.entity_set : emptyObject();
const noteTweetUrls = collectTweetUrlEntities(Array.isArray(noteTweetEntitySet.urls) ? noteTweetEntitySet.urls : []);
const legacy = getLegacy(tweet);
const legacyEntities = isRecord(legacy.entities) ? legacy.entities : emptyObject();
const legacyUrls = collectTweetUrlEntities(Array.isArray(legacyEntities.urls) ? legacyEntities.urls : []);
const seen = new Set<string>();
return [...noteTweetUrls, ...legacyUrls].filter((value) => {
if (seen.has(value.url)) {
return false;
}
seen.add(value.url);
return true;
});
}
export function getTweetText(tweet: JsonObject): string {
const legacy = getLegacy(tweet);
let text =
getNoteTweetText(tweet) ?? (typeof legacy.full_text === "string" ? legacy.full_text : "");
for (const value of getTweetUrlEntities(tweet)) {
const replacement =
(typeof value.expandedUrl === "string" && value.expandedUrl) ||
(typeof value.displayUrl === "string" && value.displayUrl) ||
value.url;
text = text.replaceAll(value.url, replacement);
}
const extendedEntities = isRecord(legacy.extended_entities) ? legacy.extended_entities : emptyObject();
const media = Array.isArray(extendedEntities.media) ? extendedEntities.media : [];
for (const value of media) {
if (isRecord(value) && typeof value.url === "string") {
text = text.replaceAll(value.url, "").trim();
}
}
return text.replace(/\n{3,}/g, "\n\n").trim();
}
function normalizeXImageExtension(raw: string | undefined | null): string | undefined {
if (!raw) {
return undefined;
}
const normalized = raw.replace(/^\./, "").trim().toLowerCase();
if (!normalized) {
return undefined;
}
return normalized === "jpeg" ? "jpg" : normalized;
}
export function toHighResXImageUrl(rawUrl: string): string {
try {
const parsed = new URL(rawUrl);
if (parsed.hostname.toLowerCase() !== "pbs.twimg.com") {
return rawUrl;
}
const pathExtension = normalizeXImageExtension(path.posix.extname(parsed.pathname));
const format = normalizeXImageExtension(parsed.searchParams.get("format")) ?? pathExtension;
if (!format || !X_IMAGE_EXTENSIONS.has(format)) {
return rawUrl;
}
if (pathExtension) {
parsed.pathname = parsed.pathname.replace(new RegExp(`\\.${pathExtension}$`, "i"), "");
}
parsed.searchParams.set("format", format);
parsed.searchParams.set("name", "4096x4096");
return parsed.toString();
} catch {
return rawUrl;
}
}
export function getTweetMedia(tweet: JsonObject): XMedia[] {
const legacy = getLegacy(tweet);
const extendedEntities = isRecord(legacy.extended_entities) ? legacy.extended_entities : emptyObject();
const media = Array.isArray(extendedEntities.media) ? extendedEntities.media : [];
return media
.map((value) => {
if (!isRecord(value) || typeof value.type !== "string") {
return null;
}
if (value.type === "photo" && typeof value.media_url_https === "string") {
return {
type: value.type,
url: toHighResXImageUrl(value.media_url_https),
alt: typeof value.ext_alt_text === "string" ? value.ext_alt_text : undefined,
};
}
if ((value.type === "video" || value.type === "animated_gif") && typeof value.media_url_https === "string") {
return {
type: value.type,
url: value.media_url_https,
};
}
return null;
})
.filter((value): value is XMedia => value !== null);
}
export function getTweetUrl(tweet: JsonObject, fallbackUrl: string): string {
const user = getUser(tweet);
const fallbackScreenName = extractScreenNameFromUrl(fallbackUrl);
const id = typeof tweet.rest_id === "string" ? tweet.rest_id : "";
const screenName = user.screenName ?? fallbackScreenName;
if (screenName && id) {
return `https://x.com/${screenName}/status/${id}`;
}
return fallbackUrl;
}
export function getQuotedTweet(tweet: JsonObject, fallbackUrl: string): XQuotedTweet | undefined {
const quoted = unwrapTweetResult(
isRecord(tweet.quoted_status_result) ? tweet.quoted_status_result.result : null,
);
if (!quoted) {
return undefined;
}
const user = getUser(quoted);
return {
id: typeof quoted.rest_id === "string" ? quoted.rest_id : "",
author: user.screenName,
authorName: user.name,
text: getTweetText(quoted),
url: getTweetUrl(quoted, fallbackUrl),
media: getTweetMedia(quoted),
};
}
export function extractScreenNameFromUrl(url: string): string | undefined {
try {
const parsed = new URL(url);
const match = parsed.pathname.match(/^\/([^/]+)\/(?:status|article)\//);
if (!match) {
return undefined;
}
if (match[1] === "i") {
return undefined;
}
return match[1];
} catch {
return undefined;
}
}
export function toXTweet(tweet: JsonObject, fallbackUrl: string): XTweet {
const legacy = getLegacy(tweet);
const user = getUser(tweet);
const fallbackScreenName = extractScreenNameFromUrl(fallbackUrl);
const screenName = user.screenName ?? fallbackScreenName;
return {
id: typeof tweet.rest_id === "string" ? tweet.rest_id : "",
author: screenName,
authorName: user.name,
text: getTweetText(tweet),
likes: typeof legacy.favorite_count === "number" ? legacy.favorite_count : 0,
retweets: typeof legacy.retweet_count === "number" ? legacy.retweet_count : 0,
replies: typeof legacy.reply_count === "number" ? legacy.reply_count : 0,
createdAt: typeof legacy.created_at === "string" ? legacy.created_at : undefined,
inReplyTo: typeof legacy.in_reply_to_status_id_str === "string" ? legacy.in_reply_to_status_id_str : undefined,
url: getTweetUrl(tweet, fallbackUrl),
media: getTweetMedia(tweet),
quotedTweet: getQuotedTweet(tweet, fallbackUrl),
};
}
export function normalizeTitle(text: string, fallback: string): string {
const firstLine = text.split("\n")[0]?.trim();
if (!firstLine) {
return fallback;
}
return firstLine.slice(0, 120);
}
export function formatTweetAuthor(tweet: XTweet): string | undefined {
if (tweet.author && tweet.authorName) {
return `${tweet.authorName} (@${tweet.author})`;
}
if (tweet.author) {
return `@${tweet.author}`;
}
return tweet.authorName;
}
export function getTweetAuthorMetadata(tweet: XTweet): Record<string, unknown> {
return {
authorName: tweet.authorName,
authorUsername: tweet.author,
authorUrl: tweet.author ? `https://x.com/${tweet.author}` : undefined,
};
}
export function formatMediaList(media: XMedia[]): string[] {
return media.map((item) => {
if (item.type === "photo") {
return `photo: ${item.url}`;
}
return `${item.type}: ${item.url}`;
});
}
export function filterXGraphQlEntries(entries: NetworkEntry[]): NetworkEntry[] {
return entries.filter((entry) => entry.url.includes("/graphql/"));
}
@@ -0,0 +1,87 @@
import type { ExtractedDocument, ContentBlock } from "../../extract/document";
import { findTweetNode, formatMediaList, formatTweetAuthor, getTweetAuthorMetadata, normalizeTitle, toXTweet } from "./shared";
export function extractSingleTweetDocumentFromPayload(
payload: unknown,
statusId: string,
pageUrl: string,
): ExtractedDocument | null {
const tweet = findTweetNode(payload, statusId);
if (!tweet) {
return null;
}
const xTweet = toXTweet(tweet, pageUrl);
const content: ContentBlock[] = [];
if (xTweet.text) {
content.push({ type: "paragraph", text: xTweet.text });
}
for (const mediaLine of formatMediaList(xTweet.media)) {
if (mediaLine.startsWith("photo: ")) {
content.push({
type: "image",
url: mediaLine.slice("photo: ".length),
});
} else {
content.push({
type: "list",
ordered: false,
items: [mediaLine],
});
}
}
if (xTweet.quotedTweet) {
const quotedLines: string[] = [];
const quotedAuthor =
xTweet.quotedTweet.author && xTweet.quotedTweet.authorName
? `${xTweet.quotedTweet.authorName} (@${xTweet.quotedTweet.author})`
: xTweet.quotedTweet.author
? `@${xTweet.quotedTweet.author}`
: xTweet.quotedTweet.authorName;
if (quotedAuthor) {
quotedLines.push(quotedAuthor);
}
if (xTweet.quotedTweet.text) {
quotedLines.push(xTweet.quotedTweet.text);
}
quotedLines.push(...formatMediaList(xTweet.quotedTweet.media));
if (quotedLines.length > 0) {
content.push({ type: "heading", depth: 2, text: "Quoted Tweet" });
content.push({ type: "quote", text: quotedLines.join("\n\n") });
}
}
return {
url: pageUrl,
canonicalUrl: xTweet.url,
title: normalizeTitle(
xTweet.author ? `@${xTweet.author}: ${xTweet.text}` : xTweet.text,
"Tweet",
),
author: formatTweetAuthor(xTweet),
siteName: "X",
publishedAt: xTweet.createdAt,
summary: xTweet.text.slice(0, 200) || undefined,
adapter: "x",
metadata: {
kind: "x/post",
tweetId: xTweet.id,
...getTweetAuthorMetadata(xTweet),
conversationId:
typeof tweet.legacy === "object" &&
tweet.legacy !== null &&
typeof (tweet.legacy as Record<string, unknown>).conversation_id_str === "string"
? (tweet.legacy as Record<string, unknown>).conversation_id_str
: undefined,
favoriteCount: xTweet.likes,
replyCount: xTweet.replies,
retweetCount: xTweet.retweets,
},
content,
};
}
@@ -0,0 +1,286 @@
import type { AdapterContext } from "../types";
import { extractThreadTweetsFromPayloads } from "./thread";
import { collectXJsonPayloads, getRelevantXThreadEntries, prefetchRelevantXThreadBodies } from "./payloads";
interface ClickTextResult {
clicked: boolean;
text?: string;
}
interface ScrollStepResult {
moved: boolean;
atTop: boolean;
atBottom: boolean;
}
interface ThreadProgress {
tweetCount: number;
firstTweetId?: string;
lastTweetId?: string;
requestCount: number;
tweetDetailCount: number;
}
interface TopProbeState {
requestCount: number;
tweetDetailCount: number;
scrollHeight: number;
}
function sleep(ms: number): Promise<void> {
return new Promise((resolve) => setTimeout(resolve, ms));
}
async function waitForXNetworkSettle(context: AdapterContext, reason: string): Promise<void> {
try {
await context.network.waitForIdle({
idleMs: 650,
timeoutMs: Math.min(context.timeoutMs, 5_000),
});
} catch {
context.log.debug(`Network idle timed out after ${reason}.`);
}
}
async function captureTopProbeState(context: AdapterContext): Promise<TopProbeState> {
const entries = getRelevantXThreadEntries(context);
const scrollHeight = await context.browser.evaluate<number>(`
(() => {
const scrollRoot = document.scrollingElement ?? document.documentElement ?? document.body;
return scrollRoot.scrollHeight;
})()
`);
return {
requestCount: entries.length,
tweetDetailCount: entries.filter((entry) => entry.url.includes("TweetDetail")).length,
scrollHeight,
};
}
async function waitForTopProbe(context: AdapterContext): Promise<boolean> {
const initial = await captureTopProbeState(context);
const deadline = Date.now() + 1_200;
while (Date.now() < deadline) {
try {
await context.network.waitForIdle({
idleMs: 250,
timeoutMs: 350,
});
} catch {
// Keep polling until the shorter top-probe budget expires.
}
await prefetchRelevantXThreadBodies(context);
const next = await captureTopProbeState(context);
if (
next.requestCount > initial.requestCount ||
next.tweetDetailCount > initial.tweetDetailCount ||
next.scrollHeight > initial.scrollHeight + 4
) {
context.log.debug("Observed additional X thread activity while probing the page top.");
return true;
}
await sleep(120);
}
return false;
}
async function scrollThreadToTop(context: AdapterContext): Promise<void> {
let settledTopChecks = 0;
while (settledTopChecks < 2) {
const scroll = await context.browser.evaluate<ScrollStepResult>(`
(() => {
const scrollRoot = document.scrollingElement ?? document.documentElement ?? document.body;
const before = window.scrollY;
window.scrollTo({ top: 0, left: 0, behavior: "instant" });
const after = window.scrollY;
return {
moved: after !== before,
atTop: after <= 4,
atBottom: window.innerHeight + after >= scrollRoot.scrollHeight - 4,
};
})()
`);
await sleep(140);
await waitForXNetworkSettle(context, "scrolling X thread to top");
await prefetchRelevantXThreadBodies(context);
if (scroll.moved) {
settledTopChecks = 0;
continue;
}
const observedTopActivity = await waitForTopProbe(context);
if (observedTopActivity) {
settledTopChecks = 0;
continue;
}
settledTopChecks += 1;
}
}
async function clickVisibleShowReplies(context: AdapterContext): Promise<ClickTextResult> {
return context.browser.evaluate<ClickTextResult>(`
(() => {
const normalize = (value) => value.replace(/\\s+/g, " ").trim();
const matches = [
/^Show replies$/i,
/^Show more replies$/i,
/^Show additional replies$/i,
/^显示回复$/,
/^展开回复$/,
];
const isVisible = (element) => {
if (!(element instanceof HTMLElement)) {
return false;
}
const rect = element.getBoundingClientRect();
const style = window.getComputedStyle(element);
return (
rect.width > 0 &&
rect.height > 0 &&
style.visibility !== "hidden" &&
style.display !== "none"
);
};
const selectors = [
"a",
"button",
'[role="button"]',
'[role="link"]',
];
for (const element of document.querySelectorAll(selectors.join(","))) {
if (!isVisible(element)) {
continue;
}
const text = normalize(element.textContent ?? "");
if (!text || !matches.some((pattern) => pattern.test(text))) {
continue;
}
element.scrollIntoView({ block: "center", inline: "nearest" });
if (element instanceof HTMLElement) {
element.click();
return { clicked: true, text };
}
}
return { clicked: false };
})()
`);
}
async function expandVisibleShowReplies(context: AdapterContext): Promise<number> {
let clickCount = 0;
while (clickCount < 8) {
const result = await clickVisibleShowReplies(context).catch<ClickTextResult>(() => ({ clicked: false }));
if (!result.clicked) {
break;
}
clickCount += 1;
context.log.debug(`Expanded X thread replies via "${result.text ?? "Show replies"}".`);
await sleep(250);
await waitForXNetworkSettle(context, "expanding Show replies");
await prefetchRelevantXThreadBodies(context);
}
return clickCount;
}
async function scrollThreadBy(context: AdapterContext, stepPx: number): Promise<ScrollStepResult> {
const result = await context.browser.evaluate<ScrollStepResult>(`
(() => {
const scrollRoot = document.scrollingElement ?? document.documentElement ?? document.body;
const before = window.scrollY;
window.scrollBy({ top: ${stepPx}, left: 0, behavior: "instant" });
const after = window.scrollY;
return {
moved: after !== before,
atTop: after <= 4,
atBottom: window.innerHeight + after >= scrollRoot.scrollHeight - 4,
};
})()
`);
await sleep(140);
await waitForXNetworkSettle(context, "scrolling X thread");
await prefetchRelevantXThreadBodies(context);
return result;
}
async function captureThreadProgress(context: AdapterContext, statusId: string): Promise<ThreadProgress> {
const entries = getRelevantXThreadEntries(context);
const payloads = await collectXJsonPayloads(context);
const tweets = extractThreadTweetsFromPayloads(payloads, statusId, context.input.url.toString());
return {
tweetCount: tweets.length,
firstTweetId: tweets[0]?.id,
lastTweetId: tweets[tweets.length - 1]?.id,
requestCount: entries.length,
tweetDetailCount: entries.filter((entry) => entry.url.includes("TweetDetail")).length,
};
}
export async function loadFullXThread(context: AdapterContext, statusId: string): Promise<void> {
await scrollThreadToTop(context);
let progress = await captureThreadProgress(context, statusId);
let stagnantRounds = 0;
let roundsWithoutMovement = 0;
let distanceWithoutThreadActivityPx = 0;
for (let round = 0; ; round += 1) {
const stepPx = round < 12 ? 1_200 : 1_600;
let expandedCount = await expandVisibleShowReplies(context);
const scroll = await scrollThreadBy(context, stepPx);
expandedCount += await expandVisibleShowReplies(context);
const nextProgress = await captureThreadProgress(context, statusId);
const grew =
nextProgress.tweetCount > progress.tweetCount ||
nextProgress.firstTweetId !== progress.firstTweetId ||
nextProgress.lastTweetId !== progress.lastTweetId ||
nextProgress.requestCount > progress.requestCount ||
nextProgress.tweetDetailCount > progress.tweetDetailCount;
if (grew) {
context.log.debug(
`X thread progress: ${nextProgress.tweetCount} tweets (${nextProgress.firstTweetId ?? "unknown"} -> ${nextProgress.lastTweetId ?? "unknown"}), ${nextProgress.requestCount} requests, ${nextProgress.tweetDetailCount} TweetDetail.`,
);
stagnantRounds = 0;
distanceWithoutThreadActivityPx = 0;
} else if (expandedCount > 0) {
stagnantRounds = 0;
distanceWithoutThreadActivityPx = 0;
} else {
stagnantRounds += 1;
distanceWithoutThreadActivityPx += stepPx;
}
roundsWithoutMovement = scroll.moved ? 0 : roundsWithoutMovement + 1;
progress = nextProgress;
if (scroll.atBottom && stagnantRounds >= 6) {
context.log.debug("Stopping X thread scroll after reaching page bottom with no further thread progress.");
break;
}
if (roundsWithoutMovement >= 2 && stagnantRounds >= 4) {
context.log.debug("Stopping X thread scroll after repeated downward scrolls no longer move the page.");
break;
}
if (distanceWithoutThreadActivityPx >= 24_000 && stagnantRounds >= 12) {
context.log.debug("Stopping X thread scroll after a long stretch with no thread-related progress.");
break;
}
}
}
@@ -0,0 +1,316 @@
import type { ExtractedDocument } from "../../extract/document";
import {
formatMediaList,
formatTweetAuthor,
getLegacy,
getTweetAuthorMetadata,
isRecord,
normalizeTitle,
toXTweet,
unwrapTweetResult,
} from "./shared";
import type { JsonObject, XQuotedTweet, XTweet } from "./types";
interface ParsedThreadTweet extends XTweet {
userId?: string;
conversationId?: string;
inReplyToUserId?: string;
sortTimestamp: number;
}
function compareTweetIds(left: string, right: string): number {
try {
const leftId = BigInt(left);
const rightId = BigInt(right);
if (leftId === rightId) {
return 0;
}
return leftId < rightId ? -1 : 1;
} catch {
return left.localeCompare(right);
}
}
function toTimestamp(value: string | undefined): number {
if (!value) {
return 0;
}
const parsed = Date.parse(value);
return Number.isNaN(parsed) ? 0 : parsed;
}
function scoreParsedTweet(tweet: ParsedThreadTweet): number {
return (
(tweet.text ? 4 : 0) +
(tweet.author ? 2 : 0) +
(tweet.authorName ? 2 : 0) +
(tweet.media.length > 0 ? 1 : 0)
);
}
function toParsedThreadTweet(tweet: JsonObject, pageUrl: string): ParsedThreadTweet {
const legacy = getLegacy(tweet);
const xTweet = toXTweet(tweet, pageUrl);
return {
...xTweet,
userId: typeof legacy.user_id_str === "string" ? legacy.user_id_str : undefined,
conversationId: typeof legacy.conversation_id_str === "string" ? legacy.conversation_id_str : undefined,
inReplyToUserId: typeof legacy.in_reply_to_user_id_str === "string" ? legacy.in_reply_to_user_id_str : undefined,
sortTimestamp: toTimestamp(xTweet.createdAt),
};
}
function collectTweetFromItemContent(
itemContent: unknown,
pageUrl: string,
tweets: Map<string, ParsedThreadTweet>,
): void {
if (!isRecord(itemContent)) {
return;
}
const tweet = unwrapTweetResult(
isRecord(itemContent.tweet_results) ? itemContent.tweet_results.result : null,
);
if (!tweet || typeof tweet.rest_id !== "string") {
return;
}
const parsed = toParsedThreadTweet(tweet, pageUrl);
const existing = tweets.get(parsed.id);
if (!existing || scoreParsedTweet(parsed) >= scoreParsedTweet(existing)) {
tweets.set(parsed.id, parsed);
}
}
function collectTweetsFromItems(
items: unknown,
pageUrl: string,
tweets: Map<string, ParsedThreadTweet>,
): void {
if (!Array.isArray(items)) {
return;
}
for (const item of items) {
if (!isRecord(item)) {
continue;
}
if (isRecord(item.item) && isRecord(item.item.itemContent)) {
collectTweetFromItemContent(item.item.itemContent, pageUrl, tweets);
continue;
}
if (isRecord(item.itemContent)) {
collectTweetFromItemContent(item.itemContent, pageUrl, tweets);
}
}
}
function getInstructions(payload: unknown): unknown[] {
if (!isRecord(payload) || !isRecord(payload.data)) {
return [];
}
const { data } = payload;
return (
(isRecord(data.threaded_conversation_with_injections_v2) &&
Array.isArray(data.threaded_conversation_with_injections_v2.instructions)
? data.threaded_conversation_with_injections_v2.instructions
: undefined) ??
(isRecord(data.threaded_conversation_with_injections) &&
Array.isArray(data.threaded_conversation_with_injections.instructions)
? data.threaded_conversation_with_injections.instructions
: undefined) ??
(isRecord(data.tweetResult) &&
isRecord(data.tweetResult.result) &&
isRecord(data.tweetResult.result.timeline) &&
Array.isArray(data.tweetResult.result.timeline.instructions)
? data.tweetResult.result.timeline.instructions
: [])
);
}
function parseTweetDetailPayload(payload: unknown, pageUrl: string): ParsedThreadTweet[] {
const tweets = new Map<string, ParsedThreadTweet>();
const instructions = getInstructions(payload);
for (const instruction of instructions) {
if (!isRecord(instruction)) {
continue;
}
collectTweetsFromItems(instruction.moduleItems, pageUrl, tweets);
if (!Array.isArray(instruction.entries)) {
continue;
}
for (const entry of instruction.entries) {
if (!isRecord(entry)) {
continue;
}
const content = isRecord(entry.content) ? entry.content : {};
collectTweetFromItemContent(content.itemContent, pageUrl, tweets);
collectTweetsFromItems(content.items, pageUrl, tweets);
}
}
return Array.from(tweets.values());
}
function buildContinuousThread(tweets: ParsedThreadTweet[], statusId: string): ParsedThreadTweet[] {
const byId = new Map<string, ParsedThreadTweet>();
for (const tweet of tweets) {
const existing = byId.get(tweet.id);
if (!existing || scoreParsedTweet(tweet) >= scoreParsedTweet(existing)) {
byId.set(tweet.id, tweet);
}
}
const rootTweet = byId.get(statusId);
if (!rootTweet?.userId || !rootTweet.conversationId) {
return [];
}
const candidates = Array.from(byId.values()).filter(
(tweet) =>
tweet.id === statusId ||
(tweet.userId === rootTweet.userId && tweet.conversationId === rootTweet.conversationId),
);
const repliesByParent = new Map<string, ParsedThreadTweet[]>();
for (const tweet of candidates) {
if (!tweet.inReplyTo || tweet.id === statusId) {
continue;
}
const bucket = repliesByParent.get(tweet.inReplyTo) ?? [];
bucket.push(tweet);
bucket.sort((left, right) => {
if (left.sortTimestamp !== right.sortTimestamp) {
return left.sortTimestamp - right.sortTimestamp;
}
return compareTweetIds(left.id, right.id);
});
repliesByParent.set(tweet.inReplyTo, bucket);
}
const ancestorPath: ParsedThreadTweet[] = [rootTweet];
const ancestorSeen = new Set<string>([rootTweet.id]);
let currentAncestor = rootTweet;
while (currentAncestor.inReplyTo) {
const parent = byId.get(currentAncestor.inReplyTo);
if (!parent || ancestorSeen.has(parent.id)) {
break;
}
ancestorPath.unshift(parent);
ancestorSeen.add(parent.id);
currentAncestor = parent;
}
const chain = ancestorPath.slice();
const seen = new Set<string>(chain.map((tweet) => tweet.id));
let currentId = rootTweet.id;
while (true) {
const next = (repliesByParent.get(currentId) ?? []).find((tweet) => !seen.has(tweet.id));
if (!next) {
break;
}
chain.push(next);
seen.add(next.id);
currentId = next.id;
}
return chain;
}
export function extractThreadTweetsFromPayloads(
payloads: unknown[],
statusId: string,
pageUrl: string,
): XTweet[] {
const parsedTweets: ParsedThreadTweet[] = [];
for (const payload of payloads) {
parsedTweets.push(...parseTweetDetailPayload(payload, pageUrl));
}
return buildContinuousThread(parsedTweets, statusId).map(({ sortTimestamp: _sortTimestamp, ...tweet }) => tweet);
}
function buildQuotedTweetMarkdown(quotedTweet: XQuotedTweet): string {
const author = quotedTweet.author ? `@${quotedTweet.author}` : "Unknown";
const name = quotedTweet.authorName ? `${quotedTweet.authorName} ` : "";
const lines: string[] = [`Quoted Tweet${quotedTweet.author || quotedTweet.authorName ? `: ${name}${author}`.trim() : ""}`];
if (quotedTweet.text) {
lines.push(...quotedTweet.text.split("\n"));
}
for (const mediaLine of formatMediaList(quotedTweet.media)) {
lines.push(mediaLine);
}
return lines.map((line) => (line ? `> ${line}` : ">")).join("\n");
}
function buildThreadMarkdown(tweets: XTweet[]): string {
return tweets
.map((tweet, index) => {
const lines: string[] = [];
const author = tweet.author ? `@${tweet.author}` : "Unknown";
const name = tweet.authorName ? `${tweet.authorName} ` : "";
lines.push(`## ${index + 1}. ${name}${author}`.trim());
if (tweet.createdAt) {
lines.push(`_Published: ${tweet.createdAt}_`);
}
lines.push(tweet.text || "(No text)");
const mediaLines = formatMediaList(tweet.media);
if (mediaLines.length > 0) {
lines.push(mediaLines.map((line) => `- ${line}`).join("\n"));
}
if (tweet.quotedTweet) {
lines.push(buildQuotedTweetMarkdown(tweet.quotedTweet));
}
return lines.join("\n\n");
})
.join("\n\n");
}
export function extractThreadDocumentFromPayloads(
payloads: unknown[],
statusId: string,
pageUrl: string,
): ExtractedDocument | null {
const tweets = extractThreadTweetsFromPayloads(payloads, statusId, pageUrl);
if (tweets.length <= 1) {
return null;
}
const rootTweet = tweets[0];
const rootAuthor = formatTweetAuthor(rootTweet);
return {
url: pageUrl,
canonicalUrl: rootTweet.url,
title: normalizeTitle(rootTweet.text, "X Thread"),
author: rootAuthor,
siteName: "X",
publishedAt: rootTweet.createdAt,
summary: rootTweet.text.slice(0, 200) || undefined,
adapter: "x",
metadata: {
kind: "x/thread",
tweetId: rootTweet.id,
tweetCount: tweets.length,
lastTweetId: tweets[tweets.length - 1]?.id,
...getTweetAuthorMetadata(rootTweet),
},
content: [{ type: "markdown", markdown: buildThreadMarkdown(tweets) }],
};
}
@@ -0,0 +1,36 @@
export type JsonObject = Record<string, unknown>;
export interface XUser {
name?: string;
screenName?: string;
}
export interface XMedia {
type: string;
url: string;
alt?: string;
}
export interface XQuotedTweet {
id: string;
author?: string;
authorName?: string;
text: string;
url: string;
media: XMedia[];
}
export interface XTweet {
id: string;
author?: string;
authorName?: string;
text: string;
likes: number;
retweets: number;
replies: number;
createdAt?: string;
inReplyTo?: string;
url: string;
media: XMedia[];
quotedTweet?: XQuotedTweet;
}
@@ -0,0 +1,33 @@
import type { Adapter } from "../types";
import { collectMediaFromDocument } from "../../media/markdown-media";
import { extractYouTubeTranscriptDocument } from "./transcript";
import { isYouTubeHost, parseYouTubeVideoId } from "./utils";
export const youtubeAdapter: Adapter = {
name: "youtube",
match(input) {
return isYouTubeHost(input.url.hostname);
},
async process(context) {
const videoId = parseYouTubeVideoId(context.input.url);
if (!videoId) {
return {
status: "no_document",
};
}
context.log.info(`Loading ${context.input.url.toString()} with youtube adapter`);
const document = await extractYouTubeTranscriptDocument(context, videoId);
if (!document) {
return {
status: "no_document",
};
}
return {
status: "ok",
document,
media: collectMediaFromDocument(document),
};
},
};
@@ -0,0 +1,392 @@
import type { ExtractedDocument } from "../../extract/document";
import { detectInteractionGate } from "../../browser/interaction-gates";
import {
buildYouTubeThumbnailCandidates,
parseYouTubeDescriptionChapters,
renderYouTubeTranscriptMarkdown,
type YouTubeChapter,
type YouTubeTranscriptSegment,
} from "./utils";
interface CaptionInfo {
captionUrl: string;
language: string;
kind: string;
available: string[];
title?: string;
author?: string;
authorUrl?: string;
channelId?: string;
description?: string;
publishedAt?: string;
viewCount?: number;
durationSeconds?: number;
keywords: string[];
category?: string;
isLiveContent?: boolean;
coverImages: string[];
}
function normalizeUrl(url: string | undefined): string | undefined {
if (!url) {
return undefined;
}
try {
const parsed = new URL(url);
if (parsed.protocol === "http:") {
parsed.protocol = "https:";
}
return parsed.toString();
} catch {
return url;
}
}
function buildSummary(description: string | undefined, segments: YouTubeTranscriptSegment[]): string | undefined {
const descriptionSummary = description
?.replace(/\r\n/g, "\n")
.split("\n")
.map((line) => line.trim())
.find((line) => line && !/^https?:\/\//i.test(line));
if (descriptionSummary) {
return descriptionSummary.slice(0, 240);
}
const transcriptSummary = segments
.slice(0, 8)
.map((segment) => segment.text)
.join(" ")
.slice(0, 240)
.trim();
return transcriptSummary || undefined;
}
async function canFetchThumbnail(url: string): Promise<boolean> {
try {
const response = await fetch(url, { method: "HEAD", redirect: "follow" });
if (response.ok) {
return true;
}
if (response.status === 405) {
const fallbackResponse = await fetch(url, {
method: "GET",
headers: { Range: "bytes=0-0" },
redirect: "follow",
});
return fallbackResponse.ok;
}
} catch {
return false;
}
return false;
}
async function resolveBestCoverImage(videoId: string, coverImages: string[]): Promise<string | undefined> {
const candidates = buildYouTubeThumbnailCandidates(videoId, coverImages);
for (const candidate of candidates) {
if (await canFetchThumbnail(candidate)) {
return candidate;
}
}
return candidates[0];
}
export async function extractYouTubeTranscriptDocument(
context: Parameters<import("../types").Adapter["process"]>[0],
videoId: string,
): Promise<ExtractedDocument | null> {
const videoUrl = `https://www.youtube.com/watch?v=${videoId}`;
await context.browser.goto(videoUrl, context.timeoutMs);
const interaction = await detectInteractionGate(context.browser);
if (interaction) {
context.log.debug(`Interaction gate detected on YouTube: ${interaction.provider}`);
return null;
}
try {
await context.network.waitForIdle({
idleMs: 1_000,
timeoutMs: Math.min(context.timeoutMs, 8_000),
});
} catch {
context.log.debug("Network idle timed out on YouTube load.");
}
const captionInfo = await context.browser.evaluate<CaptionInfo | { error: string }>(`
(async () => {
function readText(value) {
if (!value) return undefined;
if (typeof value === 'string') {
const text = value.trim();
return text || undefined;
}
if (typeof value.simpleText === 'string') {
const text = value.simpleText.trim();
return text || undefined;
}
if (Array.isArray(value.runs)) {
const text = value.runs
.map((run) => typeof run?.text === 'string' ? run.text : '')
.join('')
.trim();
return text || undefined;
}
return undefined;
}
function parsePositiveInteger(value) {
if (typeof value === 'number' && Number.isFinite(value) && value >= 0) {
return Math.floor(value);
}
if (typeof value !== 'string') {
return undefined;
}
const normalized = value.replace(/[^\\d]/g, '');
if (!normalized) {
return undefined;
}
const parsed = Number.parseInt(normalized, 10);
return Number.isFinite(parsed) ? parsed : undefined;
}
const apiKey = window.ytcfg?.data_?.INNERTUBE_API_KEY;
const playerResponse = window.ytInitialPlayerResponse;
const videoDetails = playerResponse?.videoDetails || {};
const microformat = playerResponse?.microformat?.playerMicroformatRenderer || {};
const title =
videoDetails.title ||
readText(microformat.title) ||
document.title.replace(/ - YouTube$/, '').trim();
const author =
videoDetails.author ||
microformat.ownerChannelName ||
document.querySelector('link[itemprop="name"]')?.getAttribute('content') ||
undefined;
const authorUrl =
microformat.ownerProfileUrl ||
(typeof videoDetails.channelId === 'string' && videoDetails.channelId
? 'https://www.youtube.com/channel/' + videoDetails.channelId
: undefined);
const description =
readText(microformat.description) ||
(typeof videoDetails.shortDescription === 'string' ? videoDetails.shortDescription.trim() : undefined);
const keywords = Array.isArray(videoDetails.keywords)
? videoDetails.keywords.filter((keyword) => typeof keyword === 'string' && keyword.trim())
: [];
const thumbnails = [
...(Array.isArray(videoDetails.thumbnail?.thumbnails) ? videoDetails.thumbnail.thumbnails : []),
...(Array.isArray(microformat.thumbnail?.thumbnails) ? microformat.thumbnail.thumbnails : []),
]
.filter((thumbnail) => typeof thumbnail?.url === 'string' && thumbnail.url)
.sort((left, right) => ((right?.width || 0) * (right?.height || 0)) - ((left?.width || 0) * (left?.height || 0)))
.map((thumbnail) => thumbnail.url);
if (!apiKey) {
return { error: 'INNERTUBE_API_KEY not found on page' };
}
const response = await fetch('/youtubei/v1/player?key=' + apiKey + '&prettyPrint=false', {
method: 'POST',
credentials: 'include',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
context: { client: { clientName: 'ANDROID', clientVersion: '20.10.38' } },
videoId: ${JSON.stringify(videoId)}
})
});
if (!response.ok) {
return { error: 'InnerTube player API returned HTTP ' + response.status };
}
const data = await response.json();
const renderer = data.captions?.playerCaptionsTracklistRenderer;
if (!renderer?.captionTracks?.length) {
return { error: 'No captions available for this video' };
}
const tracks = renderer.captionTracks;
const track = tracks.find((item) => item.kind !== 'asr') || tracks[0];
return {
captionUrl: track.baseUrl,
language: track.languageCode,
kind: track.kind || 'manual',
available: tracks.map((item) => {
const languageLabel = readText(item.name) || item.languageCode;
return item.kind === 'asr'
? languageLabel + ' [' + item.languageCode + ', auto]'
: languageLabel + ' [' + item.languageCode + ']';
}),
title,
author,
authorUrl,
channelId: typeof videoDetails.channelId === 'string' ? videoDetails.channelId : undefined,
description,
publishedAt:
(typeof microformat.publishDate === 'string' && microformat.publishDate) ||
(typeof microformat.uploadDate === 'string' && microformat.uploadDate) ||
document.querySelector('meta[itemprop="datePublished"]')?.getAttribute('content') ||
undefined,
viewCount: parsePositiveInteger(videoDetails.viewCount) ?? parsePositiveInteger(microformat.viewCount),
durationSeconds: parsePositiveInteger(videoDetails.lengthSeconds),
keywords,
category: typeof microformat.category === 'string' ? microformat.category : undefined,
isLiveContent: Boolean(videoDetails.isLiveContent || microformat.isLiveContent),
coverImages: thumbnails,
};
})()
`);
if ("error" in captionInfo) {
context.log.debug(`YouTube transcript unavailable: ${captionInfo.error}`);
return null;
}
const segments = await context.browser.evaluate<YouTubeTranscriptSegment[] | { error: string }>(`
(async () => {
const response = await fetch(${JSON.stringify(captionInfo.captionUrl)});
const xml = await response.text();
if (!xml) {
return { error: 'Caption XML is empty' };
}
function getAttr(tag, name) {
const needle = name + '="';
const index = tag.indexOf(needle);
if (index === -1) return '';
const valueStart = index + needle.length;
const valueEnd = tag.indexOf('"', valueStart);
if (valueEnd === -1) return '';
return tag.substring(valueStart, valueEnd);
}
function decodeEntities(value) {
return value
.replaceAll('&amp;', '&')
.replaceAll('&lt;', '<')
.replaceAll('&gt;', '>')
.replaceAll('&quot;', '"')
.replaceAll('&#39;', "'");
}
const marker = xml.includes('<p t="') ? '<p ' : '<text ';
const endMarker = marker === '<p ' ? '</p>' : '</text>';
const results = [];
let position = 0;
while (true) {
const tagStart = xml.indexOf(marker, position);
if (tagStart === -1) break;
let contentStart = xml.indexOf('>', tagStart);
if (contentStart === -1) break;
contentStart += 1;
const tagEnd = xml.indexOf(endMarker, contentStart);
if (tagEnd === -1) break;
const attrString = xml.substring(tagStart + marker.length, contentStart - 1);
const content = xml.substring(contentStart, tagEnd);
const start = marker === '<p '
? (parseFloat(getAttr(attrString, 't')) || 0) / 1000
: (parseFloat(getAttr(attrString, 'start')) || 0);
const duration = marker === '<p '
? (parseFloat(getAttr(attrString, 'd')) || 0) / 1000
: (parseFloat(getAttr(attrString, 'dur')) || 0);
const text = decodeEntities(content.replace(/<[^>]+>/g, '')).split('\\n').join(' ').trim();
if (text) {
results.push({ start, end: start + duration, text });
}
position = tagEnd + endMarker.length;
}
if (results.length === 0) {
return { error: 'Parsed 0 transcript segments' };
}
return results;
})()
`);
if (!Array.isArray(segments) || segments.length === 0) {
context.log.debug("Parsed no YouTube transcript segments.");
return null;
}
const extractedChapters = await context.browser.evaluate<YouTubeChapter[]>(`
(() => {
const data = window.ytInitialData;
const markers = data?.playerOverlays?.playerOverlayRenderer
?.decoratedPlayerBarRenderer?.decoratedPlayerBarRenderer
?.playerBar?.multiMarkersPlayerBarRenderer?.markersMap || [];
const results = [];
for (const marker of markers) {
const chapters = marker?.value?.chapters;
if (!Array.isArray(chapters)) continue;
for (const chapter of chapters) {
const renderer = chapter?.chapterRenderer;
const title = renderer?.title?.simpleText;
const timeRangeStartMillis = renderer?.timeRangeStartMillis;
if (title && typeof timeRangeStartMillis === 'number') {
results.push({ title, time: Math.floor(timeRangeStartMillis / 1000) });
}
}
}
return results;
})()
`).catch(() => []);
const descriptionChapters = parseYouTubeDescriptionChapters(captionInfo.description);
const chapters = extractedChapters.length > 0 ? extractedChapters : descriptionChapters;
const markdown = renderYouTubeTranscriptMarkdown({
description: captionInfo.description,
segments,
chapters,
});
if (!markdown) {
return null;
}
const pageUrl = await context.browser.getURL();
const coverImage = await resolveBestCoverImage(videoId, captionInfo.coverImages);
const summary = buildSummary(captionInfo.description, segments);
return {
url: pageUrl,
canonicalUrl: pageUrl,
title: captionInfo.title || "YouTube Transcript",
author: captionInfo.author,
publishedAt: captionInfo.publishedAt,
siteName: "YouTube",
summary,
adapter: "youtube",
metadata: {
kind: "youtube/transcript",
videoId,
authorUrl: normalizeUrl(captionInfo.authorUrl),
channelId: captionInfo.channelId,
coverImage,
description: captionInfo.description,
durationSeconds: captionInfo.durationSeconds,
language: captionInfo.language,
captionKind: captionInfo.kind,
availableLanguages: captionInfo.available,
viewCount: captionInfo.viewCount,
keywords: captionInfo.keywords,
category: captionInfo.category,
isLiveContent: captionInfo.isLiveContent,
chapterCount: chapters.length,
},
content: [{ type: "markdown", markdown }],
};
}
@@ -0,0 +1,253 @@
export interface YouTubeTranscriptSegment {
start: number;
end: number;
text: string;
}
export interface YouTubeChapter {
title: string;
time: number;
}
interface RenderYouTubeTranscriptMarkdownInput {
description?: string;
segments: YouTubeTranscriptSegment[];
chapters: YouTubeChapter[];
}
const DESCRIPTION_CHAPTER_RE = /^((?:\d{1,2}:)?\d{1,2}:\d{2})(?:\s+[-|:]\s+|\s+)(.+)$/;
const YOUTUBE_THUMBNAIL_VARIANTS = [
"maxresdefault.jpg",
"sddefault.jpg",
"hqdefault.jpg",
"mqdefault.jpg",
"default.jpg",
];
export function isYouTubeHost(hostname: string): boolean {
return [
"youtube.com",
"www.youtube.com",
"m.youtube.com",
"youtu.be",
].includes(hostname);
}
export function parseYouTubeVideoId(url: URL): string | null {
if (url.hostname === "youtu.be") {
return url.pathname.split("/").filter(Boolean)[0] ?? null;
}
if (url.pathname === "/watch") {
return url.searchParams.get("v");
}
const shortsMatch = url.pathname.match(/^\/shorts\/([^/?#]+)/);
if (shortsMatch) {
return shortsMatch[1];
}
const liveMatch = url.pathname.match(/^\/live\/([^/?#]+)/);
if (liveMatch) {
return liveMatch[1];
}
return null;
}
function parseTimestampValue(raw: string): number | null {
const parts = raw
.split(":")
.map((part) => Number.parseInt(part, 10))
.filter((part) => Number.isFinite(part));
if (parts.length < 2 || parts.length > 3) {
return null;
}
if (parts.some((part) => part < 0)) {
return null;
}
if (parts.length === 2) {
const [minutes, seconds] = parts;
return minutes * 60 + seconds;
}
const [hours, minutes, seconds] = parts;
return hours * 3600 + minutes * 60 + seconds;
}
export function formatTimestamp(totalSeconds: number): string {
const rounded = Math.max(0, Math.floor(totalSeconds));
const hours = Math.floor(rounded / 3600);
const minutes = Math.floor((rounded % 3600) / 60);
const seconds = rounded % 60;
if (hours > 0) {
return `${hours}:${String(minutes).padStart(2, "0")}:${String(seconds).padStart(2, "0")}`;
}
return `${minutes}:${String(seconds).padStart(2, "0")}`;
}
export function formatTimestampRange(start: number, end: number): string {
const safeStart = Math.max(0, start);
const safeEnd = Math.max(safeStart, end);
return `[${formatTimestamp(safeStart)} -> ${formatTimestamp(safeEnd)}]`;
}
export function normalizeYouTubeChapters(chapters: YouTubeChapter[]): YouTubeChapter[] {
const seenTimes = new Set<number>();
return chapters
.map((chapter) => ({
title: chapter.title.trim(),
time: Math.max(0, Math.floor(chapter.time)),
}))
.filter((chapter) => chapter.title)
.sort((left, right) => left.time - right.time)
.filter((chapter) => {
if (seenTimes.has(chapter.time)) {
return false;
}
seenTimes.add(chapter.time);
return true;
});
}
export function parseYouTubeDescriptionChapters(description?: string | null): YouTubeChapter[] {
if (!description) {
return [];
}
const chapters: YouTubeChapter[] = [];
const seen = new Set<string>();
for (const rawLine of description.replace(/\r\n/g, "\n").split("\n")) {
const line = rawLine.trim();
if (!line) {
continue;
}
const match = line.match(DESCRIPTION_CHAPTER_RE);
if (!match) {
continue;
}
const time = parseTimestampValue(match[1]);
const title = match[2]?.trim();
if (time === null || !title) {
continue;
}
const key = `${time}:${title.toLowerCase()}`;
if (seen.has(key)) {
continue;
}
seen.add(key);
chapters.push({ title, time });
}
const normalized = normalizeYouTubeChapters(chapters);
if (normalized.length >= 2) {
return normalized;
}
if (normalized.length === 1 && normalized[0]?.time === 0) {
return normalized;
}
return [];
}
function renderDescriptionMarkdown(description: string): string {
return description
.replace(/\r\n/g, "\n")
.trim()
.split(/\n{2,}/)
.map((block) => block.split("\n").map((line) => line.trimEnd()).join(" \n"))
.join("\n\n")
.trim();
}
function renderSegmentLine(segment: YouTubeTranscriptSegment): string {
return `${formatTimestampRange(segment.start, segment.end)} ${segment.text}`;
}
export function renderYouTubeTranscriptMarkdown({
description,
segments,
chapters,
}: RenderYouTubeTranscriptMarkdownInput): string {
if (segments.length === 0) {
return "";
}
const parts: string[] = [];
const normalizedDescription = description?.trim();
const transcriptEnd = segments.reduce((maxEnd, segment) => Math.max(maxEnd, segment.end, segment.start), 0);
const normalizedChapters = normalizeYouTubeChapters(chapters).filter(
(chapter) => transcriptEnd <= 0 || chapter.time < transcriptEnd,
);
if (normalizedDescription) {
parts.push("## Description");
parts.push(renderDescriptionMarkdown(normalizedDescription));
}
if (normalizedChapters.length > 0) {
parts.push("## Chapters");
for (let index = 0; index < normalizedChapters.length; index += 1) {
const chapter = normalizedChapters[index];
const nextChapter = normalizedChapters[index + 1];
const chapterEnd = nextChapter ? nextChapter.time : transcriptEnd;
const chapterSegments = segments.filter(
(segment) => segment.start >= chapter.time && segment.start < chapterEnd,
);
parts.push(`### ${chapter.title} ${formatTimestampRange(chapter.time, chapterEnd)}`);
if (chapterSegments.length > 0) {
parts.push(chapterSegments.map(renderSegmentLine).join("\n"));
}
}
} else {
parts.push("## Transcript");
parts.push(segments.map(renderSegmentLine).join("\n"));
}
return parts.filter(Boolean).join("\n\n").trim();
}
function normalizeThumbnailKey(url: string): string {
try {
const parsed = new URL(url);
return `${parsed.origin}${parsed.pathname}`;
} catch {
return url;
}
}
export function buildYouTubeThumbnailCandidates(videoId: string, listedUrls: string[]): string[] {
const candidates = [
...YOUTUBE_THUMBNAIL_VARIANTS.map((variant) => `https://i.ytimg.com/vi/${videoId}/${variant}`),
...listedUrls,
];
const seen = new Set<string>();
return candidates.filter((candidate) => {
if (!candidate) {
return false;
}
const key = normalizeThumbnailKey(candidate);
if (seen.has(key)) {
return false;
}
seen.add(key);
return true;
});
}
@@ -0,0 +1,258 @@
import { EventEmitter } from "node:events";
import WebSocket from "ws";
type JsonObject = Record<string, unknown>;
interface CdpPendingCommand {
resolve(value: unknown): void;
reject(error: unknown): void;
method: string;
}
interface CdpErrorShape {
message?: string;
}
interface CdpCommandResult<T> {
result?: T;
error?: CdpErrorShape;
}
interface CreatePageSessionOptions {
initialUrl?: string;
visible?: boolean;
}
export class TargetSession extends EventEmitter {
constructor(
private readonly client: CdpClient,
public readonly targetId: string,
public readonly sessionId: string,
) {
super();
}
async send<T>(method: string, params: JsonObject = {}): Promise<T> {
return this.client.sendSessionCommand<T>(this.sessionId, method, params);
}
handleEvent(method: string, params: JsonObject): void {
this.emit(method, params);
this.emit("event", { method, params });
}
async waitForEvent<T extends JsonObject>(
method: string,
predicate?: (params: T) => boolean,
timeoutMs = 30_000,
): Promise<T> {
return new Promise<T>((resolve, reject) => {
const timeout = setTimeout(() => {
this.off(method, listener);
reject(new Error(`Timed out waiting for ${method}`));
}, timeoutMs);
const listener = (params: T): void => {
if (predicate && !predicate(params)) {
return;
}
clearTimeout(timeout);
this.off(method, listener);
resolve(params);
};
this.on(method, listener);
});
}
}
export class CdpClient {
private readonly ws: WebSocket;
private readonly pending = new Map<number, CdpPendingCommand>();
private readonly sessions = new Map<string, TargetSession>();
private nextId = 1;
private constructor(ws: WebSocket) {
this.ws = ws;
this.ws.on("message", (raw) => {
this.handleMessage(raw.toString());
});
}
static async connect(browserWsUrl: string): Promise<CdpClient> {
const ws = await new Promise<WebSocket>((resolve, reject) => {
const socket = new WebSocket(browserWsUrl);
socket.once("open", () => resolve(socket));
socket.once("error", (error) => reject(error));
});
return new CdpClient(ws);
}
private handleMessage(rawMessage: string): void {
const message = JSON.parse(rawMessage) as {
id?: number;
sessionId?: string;
method?: string;
params?: JsonObject;
result?: unknown;
error?: CdpErrorShape;
};
if (typeof message.id === "number") {
const pending = this.pending.get(message.id);
if (!pending) {
return;
}
this.pending.delete(message.id);
if (message.error) {
pending.reject(new Error(`${pending.method}: ${message.error.message ?? "Unknown CDP error"}`));
return;
}
pending.resolve(message.result);
return;
}
if (typeof message.sessionId === "string" && typeof message.method === "string") {
const session = this.sessions.get(message.sessionId);
if (session) {
session.handleEvent(message.method, (message.params ?? {}) as JsonObject);
}
}
}
private async sendCommand<T>(
method: string,
params: JsonObject = {},
sessionId?: string,
): Promise<T> {
const id = this.nextId;
this.nextId += 1;
const payload = sessionId ? { id, method, params, sessionId } : { id, method, params };
const result = new Promise<T>((resolve, reject) => {
this.pending.set(id, {
resolve: (value) => resolve(value as T),
reject,
method,
});
});
this.ws.send(JSON.stringify(payload));
return result;
}
async sendBrowserCommand<T>(method: string, params: JsonObject = {}): Promise<T> {
return this.sendCommand<T>(method, params);
}
async sendSessionCommand<T>(sessionId: string, method: string, params: JsonObject = {}): Promise<T> {
return this.sendCommand<T>(method, params, sessionId);
}
private async createPageTarget(initialUrl: string, visible = false): Promise<{ targetId: string }> {
const attempts: JsonObject[] = visible
? [
{
url: initialUrl,
newWindow: true,
focus: true,
},
{
url: initialUrl,
focus: true,
},
{
url: initialUrl,
},
]
: [
{
url: initialUrl,
hidden: true,
},
{
url: initialUrl,
background: true,
focus: false,
},
{
url: initialUrl,
},
];
let lastError: unknown;
for (const params of attempts) {
try {
return await this.sendBrowserCommand<{ targetId: string }>("Target.createTarget", params);
} catch (error) {
lastError = error;
}
}
throw lastError instanceof Error ? lastError : new Error("Target.createTarget failed");
}
async createPageSession(options: CreatePageSessionOptions = {}): Promise<TargetSession> {
const initialUrl = options.initialUrl ?? "about:blank";
const created = await this.createPageTarget(initialUrl, Boolean(options.visible));
const attached = await this.sendBrowserCommand<{ sessionId: string }>("Target.attachToTarget", {
targetId: created.targetId,
flatten: true,
});
const session = new TargetSession(this, created.targetId, attached.sessionId);
this.sessions.set(attached.sessionId, session);
if (options.visible) {
await this.sendBrowserCommand("Target.activateTarget", {
targetId: created.targetId,
}).catch(() => {});
}
await session.send("Page.enable");
await session.send("Runtime.enable");
await session.send("DOM.enable");
if (options.visible) {
await session.send("Page.bringToFront").catch(() => {});
}
return session;
}
async closeTarget(targetId: string): Promise<void> {
try {
await this.sendBrowserCommand("Target.closeTarget", { targetId });
} catch {
// Target may already be gone.
}
}
async close(): Promise<void> {
await new Promise<void>((resolve) => {
if (this.ws.readyState === WebSocket.CLOSED) {
resolve();
return;
}
this.ws.once("close", () => resolve());
this.ws.close();
});
}
}
export async function evaluateRuntime<T>(session: TargetSession, expression: string): Promise<T> {
const response = await session.send<CdpCommandResult<{ value?: T; description?: string }>>("Runtime.evaluate", {
expression,
awaitPromise: true,
returnByValue: true,
});
if (response.error) {
throw new Error(response.error.message ?? "Runtime.evaluate failed");
}
return (response.result?.value as T | undefined) ?? (undefined as T);
}
@@ -0,0 +1,117 @@
import { launch, type LaunchedChrome } from "chrome-launcher";
import type { Logger } from "../utils/logger";
import { ensureChromeProfileDir, findExistingChromeDebugPort, resolveChromeProfileDir } from "./profile";
interface ChromeVersionResponse {
webSocketDebuggerUrl: string;
}
export interface ChromeConnectOptions {
cdpUrl?: string;
browserPath?: string;
headless?: boolean;
logger?: Logger;
profileDir?: string;
}
export interface ChromeConnection {
browserWsUrl: string;
origin?: string;
port?: number;
profileDir?: string;
launched: boolean;
close(): Promise<void>;
}
async function fetchJson<T>(url: string): Promise<T> {
const response = await fetch(url);
if (!response.ok) {
throw new Error(`Failed to fetch ${url}: HTTP ${response.status}`);
}
return (await response.json()) as T;
}
async function connectToHttpEndpoint(origin: string): Promise<ChromeConnection> {
const normalizedOrigin = origin.replace(/\/$/, "");
const version = await fetchJson<ChromeVersionResponse>(`${normalizedOrigin}/json/version`);
return {
browserWsUrl: version.webSocketDebuggerUrl,
origin: normalizedOrigin,
port: Number(new URL(normalizedOrigin).port || 80),
launched: false,
async close() {
// Reused external Chrome, nothing to close here.
},
};
}
async function tryReuseChrome(profileDir: string, logger?: Logger): Promise<ChromeConnection | null> {
const port = await findExistingChromeDebugPort({ profileDir });
if (!port) {
return null;
}
const origin = `http://127.0.0.1:${port}`;
try {
const connection = await connectToHttpEndpoint(origin);
logger?.info(`Reusing Chrome debugger at ${origin} for profile ${profileDir}`);
return {
...connection,
profileDir,
};
} catch {
// Debugger disappeared between detection and connect.
}
return null;
}
export async function connectChrome(options: ChromeConnectOptions): Promise<ChromeConnection> {
if (options.cdpUrl) {
if (options.cdpUrl.startsWith("ws://") || options.cdpUrl.startsWith("wss://")) {
return {
browserWsUrl: options.cdpUrl,
launched: false,
async close() {},
};
}
return connectToHttpEndpoint(options.cdpUrl);
}
const profileDir = ensureChromeProfileDir(resolveChromeProfileDir(options.profileDir));
const reused = await tryReuseChrome(profileDir, options.logger);
if (reused) {
return reused;
}
options.logger?.warn(`No running Chrome debugger found for profile ${profileDir}. Launching Chrome with that profile.`);
const launchedChrome: LaunchedChrome = await launch({
chromePath: options.browserPath,
userDataDir: profileDir,
chromeFlags: [
"--disable-background-networking",
"--disable-default-apps",
"--disable-popup-blocking",
"--disable-sync",
"--no-first-run",
"--no-default-browser-check",
"--remote-allow-origins=*",
...(!options.headless ? ["--no-startup-window"] : []),
...(options.headless ? ["--headless=new"] : []),
],
});
const origin = `http://127.0.0.1:${launchedChrome.port}`;
const version = await fetchJson<ChromeVersionResponse>(`${origin}/json/version`);
return {
browserWsUrl: version.webSocketDebuggerUrl,
origin,
port: launchedChrome.port,
profileDir,
launched: true,
async close() {
launchedChrome.kill();
},
};
}
@@ -0,0 +1,123 @@
import type { WaitForInteractionRequest } from "../adapters/types";
import type { BrowserSession } from "./session";
interface GateSnapshot {
title: string;
currentUrl: string;
bodyText: string;
hasCloudflareTurnstile: boolean;
hasCloudflareChallenge: boolean;
hasRecaptcha: boolean;
hasRecaptchaIframe: boolean;
hasHcaptcha: boolean;
hasHcaptchaIframe: boolean;
}
export function detectInteractionGateFromSnapshot(snapshot: GateSnapshot): WaitForInteractionRequest | null {
const text = snapshot.bodyText.toLowerCase();
const title = snapshot.title.toLowerCase();
const url = snapshot.currentUrl.toLowerCase();
if (
snapshot.hasCloudflareTurnstile ||
snapshot.hasCloudflareChallenge ||
title.includes("just a moment") ||
text.includes("verify you are human") ||
text.includes("checking your browser before accessing") ||
text.includes("enable javascript and cookies to continue") ||
url.includes("/cdn-cgi/challenge-platform/")
) {
return {
type: "wait_for_interaction",
kind: "cloudflare",
provider: "cloudflare",
reason: "Cloudflare human verification detected",
prompt: "Please complete the Cloudflare verification in the opened Chrome window. Extraction will continue automatically once the challenge disappears.",
requiresVisibleBrowser: true,
};
}
if (
snapshot.hasRecaptcha ||
snapshot.hasRecaptchaIframe ||
text.includes("i'm not a robot") ||
text.includes("recaptcha")
) {
return {
type: "wait_for_interaction",
kind: "recaptcha",
provider: "google_recaptcha",
reason: "Google reCAPTCHA detected",
prompt: "Please complete the reCAPTCHA verification in the opened Chrome window. Extraction will continue automatically once the challenge disappears.",
requiresVisibleBrowser: true,
};
}
if (
snapshot.hasHcaptcha ||
snapshot.hasHcaptchaIframe ||
text.includes("hcaptcha")
) {
return {
type: "wait_for_interaction",
kind: "hcaptcha",
provider: "hcaptcha",
reason: "hCaptcha verification detected",
prompt: "Please complete the hCaptcha verification in the opened Chrome window. Extraction will continue automatically once the challenge disappears.",
requiresVisibleBrowser: true,
};
}
return null;
}
export async function detectInteractionGate(browser: BrowserSession): Promise<WaitForInteractionRequest | null> {
const snapshot = await browser.evaluate<GateSnapshot>(`
(() => {
const bodyText = (document.body?.innerText ?? "").slice(0, 4000);
return {
title: document.title ?? "",
currentUrl: window.location.href,
bodyText,
hasCloudflareTurnstile: Boolean(
document.querySelector(
'.cf-turnstile, [name="cf-turnstile-response"], iframe[src*="challenges.cloudflare.com"]'
)
),
hasCloudflareChallenge: Boolean(
document.querySelector(
'#challenge-running, #cf-challenge-running, .challenge-platform, [data-ray], [data-translate="checking_browser"]'
)
),
hasRecaptcha: Boolean(
document.querySelector(
'.g-recaptcha, textarea[name="g-recaptcha-response"], iframe[title*="reCAPTCHA"]'
)
),
hasRecaptchaIframe: Boolean(
document.querySelector('iframe[src*="google.com/recaptcha"], iframe[src*="recaptcha/api2"]')
),
hasHcaptcha: Boolean(
document.querySelector(
'.h-captcha, textarea[name="h-captcha-response"], iframe[title*="hCaptcha"]'
)
),
hasHcaptchaIframe: Boolean(
document.querySelector('iframe[src*="hcaptcha.com"]')
),
};
})()
`).catch(() => ({
title: "",
currentUrl: "",
bodyText: "",
hasCloudflareTurnstile: false,
hasCloudflareChallenge: false,
hasRecaptcha: false,
hasRecaptchaIframe: false,
hasHcaptcha: false,
hasHcaptchaIframe: false,
}));
return detectInteractionGateFromSnapshot(snapshot);
}
@@ -0,0 +1,235 @@
import type { TargetSession } from "./cdp-client";
import type { Logger } from "../utils/logger";
type JsonObject = Record<string, unknown>;
export interface NetworkEntry {
requestId: string;
url: string;
method: string;
resourceType: string;
timestamp: number;
requestHeaders?: Record<string, string>;
requestBody?: string;
status?: number;
statusText?: string;
responseHeaders?: Record<string, string>;
mimeType?: string;
body?: string;
bodyBase64?: boolean;
bodyError?: string;
failed?: boolean;
failureReason?: string;
finished: boolean;
}
function normalizeHeaders(headers: unknown): Record<string, string> | undefined {
if (!headers || typeof headers !== "object") {
return undefined;
}
return Object.fromEntries(
Object.entries(headers as Record<string, unknown>).map(([key, value]) => [key, String(value)]),
);
}
function sleep(ms: number): Promise<void> {
return new Promise((resolve) => setTimeout(resolve, ms));
}
export class NetworkJournal {
private readonly entries = new Map<string, NetworkEntry>();
private lastActivityAt = Date.now();
private started = false;
constructor(
private readonly session: TargetSession,
private readonly log: Logger,
) {}
async start(): Promise<void> {
if (this.started) {
return;
}
this.started = true;
this.session.on("Network.requestWillBeSent", this.handleRequestWillBeSent);
this.session.on("Network.responseReceived", this.handleResponseReceived);
this.session.on("Network.loadingFinished", this.handleLoadingFinished);
this.session.on("Network.loadingFailed", this.handleLoadingFailed);
await this.session.send("Network.enable");
}
stop(): void {
if (!this.started) {
return;
}
this.session.off("Network.requestWillBeSent", this.handleRequestWillBeSent);
this.session.off("Network.responseReceived", this.handleResponseReceived);
this.session.off("Network.loadingFinished", this.handleLoadingFinished);
this.session.off("Network.loadingFailed", this.handleLoadingFailed);
this.started = false;
}
private touch(): void {
this.lastActivityAt = Date.now();
}
private readonly handleRequestWillBeSent = (params: JsonObject): void => {
const requestId = typeof params.requestId === "string" ? params.requestId : undefined;
const request = params.request as JsonObject | undefined;
if (!requestId || !request) {
return;
}
this.touch();
this.entries.set(requestId, {
requestId,
url: String(request.url ?? ""),
method: String(request.method ?? "GET"),
resourceType: String(params.type ?? "Other"),
timestamp: Date.now(),
requestHeaders: normalizeHeaders(request.headers),
requestBody: typeof request.postData === "string" ? request.postData : undefined,
finished: false,
});
};
private readonly handleResponseReceived = (params: JsonObject): void => {
const requestId = typeof params.requestId === "string" ? params.requestId : undefined;
const response = params.response as JsonObject | undefined;
if (!requestId || !response) {
return;
}
this.touch();
const existing = this.entries.get(requestId);
if (!existing) {
return;
}
existing.status = typeof response.status === "number" ? response.status : undefined;
existing.statusText = typeof response.statusText === "string" ? response.statusText : undefined;
existing.responseHeaders = normalizeHeaders(response.headers);
existing.mimeType = typeof response.mimeType === "string" ? response.mimeType : undefined;
this.entries.set(requestId, existing);
};
private readonly handleLoadingFinished = (params: JsonObject): void => {
const requestId = typeof params.requestId === "string" ? params.requestId : undefined;
if (!requestId) {
return;
}
this.touch();
const existing = this.entries.get(requestId);
if (!existing) {
return;
}
existing.finished = true;
this.entries.set(requestId, existing);
};
private readonly handleLoadingFailed = (params: JsonObject): void => {
const requestId = typeof params.requestId === "string" ? params.requestId : undefined;
if (!requestId) {
return;
}
this.touch();
const existing = this.entries.get(requestId);
if (!existing) {
return;
}
existing.finished = true;
existing.failed = true;
existing.failureReason = typeof params.errorText === "string" ? params.errorText : "Unknown error";
this.entries.set(requestId, existing);
};
getEntries(): NetworkEntry[] {
return Array.from(this.entries.values());
}
findEntries(predicate: (entry: NetworkEntry) => boolean): NetworkEntry[] {
return this.getEntries().filter(predicate);
}
async waitForIdle(options: { idleMs?: number; timeoutMs?: number } = {}): Promise<void> {
const idleMs = options.idleMs ?? 1_200;
const timeoutMs = options.timeoutMs ?? 15_000;
const startedAt = Date.now();
while (Date.now() - startedAt < timeoutMs) {
if (Date.now() - this.lastActivityAt >= idleMs) {
return;
}
await sleep(Math.min(150, idleMs));
}
throw new Error("Timed out waiting for network idle");
}
async waitForResponse(
predicate: (entry: NetworkEntry) => boolean,
options: { timeoutMs?: number } = {},
): Promise<NetworkEntry> {
const timeoutMs = options.timeoutMs ?? 10_000;
const startedAt = Date.now();
while (Date.now() - startedAt < timeoutMs) {
const matched = this.getEntries().find((entry) => entry.finished && predicate(entry));
if (matched) {
return matched;
}
await sleep(150);
}
throw new Error("Timed out waiting for matching network response");
}
async ensureBody(entry: NetworkEntry): Promise<string | undefined> {
if (entry.body !== undefined) {
return entry.body;
}
if (entry.bodyError || entry.failed || !entry.finished) {
return undefined;
}
try {
const result = await this.session.send<{ body: string; base64Encoded: boolean }>("Network.getResponseBody", {
requestId: entry.requestId,
});
entry.bodyBase64 = result.base64Encoded;
entry.body = result.base64Encoded ? Buffer.from(result.body, "base64").toString("utf8") : result.body;
return entry.body;
} catch (error) {
entry.bodyError = error instanceof Error ? error.message : String(error);
this.log.debug(`Failed to fetch response body for ${entry.url}: ${entry.bodyError}`);
return undefined;
}
}
async getJsonBody(entry: NetworkEntry): Promise<unknown | null> {
const body = await this.ensureBody(entry);
if (!body) {
return null;
}
try {
return JSON.parse(body);
} catch {
return null;
}
}
async toJSON(options: { includeBodies?: boolean } = {}): Promise<NetworkEntry[]> {
const entries = this.getEntries();
if (!options.includeBodies) {
return entries;
}
await Promise.all(entries.map((entry) => this.ensureBody(entry)));
return entries;
}
}
@@ -0,0 +1,105 @@
import type { BrowserSession } from "./session";
export interface CapturedPageSnapshot {
html: string;
finalUrl: string;
}
export const CAPTURE_NORMALIZED_PAGE_SCRIPT = String.raw`
(() => {
const baseUrl = document.baseURI || location.href;
const htmlClone = document.documentElement.cloneNode(true);
function materializeShadowDom(sourceRoot, cloneRoot) {
const sourceElements = Array.from(sourceRoot.querySelectorAll("*"));
const cloneElements = Array.from(cloneRoot.querySelectorAll("*"));
for (let index = sourceElements.length - 1; index >= 0; index -= 1) {
const sourceElement = sourceElements[index];
const cloneElement = cloneElements[index];
const shadowRoot = sourceElement && sourceElement.shadowRoot;
if (!shadowRoot || !cloneElement || !shadowRoot.innerHTML) {
continue;
}
if (cloneElement.tagName && cloneElement.tagName.includes("-")) {
const wrapper = document.createElement("div");
wrapper.setAttribute("data-shadow-host", cloneElement.tagName.toLowerCase());
wrapper.innerHTML = shadowRoot.innerHTML;
cloneElement.replaceWith(wrapper);
} else {
cloneElement.innerHTML = shadowRoot.innerHTML;
}
}
}
function toAbsolute(url) {
if (!url) return url;
try {
return new URL(url, baseUrl).href;
} catch {
return url;
}
}
function absolutizeAttribute(root, selector, attribute) {
root.querySelectorAll(selector).forEach((element) => {
const value = element.getAttribute(attribute);
if (!value) return;
const absolute = toAbsolute(value);
if (absolute) {
element.setAttribute(attribute, absolute);
}
});
}
function absolutizeSrcset(root, selector) {
root.querySelectorAll(selector).forEach((element) => {
const srcset = element.getAttribute("srcset");
if (!srcset) return;
element.setAttribute(
"srcset",
srcset
.split(",")
.map((part) => {
const trimmed = part.trim();
if (!trimmed) return "";
const [url, ...descriptor] = trimmed.split(/\s+/);
const absolute = toAbsolute(url);
return descriptor.length > 0 ? absolute + " " + descriptor.join(" ") : absolute;
})
.filter(Boolean)
.join(", "),
);
});
}
materializeShadowDom(document.documentElement, htmlClone);
htmlClone
.querySelectorAll("img[data-src], video[data-src], audio[data-src], source[data-src]")
.forEach((element) => {
const dataSource = element.getAttribute("data-src");
const current = element.getAttribute("src");
if (dataSource && (!current || current === "" || current.startsWith("data:"))) {
element.setAttribute("src", dataSource);
}
});
absolutizeAttribute(htmlClone, "a[href]", "href");
absolutizeAttribute(htmlClone, "img[src], video[src], audio[src], source[src], iframe[src]", "src");
absolutizeAttribute(htmlClone, "video[poster]", "poster");
absolutizeSrcset(htmlClone, "img[srcset], source[srcset]");
return {
html: "<!doctype html>\n" + htmlClone.outerHTML,
finalUrl: location.href,
};
})()
`;
export async function captureNormalizedPageSnapshot(
browser: BrowserSession,
): Promise<CapturedPageSnapshot> {
return browser.evaluate<CapturedPageSnapshot>(CAPTURE_NORMALIZED_PAGE_SCRIPT);
}
@@ -0,0 +1,148 @@
import fs from "node:fs";
import os from "node:os";
import path from "node:path";
import process from "node:process";
import { spawnSync } from "node:child_process";
export interface ResolveSharedChromeProfileDirOptions {
envNames?: string[];
appDataDirName?: string;
profileDirName?: string;
}
export interface FindExistingChromeDebugPortOptions {
profileDir: string;
timeoutMs?: number;
}
interface ChromeVersionResponse {
webSocketDebuggerUrl?: string;
}
function resolveDataBaseDir(): string {
if (process.platform === "darwin") {
return path.join(os.homedir(), "Library", "Application Support");
}
if (process.platform === "win32") {
return process.env.APPDATA ?? path.join(os.homedir(), "AppData", "Roaming");
}
return process.env.XDG_DATA_HOME ?? path.join(os.homedir(), ".local", "share");
}
export function resolveSharedChromeProfileDir(
options: ResolveSharedChromeProfileDirOptions = {},
): string {
for (const envName of options.envNames ?? []) {
const override = process.env[envName]?.trim();
if (override) {
return path.resolve(override);
}
}
const appDataDirName = options.appDataDirName ?? "baoyu-skills";
const profileDirName = options.profileDirName ?? "chrome-profile";
return path.join(resolveDataBaseDir(), appDataDirName, profileDirName);
}
export function resolveChromeProfileDir(profileDir?: string): string {
if (profileDir?.trim()) {
return path.resolve(profileDir.trim());
}
return resolveSharedChromeProfileDir({
envNames: ["BAOYU_CHROME_PROFILE_DIR"],
appDataDirName: "baoyu-skills",
profileDirName: "chrome-profile",
});
}
export function ensureChromeProfileDir(profileDir: string): string {
fs.mkdirSync(profileDir, { recursive: true });
return profileDir;
}
async function fetchWithTimeout(url: string, timeoutMs = 3_000): Promise<Response> {
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), timeoutMs);
try {
return await fetch(url, {
redirect: "follow",
signal: controller.signal,
});
} finally {
clearTimeout(timer);
}
}
async function fetchJson<T>(url: string, timeoutMs = 3_000): Promise<T> {
const response = await fetchWithTimeout(url, timeoutMs);
if (!response.ok) {
throw new Error(`Request failed: ${response.status} ${response.statusText}`);
}
return (await response.json()) as T;
}
async function isDebugPortReady(port: number, timeoutMs = 3_000): Promise<boolean> {
try {
const version = await fetchJson<ChromeVersionResponse>(`http://127.0.0.1:${port}/json/version`, timeoutMs);
return Boolean(version.webSocketDebuggerUrl);
} catch {
return false;
}
}
function parseDevToolsActivePort(filePath: string): { port: number; wsPath: string } | null {
try {
const content = fs.readFileSync(filePath, "utf8");
const lines = content.split(/\r?\n/);
const port = Number.parseInt(lines[0]?.trim() ?? "", 10);
const wsPath = lines[1]?.trim() ?? "";
if (port > 0 && wsPath) {
return { port, wsPath };
}
} catch {
// Ignore and fall back to process inspection.
}
return null;
}
export async function findExistingChromeDebugPort(
options: FindExistingChromeDebugPortOptions,
): Promise<number | null> {
const timeoutMs = options.timeoutMs ?? 3_000;
const activePort = parseDevToolsActivePort(path.join(options.profileDir, "DevToolsActivePort"));
if (activePort && await isDebugPortReady(activePort.port, timeoutMs)) {
return activePort.port;
}
if (process.platform === "win32") {
return null;
}
try {
const result = spawnSync("ps", ["aux"], {
encoding: "utf8",
timeout: 5_000,
});
if (result.status !== 0 || !result.stdout) {
return null;
}
const lines = result.stdout
.split("\n")
.filter((line) => line.includes(options.profileDir) && line.includes("--remote-debugging-port="));
for (const line of lines) {
const match = line.match(/--remote-debugging-port=(\d+)/);
const port = Number.parseInt(match?.[1] ?? "", 10);
if (port > 0 && await isDebugPortReady(port, timeoutMs)) {
return port;
}
}
} catch {
// Ignore and report no reusable debugger.
}
return null;
}
@@ -0,0 +1,155 @@
import { execFile } from "node:child_process";
import { promisify } from "node:util";
import { CdpClient, TargetSession, evaluateRuntime } from "./cdp-client";
interface NavigationResult {
errorText?: string;
}
const execFileAsync = promisify(execFile);
const MACOS_BROWSER_APP_IDS = [
"com.google.Chrome",
"org.chromium.Chromium",
"com.brave.Browser",
"com.microsoft.edgemac",
];
function sleep(ms: number): Promise<void> {
return new Promise((resolve) => setTimeout(resolve, ms));
}
async function activateBrowserApp(): Promise<void> {
if (process.platform !== "darwin") {
return;
}
for (const appId of MACOS_BROWSER_APP_IDS) {
try {
await execFileAsync("osascript", ["-e", `tell application id "${appId}" to activate`]);
return;
} catch {
// Try the next installed browser bundle id.
}
}
}
export class BrowserSession {
private constructor(
private readonly cdp: CdpClient,
public readonly targetSession: TargetSession,
public readonly interactive: boolean,
) {}
static async open(
cdp: CdpClient,
options: {
initialUrl?: string;
interactive?: boolean;
} = {},
): Promise<BrowserSession> {
const targetSession = await cdp.createPageSession({
initialUrl: options.initialUrl,
visible: options.interactive,
});
const browser = new BrowserSession(cdp, targetSession, Boolean(options.interactive));
if (browser.interactive) {
await browser.bringToFront().catch(() => {});
}
return browser;
}
async goto(url: string, timeoutMs = 30_000): Promise<void> {
const loadPromise = this.targetSession.waitForEvent("Page.loadEventFired", undefined, timeoutMs).catch(() => null);
const result = await this.targetSession.send<NavigationResult>("Page.navigate", { url });
if (result.errorText) {
throw new Error(`Navigation failed: ${result.errorText}`);
}
await loadPromise;
await this.waitForReadyState(timeoutMs);
}
async waitForReadyState(timeoutMs = 30_000): Promise<void> {
const startedAt = Date.now();
while (Date.now() - startedAt < timeoutMs) {
const state = await this.evaluate<string>("document.readyState");
if (state === "interactive" || state === "complete") {
return;
}
await sleep(150);
}
throw new Error("Timed out waiting for document.readyState");
}
async evaluate<T>(expression: string): Promise<T> {
return evaluateRuntime<T>(this.targetSession, expression);
}
async getHTML(): Promise<string> {
return this.evaluate<string>("document.documentElement.outerHTML");
}
async getTitle(): Promise<string> {
return this.evaluate<string>("document.title");
}
async getURL(): Promise<string> {
return this.evaluate<string>("window.location.href");
}
async bringToFront(): Promise<void> {
await this.targetSession.send("Page.bringToFront").catch(async () => {
await this.cdp.sendBrowserCommand("Target.activateTarget", {
targetId: this.targetSession.targetId,
});
});
if (this.interactive) {
await activateBrowserApp().catch(() => {});
}
}
async click(selector: string): Promise<void> {
const result = await this.evaluate<{ ok: boolean; error?: string }>(`
(() => {
const element = document.querySelector(${JSON.stringify(selector)});
if (!element) {
return { ok: false, error: "Element not found" };
}
element.scrollIntoView({ block: "center", inline: "center" });
if (element instanceof HTMLElement) {
element.click();
return { ok: true };
}
return { ok: false, error: "Element is not clickable" };
})()
`);
if (!result.ok) {
throw new Error(result.error ?? `Failed to click ${selector}`);
}
}
async scrollToEnd(options: { stepPx?: number; delayMs?: number; maxSteps?: number } = {}): Promise<void> {
const stepPx = options.stepPx ?? 1_400;
const delayMs = options.delayMs ?? 250;
const maxSteps = options.maxSteps ?? 6;
for (let step = 0; step < maxSteps; step += 1) {
const done = await this.evaluate<boolean>(`
(() => {
const before = window.scrollY;
window.scrollBy(0, ${stepPx});
const atBottom = window.innerHeight + window.scrollY >= document.body.scrollHeight - 4;
return atBottom || window.scrollY === before;
})()
`);
if (done) {
break;
}
await sleep(delayMs);
}
}
async close(): Promise<void> {
await this.cdp.closeTarget(this.targetSession.targetId);
}
}
+227
View File
@@ -0,0 +1,227 @@
#!/usr/bin/env bun
import {
runConvertCommand,
type ConvertCommandOptions,
type OutputFormat,
type WaitMode,
} from "./commands/convert";
export const HELP_TEXT = `
baoyu-fetch - Read a URL into Markdown or JSON with Chrome CDP
Usage:
baoyu-fetch <url> [options]
Options:
--output <file> Save output to file
--format <type> Output format: markdown | json
--json Alias for --format json
--adapter <name> Force an adapter (e.g. x, generic)
--download-media Download adapter-reported media into ./imgs and ./videos, then rewrite markdown links
--media-dir <dir> Base directory for downloaded media. Defaults to the output directory
--debug-dir <dir> Write debug artifacts
--cdp-url <url> Reuse an existing Chrome DevTools endpoint
--browser-path <path> Explicit Chrome binary path
--chrome-profile-dir <path>
Chrome user data dir. Defaults to BAOYU_CHROME_PROFILE_DIR
or baoyu-skills/chrome-profile.
--headless Launch a temporary headless Chrome if needed
--wait-for <mode> Wait mode: interaction | force
interaction: start visible Chrome and auto-wait only when login or verification is required
force: start visible Chrome, then auto-continue after it detects login/challenge progress
or continue immediately when you press Enter
--wait-for-interaction
Alias for --wait-for interaction
--wait-for-login Alias for --wait-for interaction
--interaction-timeout <ms>
How long to wait for manual interaction before failing (default: 600000)
--interaction-poll-interval <ms>
How often to poll interaction state while waiting (default: 1500)
--login-timeout <ms> Alias for --interaction-timeout
--login-poll-interval <ms>
Alias for --interaction-poll-interval
--timeout <ms> Page timeout in milliseconds (default: 30000)
--help Show help
Examples:
baoyu-fetch https://example.com
baoyu-fetch https://example.com --format markdown --output article.md --download-media
baoyu-fetch https://example.com --format json --output article.json
baoyu-fetch https://x.com/lennysan/status/2036483059407810640 --wait-for interaction
baoyu-fetch https://x.com/lennysan/status/2036483059407810640 --wait-for force
`.trim();
interface CliOptions extends ConvertCommandOptions {
url?: string;
help: boolean;
}
function normalizeWaitMode(raw: string): WaitMode {
const value = raw.toLowerCase();
if (value === "interaction" || value === "auto") {
return "interaction";
}
if (value === "force" || value === "manual" || value === "always") {
return "force";
}
throw new Error(`Invalid wait mode: ${raw}. Expected interaction or force.`);
}
function normalizeOutputFormat(raw: string): OutputFormat {
const value = raw.toLowerCase();
if (value === "markdown" || value === "json") {
return value;
}
throw new Error(`Invalid output format: ${raw}. Expected markdown or json.`);
}
export function parseArgs(argv: string[]): CliOptions {
const options: CliOptions = {
format: "markdown",
headless: false,
downloadMedia: false,
waitMode: "none",
interactionTimeoutMs: 600_000,
interactionPollIntervalMs: 1_500,
timeoutMs: 30_000,
help: false,
};
const args = argv.slice(2);
for (let index = 0; index < args.length; index += 1) {
const value = args[index];
if (value === "--help" || value === "-h") {
options.help = true;
continue;
}
if (value === "--format") {
const format = args[index + 1];
if (!format) {
throw new Error("--format requires a value");
}
options.format = normalizeOutputFormat(format);
index += 1;
continue;
}
if (value === "--json") {
options.format = "json";
continue;
}
if (value === "--download-media") {
options.downloadMedia = true;
continue;
}
if (value === "--headless") {
options.headless = true;
continue;
}
if (value === "--wait-for") {
const mode = args[index + 1];
if (!mode) {
throw new Error("--wait-for requires a mode");
}
options.waitMode = normalizeWaitMode(mode);
index += 1;
continue;
}
if (value === "--wait-for-interaction" || value === "--wait-for-login") {
options.waitMode = "interaction";
continue;
}
if (value === "--output") {
options.output = args[index + 1];
index += 1;
continue;
}
if (value === "--adapter") {
options.adapter = args[index + 1];
index += 1;
continue;
}
if (value === "--debug-dir") {
options.debugDir = args[index + 1];
index += 1;
continue;
}
if (value === "--media-dir") {
options.mediaDir = args[index + 1];
index += 1;
continue;
}
if (value === "--cdp-url") {
options.cdpUrl = args[index + 1];
index += 1;
continue;
}
if (value === "--browser-path") {
options.browserPath = args[index + 1];
index += 1;
continue;
}
if (value === "--chrome-profile-dir") {
options.chromeProfileDir = args[index + 1];
index += 1;
continue;
}
if (value === "--timeout") {
const parsed = Number(args[index + 1]);
if (!Number.isFinite(parsed) || parsed <= 0) {
throw new Error(`Invalid timeout: ${args[index + 1]}`);
}
options.timeoutMs = parsed;
index += 1;
continue;
}
if (value === "--interaction-timeout" || value === "--login-timeout") {
const parsed = Number(args[index + 1]);
if (!Number.isFinite(parsed) || parsed <= 0) {
throw new Error(`Invalid interaction timeout: ${args[index + 1]}`);
}
options.interactionTimeoutMs = parsed;
index += 1;
continue;
}
if (value === "--interaction-poll-interval" || value === "--login-poll-interval") {
const parsed = Number(args[index + 1]);
if (!Number.isFinite(parsed) || parsed <= 0) {
throw new Error(`Invalid interaction poll interval: ${args[index + 1]}`);
}
options.interactionPollIntervalMs = parsed;
index += 1;
continue;
}
if (value.startsWith("-")) {
throw new Error(`Unknown option: ${value}`);
}
if (!options.url) {
options.url = value;
continue;
}
throw new Error(`Unexpected argument: ${value}`);
}
return options;
}
async function main(): Promise<void> {
try {
const options = parseArgs(process.argv);
if (options.help || !options.url) {
console.log(HELP_TEXT);
return;
}
await runConvertCommand(options);
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
console.error(message);
process.exitCode = 1;
}
}
if (import.meta.main) {
void main();
}
@@ -0,0 +1,521 @@
import { mkdir, writeFile } from "node:fs/promises";
import { join } from "node:path";
import { createInterface } from "node:readline";
import { connectChrome, type ChromeConnection } from "../browser/chrome-launcher";
import { CdpClient } from "../browser/cdp-client";
import { detectInteractionGate } from "../browser/interaction-gates";
import { NetworkJournal } from "../browser/network-journal";
import { BrowserSession } from "../browser/session";
import { genericAdapter, resolveAdapter } from "../adapters";
import type { ExtractedDocument } from "../extract/document";
import { renderMarkdown } from "../extract/markdown-renderer";
import { downloadMediaAssets } from "../media/default-downloader";
import { rewriteMarkdownMediaLinks } from "../media/markdown-media";
import { createLogger } from "../utils/logger";
import { normalizeUrl } from "../utils/url";
import type {
Adapter,
AdapterContext,
AdapterLoginInfo,
LoginState,
MediaAsset,
WaitForInteractionRequest,
} from "../adapters/types";
export type WaitMode = "none" | "interaction" | "force";
export type OutputFormat = "markdown" | "json";
export interface ConvertCommandOptions {
url?: string;
output?: string;
format: OutputFormat;
adapter?: string;
debugDir?: string;
cdpUrl?: string;
browserPath?: string;
chromeProfileDir?: string;
headless: boolean;
downloadMedia: boolean;
mediaDir?: string;
waitMode: WaitMode;
interactionTimeoutMs: number;
interactionPollIntervalMs: number;
timeoutMs: number;
}
interface RuntimeResources {
chrome: ChromeConnection;
cdp: CdpClient;
browser: BrowserSession;
network: NetworkJournal;
interactive: boolean;
}
interface ForceWaitSnapshot {
url: string;
hasGate: boolean;
loginState: LoginState | "unavailable";
}
interface SuccessfulConvertOutput {
adapter: string;
status: "ok";
login?: AdapterLoginInfo;
media: MediaAsset[];
downloads: Awaited<ReturnType<typeof downloadMediaAssets>> | null;
document: ExtractedDocument;
markdown: string;
}
interface InteractionRequiredOutput {
adapter: string;
status: "needs_interaction";
login?: AdapterLoginInfo;
interaction: WaitForInteractionRequest;
}
function sleep(ms: number): Promise<void> {
return new Promise((resolve) => setTimeout(resolve, ms));
}
export function shouldAutoContinueForceWait(
initial: ForceWaitSnapshot,
current: ForceWaitSnapshot,
): boolean {
if (initial.hasGate && !current.hasGate) {
return true;
}
if (initial.loginState === "logged_out" && current.loginState !== "logged_out") {
return true;
}
if (initial.loginState !== "logged_in" && current.loginState === "logged_in") {
return true;
}
if (current.url !== initial.url && !current.hasGate && current.loginState !== "logged_out") {
return true;
}
return false;
}
async function writeOutput(path: string, content: string): Promise<void> {
const directory = path.includes("/") ? path.slice(0, path.lastIndexOf("/")) : "";
if (directory) {
await mkdir(directory, { recursive: true });
}
await writeFile(path, content, "utf8");
}
async function writeDebugArtifacts(
debugDir: string,
document: ExtractedDocument,
markdown: string,
browser: BrowserSession,
network: NetworkJournal,
): Promise<void> {
await mkdir(debugDir, { recursive: true });
const html = await browser.getHTML().catch(() => "");
const networkDump = await network.toJSON({ includeBodies: true });
await Promise.all([
writeFile(join(debugDir, "document.json"), JSON.stringify(document, null, 2), "utf8"),
writeFile(join(debugDir, "markdown.md"), markdown, "utf8"),
writeFile(join(debugDir, "page.html"), html, "utf8"),
writeFile(join(debugDir, "network.json"), JSON.stringify(networkDump, null, 2), "utf8"),
]);
}
async function openRuntime(
options: ConvertCommandOptions,
interactive: boolean,
debugEnabled: boolean,
): Promise<RuntimeResources> {
const logger = createLogger(debugEnabled);
if (interactive) {
logger.info("Opening Chrome in interactive mode.");
}
const chrome = await connectChrome({
cdpUrl: options.cdpUrl,
browserPath: options.browserPath,
profileDir: options.chromeProfileDir,
headless: interactive ? false : options.headless,
logger,
});
const cdp = await CdpClient.connect(chrome.browserWsUrl);
const browser = await BrowserSession.open(cdp, { interactive });
if (interactive) {
await browser.bringToFront().catch(() => {});
}
const network = new NetworkJournal(browser.targetSession, logger);
await network.start();
return {
chrome,
cdp,
browser,
network,
interactive,
};
}
async function closeRuntime(runtime: RuntimeResources | null | undefined): Promise<void> {
if (!runtime) {
return;
}
runtime.network.stop();
await runtime.browser.close().catch(() => {});
await runtime.cdp.close().catch(() => {});
await runtime.chrome.close().catch(() => {});
}
async function reopenInteractiveRuntime(
runtime: RuntimeResources,
options: ConvertCommandOptions,
debugEnabled: boolean,
): Promise<RuntimeResources> {
if (runtime.interactive) {
return runtime;
}
await closeRuntime(runtime);
return openRuntime(options, true, debugEnabled);
}
async function captureForceWaitSnapshot(
adapter: Adapter,
context: AdapterContext,
): Promise<ForceWaitSnapshot> {
const [gate, url, login] = await Promise.all([
detectInteractionGate(context.browser).catch(() => null),
context.browser.getURL().catch(() => context.input.url.toString()),
adapter.checkLogin?.(context).catch(() => ({
provider: adapter.name,
state: "unknown" as const,
})),
]);
return {
url,
hasGate: Boolean(gate),
loginState: login?.state ?? "unavailable",
};
}
async function waitForForceResume(
adapter: Adapter,
context: AdapterContext,
options: ConvertCommandOptions,
): Promise<void> {
if (context.interactive) {
await context.browser.bringToFront().catch(() => {});
}
const prompt =
"Chrome is ready. Complete any manual login or verification. Extraction will continue automatically after it detects progress, or press Enter to continue immediately.";
context.log.info(prompt);
const rl = createInterface({
input: process.stdin,
output: process.stderr,
});
let manualContinue = false;
let closed = false;
const closeReadline = (): void => {
if (!closed) {
closed = true;
rl.close();
}
};
rl.once("line", () => {
manualContinue = true;
closeReadline();
});
const initial = await captureForceWaitSnapshot(adapter, context);
const startedAt = Date.now();
try {
while (Date.now() - startedAt < options.interactionTimeoutMs) {
if (manualContinue) {
return;
}
const current = await captureForceWaitSnapshot(adapter, context);
if (shouldAutoContinueForceWait(initial, current)) {
return;
}
await sleep(options.interactionPollIntervalMs);
}
} finally {
closeReadline();
}
throw new Error("Timed out waiting for force-mode interaction to complete");
}
async function waitForInteraction(
adapter: Adapter,
context: AdapterContext,
interaction: WaitForInteractionRequest,
options: ConvertCommandOptions,
): Promise<AdapterLoginInfo> {
const timeoutMs = interaction.timeoutMs ?? options.interactionTimeoutMs;
const pollIntervalMs = interaction.pollIntervalMs ?? options.interactionPollIntervalMs;
if (context.interactive) {
await context.browser.bringToFront().catch(() => {});
}
context.log.info(interaction.prompt);
const startedAt = Date.now();
let lastLogin: AdapterLoginInfo | null = null;
while (Date.now() - startedAt < timeoutMs) {
if (interaction.kind === "login" && adapter.checkLogin) {
lastLogin = await adapter.checkLogin(context);
if (lastLogin.state === "logged_in") {
return lastLogin;
}
}
const gate = await detectInteractionGate(context.browser);
if (!gate) {
if (interaction.kind !== "login") {
return lastLogin ?? {
provider: interaction.provider,
state: "unknown",
reason: `${interaction.provider} challenge cleared`,
};
}
if (!adapter.checkLogin) {
return {
provider: interaction.provider,
state: "unknown",
};
}
lastLogin = await adapter.checkLogin(context);
if (lastLogin.state !== "logged_out") {
return lastLogin;
}
}
await sleep(pollIntervalMs);
}
const reason = lastLogin?.reason ? ` (${lastLogin.reason})` : "";
throw new Error(`Timed out waiting for ${interaction.provider} interaction${reason}`);
}
export function formatOutputContent(
format: OutputFormat,
payload: SuccessfulConvertOutput | InteractionRequiredOutput,
): string {
if (format === "json") {
return JSON.stringify(payload, null, 2);
}
if (payload.status !== "ok") {
throw new Error("Markdown output is only available for successful extraction results");
}
return payload.markdown;
}
function printOutput(content: string): void {
process.stdout.write(content);
if (!content.endsWith("\n")) {
process.stdout.write("\n");
}
}
export async function runConvertCommand(options: ConvertCommandOptions): Promise<void> {
if (!options.url) {
throw new Error("URL is required");
}
if (options.downloadMedia && !options.output) {
throw new Error("--download-media requires --output so media paths can be rewritten relative to the saved output file");
}
const url = normalizeUrl(options.url);
let runtime = await openRuntime(options, options.waitMode !== "none", Boolean(options.debugDir));
const logger = createLogger(Boolean(options.debugDir));
try {
const adapter = resolveAdapter({ url }, options.adapter);
let context: AdapterContext = {
input: { url },
browser: runtime.browser,
network: runtime.network,
cdp: runtime.cdp,
log: logger,
outputFormat: options.format,
timeoutMs: options.timeoutMs,
interactive: runtime.interactive,
downloadMedia: options.downloadMedia,
};
if (options.waitMode === "force") {
await context.browser.goto(url.toString(), options.timeoutMs).catch(() => {});
await waitForForceResume(adapter, context, options);
}
let result = await adapter.process(context);
if (result.status === "no_document") {
const interaction = await detectInteractionGate(context.browser);
if (interaction) {
result = {
status: "needs_interaction",
interaction,
login: result.login,
};
}
}
while (result.status === "needs_interaction") {
if (options.waitMode === "none") {
if (options.format === "json") {
printOutput(
formatOutputContent(options.format, {
adapter: adapter.name,
status: result.status,
login: result.login,
interaction: result.interaction,
}),
);
return;
}
throw new Error(`${adapter.name} requires manual interaction. Re-run with --wait-for interaction to continue after completing it.`);
}
if (result.interaction.requiresVisibleBrowser !== false) {
runtime = await reopenInteractiveRuntime(runtime, options, Boolean(options.debugDir));
}
context = {
input: { url },
browser: runtime.browser,
network: runtime.network,
cdp: runtime.cdp,
log: logger,
outputFormat: options.format,
timeoutMs: options.timeoutMs,
interactive: runtime.interactive,
downloadMedia: options.downloadMedia,
};
await context.browser.goto(url.toString(), options.timeoutMs).catch(() => {});
await waitForInteraction(adapter, context, result.interaction, options);
result = await adapter.process(context);
if (result.status === "no_document") {
const interaction = await detectInteractionGate(context.browser);
if (interaction) {
result = {
status: "needs_interaction",
interaction,
login: result.login,
};
}
}
}
let document: ExtractedDocument | null = result.status === "ok" ? result.document : null;
let media: MediaAsset[] = result.status === "ok" ? (result.media ?? []) : [];
let login = result.login;
let mediaAdapter = adapter;
if (!document && adapter.name !== genericAdapter.name && result.status === "no_document") {
logger.info(`Adapter ${adapter.name} returned no structured document; falling back to generic extraction`);
const fallback = await genericAdapter.process(context);
if (fallback.status === "ok") {
document = fallback.document;
media = fallback.media ?? [];
mediaAdapter = genericAdapter;
}
}
if (!document) {
throw new Error("Failed to extract a document from the target URL");
}
document.requestedUrl ??= url.toString();
let markdown = renderMarkdown(document);
let downloadResult:
| Awaited<ReturnType<typeof downloadMediaAssets>>
| null = null;
if (options.downloadMedia && options.output) {
downloadResult = mediaAdapter.downloadMedia
? await mediaAdapter.downloadMedia({
media,
outputPath: options.output,
mediaDir: options.mediaDir,
log: logger,
})
: await downloadMediaAssets({
media,
outputPath: options.output,
mediaDir: options.mediaDir,
log: logger,
});
markdown = rewriteMarkdownMediaLinks(markdown, downloadResult.replacements);
if (downloadResult.downloadedImages > 0 || downloadResult.downloadedVideos > 0) {
logger.info(
`Downloaded ${downloadResult.downloadedImages} images and ${downloadResult.downloadedVideos} videos`,
);
}
}
if (options.output) {
await writeOutput(
options.output,
formatOutputContent(options.format, {
adapter: document.adapter ?? adapter.name,
status: "ok",
login,
media,
downloads: downloadResult,
document,
markdown,
}),
);
logger.info(`Saved ${options.format} to ${options.output}`);
}
if (options.debugDir) {
await writeDebugArtifacts(options.debugDir, document, markdown, runtime.browser, runtime.network);
logger.info(`Wrote debug artifacts to ${options.debugDir}`);
}
if (options.format === "json") {
printOutput(
formatOutputContent(options.format, {
adapter: document.adapter ?? adapter.name,
status: "ok",
login,
media,
downloads: downloadResult,
document,
markdown,
}),
);
return;
}
printOutput(markdown);
} finally {
await closeRuntime(runtime);
}
}
@@ -0,0 +1,51 @@
export type ContentBlock =
| {
type: "paragraph";
text: string;
}
| {
type: "heading";
depth: number;
text: string;
}
| {
type: "list";
ordered: boolean;
items: string[];
}
| {
type: "quote";
text: string;
}
| {
type: "code";
code: string;
language?: string;
}
| {
type: "image";
url: string;
alt?: string;
}
| {
type: "html";
html: string;
}
| {
type: "markdown";
markdown: string;
};
export interface ExtractedDocument {
url: string;
requestedUrl?: string;
canonicalUrl?: string;
title?: string;
author?: string;
siteName?: string;
publishedAt?: string;
summary?: string;
content: ContentBlock[];
metadata?: Record<string, unknown>;
adapter?: string;
}
@@ -1,11 +1,11 @@
import { parseHTML } from "linkedom"; import { JSDOM } from "jsdom";
export interface CleaningOptions { export interface CleanHtmlOptions {
removeAds?: boolean; removeAds?: boolean;
removeBase64Images?: boolean; removeBase64Images?: boolean;
onlyMainContent?: boolean; onlyMainContent?: boolean;
includeTags?: string[]; includeSelectors?: string[];
excludeTags?: string[]; excludeSelectors?: string[];
} }
const ALWAYS_REMOVE_SELECTORS = [ const ALWAYS_REMOVE_SELECTORS = [
@@ -143,10 +143,12 @@ const AD_SELECTORS = [
function getLinkDensity(element: Element): number { function getLinkDensity(element: Element): number {
const text = element.textContent || ""; const text = element.textContent || "";
const textLength = text.trim().length; const textLength = text.trim().length;
if (textLength === 0) return 1; if (textLength === 0) {
return 1;
}
let linkLength = 0; let linkLength = 0;
element.querySelectorAll("a").forEach((link: Element) => { element.querySelectorAll("a").forEach((link) => {
linkLength += (link.textContent || "").trim().length; linkLength += (link.textContent || "").trim().length;
}); });
@@ -167,20 +169,29 @@ function getContentScore(element: Element): number {
score -= element.querySelectorAll("li").length * 0.2; score -= element.querySelectorAll("li").length * 0.2;
const linkDensity = getLinkDensity(element); const linkDensity = getLinkDensity(element);
if (linkDensity > 0.5) score -= 30; if (linkDensity > 0.5) {
else if (linkDensity > 0.3) score -= 15; score -= 30;
} else if (linkDensity > 0.3) {
score -= 15;
}
const className = typeof element.className === "string" ? element.className : ""; const className = typeof element.className === "string" ? element.className : "";
const classAndId = `${className} ${element.id || ""}`; const classAndId = `${className} ${element.id || ""}`;
if (/article|content|post|body|main|entry/i.test(classAndId)) score += 25; if (/article|content|post|body|main|entry/i.test(classAndId)) {
if (/comment|sidebar|footer|nav|menu|header|widget|ad/i.test(classAndId)) score -= 25; score += 25;
}
if (/comment|sidebar|footer|nav|menu|header|widget|ad/i.test(classAndId)) {
score -= 25;
}
return score; return score;
} }
function looksLikeNavigation(element: Element): boolean { function looksLikeNavigation(element: Element): boolean {
const linkDensity = getLinkDensity(element); const linkDensity = getLinkDensity(element);
if (linkDensity > 0.5) return true; if (linkDensity > 0.5) {
return true;
}
const listItems = element.querySelectorAll("li"); const listItems = element.querySelectorAll("li");
const links = element.querySelectorAll("a"); const links = element.querySelectorAll("a");
@@ -190,9 +201,9 @@ function looksLikeNavigation(element: Element): boolean {
function removeElements(document: Document, selectors: string[]): void { function removeElements(document: Document, selectors: string[]): void {
for (const selector of selectors) { for (const selector of selectors) {
try { try {
document.querySelectorAll(selector).forEach((element: Element) => element.remove()); document.querySelectorAll(selector).forEach((element) => element.remove());
} catch { } catch {
// Ignore unsupported selectors from linkedom/jsdom differences. // Ignore unsupported selectors.
} }
} }
} }
@@ -200,11 +211,11 @@ function removeElements(document: Document, selectors: string[]): void {
function removeWithProtection( function removeWithProtection(
document: Document, document: Document,
selectorsToRemove: string[], selectorsToRemove: string[],
protectedSelectors: string[] protectedSelectors: string[],
): void { ): void {
for (const selector of selectorsToRemove) { for (const selector of selectorsToRemove) {
try { try {
document.querySelectorAll(selector).forEach((element: Element) => { document.querySelectorAll(selector).forEach((element) => {
const isProtected = protectedSelectors.some((protectedSelector) => { const isProtected = protectedSelectors.some((protectedSelector) => {
try { try {
return element.matches(protectedSelector); return element.matches(protectedSelector);
@@ -212,7 +223,10 @@ function removeWithProtection(
return false; return false;
} }
}); });
if (isProtected) return;
if (isProtected) {
return;
}
const containsProtected = protectedSelectors.some((protectedSelector) => { const containsProtected = protectedSelectors.some((protectedSelector) => {
try { try {
@@ -221,29 +235,40 @@ function removeWithProtection(
return false; return false;
} }
}); });
if (containsProtected) return;
if (containsProtected) {
return;
}
element.remove(); element.remove();
}); });
} catch { } catch {
// Ignore unsupported selectors from linkedom/jsdom differences. // Ignore unsupported selectors.
} }
} }
} }
function findMainContent(document: Document): Element | null { function isValidContent(element: Element | null): element is Element {
const isValidContent = (element: Element | null): element is Element => { if (!element) {
if (!element) return false; return false;
const text = element.textContent || ""; }
if (text.trim().length < 100) return false; const text = element.textContent || "";
return !looksLikeNavigation(element); if (text.trim().length < 100) {
}; return false;
}
return !looksLikeNavigation(element);
}
function findMainContent(document: Document): Element | null {
const main = document.querySelector("main"); const main = document.querySelector("main");
if (isValidContent(main) && getLinkDensity(main) < 0.4) return main; if (isValidContent(main) && getLinkDensity(main) < 0.4) {
return main;
}
const roleMain = document.querySelector('[role="main"]'); const roleMain = document.querySelector('[role="main"]');
if (isValidContent(roleMain) && getLinkDensity(roleMain) < 0.4) return roleMain; if (isValidContent(roleMain) && getLinkDensity(roleMain) < 0.4) {
return roleMain;
}
const articles = document.querySelectorAll("article"); const articles = document.querySelectorAll("article");
if (articles.length === 1 && isValidContent(articles[0] ?? null)) { if (articles.length === 1 && isValidContent(articles[0] ?? null)) {
@@ -278,10 +303,11 @@ function findMainContent(document: Document): Element | null {
} }
const candidates: Array<{ element: Element; score: number }> = []; const candidates: Array<{ element: Element; score: number }> = [];
const containers = document.querySelectorAll("div, section, article"); document.querySelectorAll("div, section, article").forEach((element) => {
containers.forEach((element: Element) => {
const text = element.textContent || ""; const text = element.textContent || "";
if (text.trim().length < 200) return; if (text.trim().length < 200) {
return;
}
const score = getContentScore(element); const score = getContentScore(element);
if (score > 0) { if (score > 0) {
@@ -298,17 +324,17 @@ function findMainContent(document: Document): Element | null {
} }
function removeBase64ImagesFromDocument(document: Document): void { function removeBase64ImagesFromDocument(document: Document): void {
document.querySelectorAll("img[src^='data:']").forEach((element: Element) => { document.querySelectorAll("img[src^='data:']").forEach((element) => element.remove());
element.remove();
});
document.querySelectorAll("[style*='data:image']").forEach((element: Element) => { document.querySelectorAll("[style*='data:image']").forEach((element) => {
const style = element.getAttribute("style"); const style = element.getAttribute("style");
if (!style) return; if (!style) {
return;
}
const cleanedStyle = style.replace( const cleanedStyle = style.replace(
/background(-image)?:\s*url\([^)]*data:image[^)]*\)[^;]*;?/gi, /background(-image)?:\s*url\([^)]*data:image[^)]*\)[^;]*;?/gi,
"" "",
); );
if (cleanedStyle.trim()) { if (cleanedStyle.trim()) {
@@ -318,9 +344,9 @@ function removeBase64ImagesFromDocument(document: Document): void {
} }
}); });
document.querySelectorAll("source[src^='data:'], source[srcset*='data:']").forEach((element: Element) => { document
element.remove(); .querySelectorAll("source[src^='data:'], source[srcset*='data:']")
}); .forEach((element) => element.remove());
} }
function makeAbsoluteUrl(value: string, baseUrl: string): string | null { function makeAbsoluteUrl(value: string, baseUrl: string): string | null {
@@ -332,15 +358,19 @@ function makeAbsoluteUrl(value: string, baseUrl: string): string | null {
} }
function convertRelativeUrls(document: Document, baseUrl: string): void { function convertRelativeUrls(document: Document, baseUrl: string): void {
document.querySelectorAll("[src]").forEach((element: Element) => { document.querySelectorAll("[src]").forEach((element) => {
const src = element.getAttribute("src"); const src = element.getAttribute("src");
if (!src || src.startsWith("http") || src.startsWith("//") || src.startsWith("data:")) return; if (!src || src.startsWith("http") || src.startsWith("//") || src.startsWith("data:")) {
return;
}
const absolute = makeAbsoluteUrl(src, baseUrl); const absolute = makeAbsoluteUrl(src, baseUrl);
if (absolute) element.setAttribute("src", absolute); if (absolute) {
element.setAttribute("src", absolute);
}
}); });
document.querySelectorAll("[href]").forEach((element: Element) => { document.querySelectorAll("[href]").forEach((element) => {
const href = element.getAttribute("href"); const href = element.getAttribute("href");
if ( if (
!href || !href ||
@@ -355,20 +385,36 @@ function convertRelativeUrls(document: Document, baseUrl: string): void {
} }
const absolute = makeAbsoluteUrl(href, baseUrl); const absolute = makeAbsoluteUrl(href, baseUrl);
if (absolute) element.setAttribute("href", absolute); if (absolute) {
element.setAttribute("href", absolute);
}
}); });
} }
export function cleanHtml(html: string, baseUrl: string, options: CleaningOptions = {}): string { function removeComments(document: Document): void {
const walker = document.createTreeWalker(document, document.defaultView?.NodeFilter.SHOW_COMMENT ?? 128);
const comments: Comment[] = [];
while (walker.nextNode()) {
comments.push(walker.currentNode as Comment);
}
comments.forEach((comment) => comment.parentNode?.removeChild(comment));
}
export function cleanHtml(
html: string,
baseUrl: string,
options: CleanHtmlOptions = {},
): string {
const { const {
removeAds = true, removeAds = true,
removeBase64Images = true, removeBase64Images = true,
onlyMainContent = true, onlyMainContent = true,
includeTags, includeSelectors,
excludeTags, excludeSelectors,
} = options; } = options;
const { document } = parseHTML(html); const dom = new JSDOM(html, { url: baseUrl });
const { document } = dom.window;
removeElements(document, ALWAYS_REMOVE_SELECTORS); removeElements(document, ALWAYS_REMOVE_SELECTORS);
removeElements(document, OVERLAY_SELECTORS); removeElements(document, OVERLAY_SELECTORS);
@@ -377,8 +423,8 @@ export function cleanHtml(html: string, baseUrl: string, options: CleaningOption
removeElements(document, AD_SELECTORS); removeElements(document, AD_SELECTORS);
} }
if (excludeTags?.length) { if (excludeSelectors?.length) {
removeElements(document, excludeTags); removeElements(document, excludeSelectors);
} }
if (onlyMainContent) { if (onlyMainContent) {
@@ -386,18 +432,17 @@ export function cleanHtml(html: string, baseUrl: string, options: CleaningOption
const mainContent = findMainContent(document); const mainContent = findMainContent(document);
if (mainContent && document.body) { if (mainContent && document.body) {
const clone = mainContent.cloneNode(true) as Element; const clone = mainContent.cloneNode(true);
document.body.innerHTML = ""; document.body.innerHTML = "";
document.body.appendChild(clone); document.body.appendChild(clone);
} }
} }
if (includeTags?.length && document.body) { if (includeSelectors?.length && document.body) {
const matchedElements: Element[] = []; const matchedElements: Element[] = [];
for (const selector of includeSelectors) {
for (const selector of includeTags) {
try { try {
document.querySelectorAll(selector).forEach((element: Element) => { document.querySelectorAll(selector).forEach((element) => {
matchedElements.push(element.cloneNode(true) as Element); matchedElements.push(element.cloneNode(true) as Element);
}); });
} catch { } catch {
@@ -415,18 +460,8 @@ export function cleanHtml(html: string, baseUrl: string, options: CleaningOption
removeBase64ImagesFromDocument(document); removeBase64ImagesFromDocument(document);
} }
const walker = document.createTreeWalker(document, 128); removeComments(document);
const comments: Node[] = [];
while (walker.nextNode()) {
comments.push(walker.currentNode);
}
comments.forEach((comment) => comment.parentNode?.removeChild(comment));
convertRelativeUrls(document, baseUrl); convertRelativeUrls(document, baseUrl);
return document.documentElement?.outerHTML || html; return document.documentElement.outerHTML || html;
}
export function cleanContent(html: string, baseUrl: string, options: CleaningOptions = {}): string {
return cleanHtml(html, baseUrl, options);
} }
@@ -0,0 +1,83 @@
import { Readability } from "@mozilla/readability";
import { JSDOM } from "jsdom";
import type { ExtractedDocument } from "./document";
function getMetaContent(document: Document, selectors: string[]): string | undefined {
for (const selector of selectors) {
const value = document.querySelector(selector)?.getAttribute("content")?.trim();
if (value) {
return value;
}
}
return undefined;
}
export function extractDocumentFromHtml(input: {
url: string;
html: string;
adapter?: string;
}): ExtractedDocument {
const dom = new JSDOM(input.html, { url: input.url });
const document = dom.window.document;
const canonicalUrl =
document.querySelector('link[rel="canonical"]')?.getAttribute("href")?.trim() ??
getMetaContent(document, ['meta[property="og:url"]']);
const siteName = getMetaContent(document, [
'meta[property="og:site_name"]',
'meta[name="application-name"]',
]);
const metadataAuthor = getMetaContent(document, [
'meta[name="author"]',
'meta[property="article:author"]',
'meta[name="twitter:creator"]',
]);
const publishedAt = getMetaContent(document, [
'meta[property="article:published_time"]',
'meta[name="pubdate"]',
'meta[name="date"]',
'meta[itemprop="datePublished"]',
]);
const article = new Readability(document).parse();
const title =
article?.title?.trim() ||
getMetaContent(document, ['meta[property="og:title"]']) ||
document.title.trim() ||
undefined;
const summary =
article?.excerpt?.trim() ||
getMetaContent(document, [
'meta[name="description"]',
'meta[property="og:description"]',
'meta[name="twitter:description"]',
]);
const contentHtml =
article?.content?.trim() ||
document.querySelector("main")?.innerHTML?.trim() ||
document.body?.innerHTML?.trim() ||
"";
const author = article?.byline?.trim() || metadataAuthor;
return {
url: input.url,
canonicalUrl,
title,
author,
siteName,
publishedAt,
summary,
adapter: input.adapter ?? "generic",
metadata: {
language: document.documentElement.lang || undefined,
},
content: contentHtml ? [{ type: "html", html: contentHtml }] : [],
};
}
@@ -0,0 +1,758 @@
import { Readability } from "@mozilla/readability";
import { Defuddle } from "defuddle/node";
import { JSDOM, VirtualConsole } from "jsdom";
import TurndownService from "turndown";
import { gfm } from "turndown-plugin-gfm";
import { collectMediaFromMarkdown } from "../media/markdown-media";
import type { MediaAsset } from "../media/types";
import { cleanHtml } from "./html-cleaner";
export interface HtmlConversionMetadata {
url: string;
canonicalUrl?: string;
siteName?: string;
title?: string;
summary?: string;
author?: string;
publishedAt?: string;
coverImage?: string;
language?: string;
capturedAt: string;
}
export interface ConvertHtmlToMarkdownOptions {
enableRemoteMarkdownFallback?: boolean;
preserveBase64Images?: boolean;
}
export interface HtmlToMarkdownResult {
metadata: HtmlConversionMetadata;
markdown: string;
rawHtml: string;
cleanedHtml: string;
media: MediaAsset[];
conversionMethod: string;
fallbackReason?: string;
}
type JsonObject = Record<string, unknown>;
const MIN_CONTENT_LENGTH = 120;
const DEFUDDLE_API_ORIGIN = "https://defuddle.md";
const LOCAL_FALLBACK_SCORE_DELTA = 120;
const REMOTE_FALLBACK_SCORE_DELTA = 20;
const LOW_QUALITY_MARKERS = [
/Join The Conversation/i,
/One Community\. Many Voices/i,
/Read our community guidelines/i,
/Create a free account to share your thoughts/i,
/Become a Forbes Member/i,
/Subscribe to trusted journalism/i,
/\bComments\b/i,
];
const ARTICLE_TYPES = new Set([
"Article",
"NewsArticle",
"BlogPosting",
"WebPage",
"ReportageNewsArticle",
]);
const turndown = new TurndownService({
headingStyle: "atx",
bulletListMarker: "-",
codeBlockStyle: "fenced",
}) as TurndownService & {
remove(selectors: string[]): void;
addRule(
key: string,
rule: {
filter: string | ((node: Node) => boolean);
replacement: (content: string) => string;
},
): void;
};
turndown.use(gfm);
turndown.remove(["script", "style", "iframe", "noscript", "template", "svg", "path"]);
turndown.addRule("collapseFigure", {
filter: "figure",
replacement(content: string) {
return `\n\n${content.trim()}\n\n`;
},
});
turndown.addRule("dropInvisibleAnchors", {
filter(node: Node) {
return (
node.nodeName === "A" &&
!(node as Element).textContent?.trim() &&
!(node as Element).querySelector("img, video, picture, source")
);
},
replacement() {
return "";
},
});
function pickString(...values: unknown[]): string | undefined {
for (const value of values) {
if (typeof value !== "string") {
continue;
}
const trimmed = value.trim();
if (trimmed) {
return trimmed;
}
}
return undefined;
}
function normalizeMarkdown(markdown: string): string {
return markdown
.replace(/\r\n/g, "\n")
.replace(/[ \t]+\n/g, "\n")
.replace(/\n{3,}/g, "\n\n")
.trim();
}
function stripWrappingQuotes(value: string): string {
const trimmed = value.trim();
if (
(trimmed.startsWith('"') && trimmed.endsWith('"')) ||
(trimmed.startsWith("'") && trimmed.endsWith("'"))
) {
return trimmed.slice(1, -1).trim();
}
return trimmed;
}
function stripMarkdownFrontmatter(markdown: string): string {
return markdown.replace(/^\uFEFF?---\n[\s\S]*?\n---(?:\n|$)/, "").trim();
}
function cleanMarkdownTitle(value: string): string | undefined {
const cleaned = stripWrappingQuotes(
value
.replace(/\s+#+\s*$/, "")
.replace(/!\[[^\]]*\]\([^)]+\)/g, "")
.replace(/\[([^\]]+)\]\([^)]+\)/g, "$1")
.replace(/[*_`~]/g, "")
.trim(),
);
return cleaned || undefined;
}
export function extractTitleFromMarkdownDocument(markdown: string): string | undefined {
const normalized = markdown.replace(/\r\n/g, "\n").trim();
if (!normalized) {
return undefined;
}
const frontmatterMatch = normalized.match(/^\uFEFF?---\n([\s\S]*?)\n---(?:\n|$)/);
if (frontmatterMatch) {
for (const line of frontmatterMatch[1].split("\n")) {
const match = line.match(/^title:\s*(.+?)\s*$/i);
if (!match) {
continue;
}
const title = cleanMarkdownTitle(match[1]);
if (title) {
return title;
}
}
}
const body = stripMarkdownFrontmatter(normalized);
const headingMatch = body.match(/^#{1,6}\s+(.+)$/m);
if (!headingMatch) {
return undefined;
}
return cleanMarkdownTitle(headingMatch[1]);
}
function trimKnownBoilerplate(markdown: string): string {
const normalized = normalizeMarkdown(markdown);
const lines = normalized.split("\n");
while (lines.length > 0) {
const lastLine = lines[lines.length - 1]?.trim();
if (!lastLine) {
lines.pop();
continue;
}
if (/^继续滑动看下一个$/.test(lastLine) || /^轻触阅读原文$/.test(lastLine)) {
lines.pop();
continue;
}
break;
}
return normalizeMarkdown(lines.join("\n"));
}
function buildDefuddleApiUrl(targetUrl: string): string {
return `${DEFUDDLE_API_ORIGIN}/${encodeURIComponent(targetUrl)}`;
}
async function fetchDefuddleApiMarkdown(
targetUrl: string,
): Promise<{ markdown: string; title?: string }> {
const response = await fetch(buildDefuddleApiUrl(targetUrl), {
headers: {
accept: "text/markdown,text/plain;q=0.9,*/*;q=0.1",
},
redirect: "follow",
});
if (!response.ok) {
throw new Error(`defuddle.md returned ${response.status} ${response.statusText}`);
}
const rawMarkdown = (await response.text()).replace(/\r\n/g, "\n").trim();
if (!rawMarkdown) {
throw new Error("defuddle.md returned empty markdown");
}
const title = extractTitleFromMarkdownDocument(rawMarkdown);
const markdown = trimKnownBoilerplate(stripMarkdownFrontmatter(rawMarkdown));
if (!markdown) {
throw new Error("defuddle.md returned empty markdown");
}
return {
markdown,
title,
};
}
function sanitizeHtmlFragment(html: string): string {
const dom = new JSDOM(`<div id="__root">${html}</div>`);
const root = dom.window.document.querySelector("#__root");
if (!root) {
return html;
}
for (const selector of ["script", "style", "iframe", "noscript", "template", "svg", "path"]) {
root.querySelectorAll(selector).forEach((element) => element.remove());
}
return root.innerHTML;
}
function extractTextFromHtml(html: string): string {
const dom = new JSDOM(`<!doctype html><html><body>${html}</body></html>`);
const { document } = dom.window;
for (const selector of ["script", "style", "noscript", "template", "iframe", "svg", "path"]) {
document.querySelectorAll(selector).forEach((element) => element.remove());
}
return document.body?.textContent?.replace(/\s+/g, " ").trim() ?? "";
}
function getMetaContent(document: Document, names: string[]): string | undefined {
for (const name of names) {
const element =
document.querySelector(`meta[name="${name}"]`) ??
document.querySelector(`meta[property="${name}"]`);
const content = element?.getAttribute("content")?.trim();
if (content) {
return content;
}
}
return undefined;
}
function normalizeLanguageTag(value: string | null | undefined): string | undefined {
if (!value) {
return undefined;
}
const trimmed = value.trim();
if (!trimmed) {
return undefined;
}
const primary = trimmed.split(/[,\s;]/, 1)[0]?.trim();
if (!primary) {
return undefined;
}
return primary.replace(/_/g, "-");
}
function flattenJsonLdItems(data: unknown): JsonObject[] {
if (!data || typeof data !== "object") {
return [];
}
if (Array.isArray(data)) {
return data.flatMap(flattenJsonLdItems);
}
const item = data as JsonObject;
if (Array.isArray(item["@graph"])) {
return (item["@graph"] as unknown[]).flatMap(flattenJsonLdItems);
}
return [item];
}
function parseJsonLdScripts(document: Document): JsonObject[] {
const results: JsonObject[] = [];
document.querySelectorAll("script[type='application/ld+json']").forEach((script) => {
try {
const data = JSON.parse(script.textContent ?? "");
results.push(...flattenJsonLdItems(data));
} catch {
// Ignore malformed json-ld blocks.
}
});
return results;
}
function extractAuthorFromJsonLd(authorData: unknown): string | undefined {
if (typeof authorData === "string") {
return authorData.trim() || undefined;
}
if (!authorData || typeof authorData !== "object") {
return undefined;
}
if (Array.isArray(authorData)) {
return authorData
.map((author) => extractAuthorFromJsonLd(author))
.filter((value): value is string => Boolean(value))
.join(", ") || undefined;
}
const author = authorData as JsonObject;
return pickString(author.name);
}
function extractPrimaryJsonLdMeta(document: Document): Partial<HtmlConversionMetadata> {
for (const item of parseJsonLdScripts(document)) {
const type = Array.isArray(item["@type"]) ? item["@type"][0] : item["@type"];
if (typeof type !== "string" || !ARTICLE_TYPES.has(type)) {
continue;
}
return {
title: pickString(item.headline, item.name),
summary: pickString(item.description),
author: extractAuthorFromJsonLd(item.author),
publishedAt: pickString(item.datePublished, item.dateCreated),
coverImage: pickString(
item.image,
(item.image as JsonObject | undefined)?.url,
Array.isArray(item.image) ? item.image[0] : undefined,
),
};
}
return {};
}
function extractPageMetadata(
html: string,
url: string,
capturedAt: string,
): HtmlConversionMetadata {
const dom = new JSDOM(html, { url });
const { document } = dom.window;
const jsonLd = extractPrimaryJsonLdMeta(document);
return {
url,
canonicalUrl:
document.querySelector('link[rel="canonical"]')?.getAttribute("href")?.trim() ??
getMetaContent(document, ["og:url"]),
siteName: pickString(
getMetaContent(document, ["og:site_name"]),
document.querySelector('meta[name="application-name"]')?.getAttribute("content"),
),
title: pickString(
getMetaContent(document, ["og:title", "twitter:title"]),
jsonLd.title,
document.querySelector("h1")?.textContent,
document.title,
),
summary: pickString(
getMetaContent(document, ["description", "og:description", "twitter:description"]),
jsonLd.summary,
),
author: pickString(
getMetaContent(document, ["author", "article:author", "twitter:creator"]),
jsonLd.author,
),
publishedAt: pickString(
document.querySelector("time[datetime]")?.getAttribute("datetime"),
getMetaContent(document, ["article:published_time", "datePublished", "publishdate", "date"]),
jsonLd.publishedAt,
),
coverImage: pickString(
getMetaContent(document, ["og:image", "twitter:image", "twitter:image:src"]),
jsonLd.coverImage,
),
language: pickString(
normalizeLanguageTag(document.documentElement.getAttribute("lang")),
normalizeLanguageTag(
pickString(
getMetaContent(document, ["language", "content-language", "og:locale"]),
document.querySelector("meta[http-equiv='content-language']")?.getAttribute("content"),
),
),
),
capturedAt,
};
}
function isMarkdownUsable(markdown: string, html: string): boolean {
const normalized = normalizeMarkdown(markdown);
if (!normalized) {
return false;
}
const htmlTextLength = extractTextFromHtml(html).length;
if (htmlTextLength < MIN_CONTENT_LENGTH) {
return true;
}
if (normalized.length >= 80) {
return true;
}
return normalized.length >= Math.min(200, Math.floor(htmlTextLength * 0.2));
}
function countMarkerHits(markdown: string, markers: RegExp[]): number {
let hits = 0;
for (const marker of markers) {
if (marker.test(markdown)) {
hits += 1;
}
}
return hits;
}
function countUsefulParagraphs(markdown: string): number {
const paragraphs = normalizeMarkdown(markdown).split(/\n{2,}/);
let count = 0;
for (const paragraph of paragraphs) {
const trimmed = paragraph.trim();
if (!trimmed) {
continue;
}
if (/^!?\[[^\]]*\]\([^)]+\)$/.test(trimmed)) {
continue;
}
if (/^#{1,6}\s+/.test(trimmed)) {
continue;
}
if ((trimmed.match(/\b[\p{L}\p{N}']+\b/gu) || []).length < 8) {
continue;
}
count += 1;
}
return count;
}
function scoreMarkdownQuality(markdown: string): number {
const normalized = normalizeMarkdown(markdown);
const wordCount = (normalized.match(/\b[\p{L}\p{N}']+\b/gu) || []).length;
const usefulParagraphs = countUsefulParagraphs(normalized);
const headingCount = (normalized.match(/^#{1,6}\s+/gm) || []).length;
const markerHits = countMarkerHits(normalized, LOW_QUALITY_MARKERS);
return Math.min(wordCount, 4000) + usefulParagraphs * 40 + headingCount * 10 - markerHits * 180;
}
function shouldCompareWithFallback(markdown: string): boolean {
const normalized = normalizeMarkdown(markdown);
return countMarkerHits(normalized, LOW_QUALITY_MARKERS) > 0 || countUsefulParagraphs(normalized) < 6;
}
function hasMeaningfulMarkdownStructure(markdown: string): boolean {
const normalized = normalizeMarkdown(markdown);
if (!normalized) {
return false;
}
return (
countUsefulParagraphs(normalized) > 0 ||
/^#{1,6}\s+/m.test(normalized) ||
/^[-*]\s+/m.test(normalized) ||
/^\d+\.\s+/m.test(normalized) ||
/!\[[^\]]*\]\([^)]+\)/.test(normalized)
);
}
function shouldTryRemoteMarkdownFallback(
markdown: string,
html: string,
options: ConvertHtmlToMarkdownOptions,
): boolean {
if (!options.enableRemoteMarkdownFallback) {
return false;
}
return !isMarkdownUsable(markdown, html) || shouldCompareWithFallback(markdown);
}
function shouldPreferRemoteMarkdown(
current: HtmlToMarkdownResult,
remote: HtmlToMarkdownResult,
html: string,
): boolean {
if (!isMarkdownUsable(current.markdown, html)) {
return true;
}
if (!hasMeaningfulMarkdownStructure(current.markdown) && hasMeaningfulMarkdownStructure(remote.markdown)) {
return true;
}
return scoreMarkdownQuality(remote.markdown) > scoreMarkdownQuality(current.markdown) + REMOTE_FALLBACK_SCORE_DELTA;
}
function buildRemoteFallbackReason(current: HtmlToMarkdownResult, html: string): string {
if (!isMarkdownUsable(current.markdown, html)) {
return current.fallbackReason
? `Used defuddle.md markdown fallback after local extraction failed: ${current.fallbackReason}`
: "Used defuddle.md markdown fallback after local extraction returned empty or incomplete markdown";
}
return "defuddle.md produced higher-quality markdown than local extraction";
}
async function tryDefuddleConversion(
html: string,
url: string,
baseMetadata: HtmlConversionMetadata,
): Promise<{ ok: true; result: HtmlToMarkdownResult } | { ok: false; reason: string }> {
try {
const virtualConsole = new VirtualConsole();
virtualConsole.on("jsdomError", (error: Error & { type?: string }) => {
if (error.type === "css parsing" || /Could not parse CSS stylesheet/i.test(error.message)) {
return;
}
});
const dom = new JSDOM(html, { url, virtualConsole });
const result = await Defuddle(dom, url, { markdown: true });
const markdown = trimKnownBoilerplate(result.content || "");
if (!isMarkdownUsable(markdown, html)) {
return { ok: false, reason: "Defuddle returned empty or incomplete markdown" };
}
const metadata: HtmlConversionMetadata = {
...baseMetadata,
title: pickString(result.title, baseMetadata.title),
summary: pickString(result.description, baseMetadata.summary),
author: pickString(result.author, baseMetadata.author),
publishedAt: pickString(result.published, baseMetadata.publishedAt),
coverImage: pickString(result.image, baseMetadata.coverImage),
language: pickString(result.language, baseMetadata.language),
};
return {
ok: true,
result: {
metadata,
markdown,
rawHtml: html,
cleanedHtml: html,
media: collectMediaFromMarkdown(markdown).concat(
metadata.coverImage
? [{ url: metadata.coverImage, kind: "image", role: "cover" as const }]
: [],
),
conversionMethod: "defuddle",
},
};
} catch (error) {
return {
ok: false,
reason: error instanceof Error ? error.message : String(error),
};
}
}
async function tryDefuddleApiConversion(
html: string,
url: string,
baseMetadata: HtmlConversionMetadata,
): Promise<{ ok: true; result: HtmlToMarkdownResult } | { ok: false; reason: string }> {
try {
const result = await fetchDefuddleApiMarkdown(url);
const markdown = result.markdown;
if (!isMarkdownUsable(markdown, html) && scoreMarkdownQuality(markdown) < 80) {
return { ok: false, reason: "defuddle.md returned empty or incomplete markdown" };
}
const metadata: HtmlConversionMetadata = {
...baseMetadata,
title: pickString(result.title, baseMetadata.title),
};
return {
ok: true,
result: {
metadata,
markdown,
rawHtml: html,
cleanedHtml: html,
media: collectMediaFromMarkdown(markdown).concat(
metadata.coverImage
? [{ url: metadata.coverImage, kind: "image", role: "cover" as const }]
: [],
),
conversionMethod: "defuddle-api",
},
};
} catch (error) {
return {
ok: false,
reason: error instanceof Error ? error.message : String(error),
};
}
}
function convertHtmlFragmentToMarkdown(html: string): string {
if (!html.trim()) {
return "";
}
try {
return turndown.turndown(sanitizeHtmlFragment(html));
} catch {
return "";
}
}
function fallbackPlainText(html: string): string {
return trimKnownBoilerplate(extractTextFromHtml(html));
}
function convertWithReadability(
rawHtml: string,
cleanedHtml: string,
url: string,
baseMetadata: HtmlConversionMetadata,
): HtmlToMarkdownResult {
const dom = new JSDOM(cleanedHtml, { url });
const document = dom.window.document;
const article = new Readability(document).parse();
const contentHtml =
article?.content?.trim() ??
document.querySelector("main")?.innerHTML?.trim() ??
document.body?.innerHTML?.trim() ??
"";
let markdown = contentHtml ? convertHtmlFragmentToMarkdown(contentHtml) : "";
if (!markdown) {
markdown = fallbackPlainText(cleanedHtml);
}
const metadata: HtmlConversionMetadata = {
...baseMetadata,
title: pickString(article?.title, baseMetadata.title),
summary: pickString(article?.excerpt, baseMetadata.summary),
author: pickString(article?.byline, baseMetadata.author),
};
const media = collectMediaFromMarkdown(markdown);
if (metadata.coverImage) {
media.unshift({
url: metadata.coverImage,
kind: "image",
role: "cover",
});
}
return {
metadata,
markdown: trimKnownBoilerplate(markdown),
rawHtml,
cleanedHtml,
media,
conversionMethod: article?.content ? "legacy:readability" : "legacy:body",
};
}
export async function convertHtmlToMarkdown(
html: string,
url: string,
options: ConvertHtmlToMarkdownOptions = {},
): Promise<HtmlToMarkdownResult> {
const capturedAt = new Date().toISOString();
const baseMetadata = extractPageMetadata(html, url, capturedAt);
let cleanedHtml = html;
try {
cleanedHtml = cleanHtml(html, url, {
removeBase64Images: !options.preserveBase64Images,
});
} catch {
cleanedHtml = html;
}
let selectedResult: HtmlToMarkdownResult;
const defuddleResult = await tryDefuddleConversion(cleanedHtml, url, baseMetadata);
if (defuddleResult.ok) {
if (shouldCompareWithFallback(defuddleResult.result.markdown)) {
const fallbackResult = convertWithReadability(html, cleanedHtml, url, baseMetadata);
if (
scoreMarkdownQuality(fallbackResult.markdown) >
scoreMarkdownQuality(defuddleResult.result.markdown) + LOCAL_FALLBACK_SCORE_DELTA
) {
selectedResult = {
...fallbackResult,
fallbackReason: "Readability/Turndown produced higher-quality markdown than Defuddle",
};
} else {
selectedResult = {
...defuddleResult.result,
rawHtml: html,
cleanedHtml,
};
}
} else {
selectedResult = {
...defuddleResult.result,
rawHtml: html,
cleanedHtml,
};
}
} else {
selectedResult = {
...convertWithReadability(html, cleanedHtml, url, baseMetadata),
fallbackReason: defuddleResult.reason,
};
}
if (!shouldTryRemoteMarkdownFallback(selectedResult.markdown, cleanedHtml, options)) {
return selectedResult;
}
const remoteDefuddleResult = await tryDefuddleApiConversion(cleanedHtml, url, baseMetadata);
if (!remoteDefuddleResult.ok || !shouldPreferRemoteMarkdown(selectedResult, remoteDefuddleResult.result, cleanedHtml)) {
return selectedResult;
}
return {
...remoteDefuddleResult.result,
rawHtml: html,
cleanedHtml,
fallbackReason: buildRemoteFallbackReason(selectedResult, cleanedHtml),
};
}
@@ -0,0 +1,169 @@
import TurndownService from "turndown";
import { gfm } from "turndown-plugin-gfm";
import { normalizeMarkdownMediaLinks } from "../media/markdown-media";
import type { ContentBlock, ExtractedDocument } from "./document";
const turndownService = new TurndownService({
codeBlockStyle: "fenced",
headingStyle: "atx",
bulletListMarker: "-",
});
turndownService.use(gfm);
function renderBlock(block: ContentBlock): string {
switch (block.type) {
case "paragraph":
return block.text.trim();
case "heading":
return `${"#".repeat(Math.min(Math.max(block.depth, 1), 6))} ${block.text.trim()}`;
case "list":
return block.items
.map((item, index) => (block.ordered ? `${index + 1}. ${item.trim()}` : `- ${item.trim()}`))
.join("\n");
case "quote":
return block.text
.split("\n")
.map((line) => `> ${line}`)
.join("\n");
case "code":
return `\`\`\`${block.language ?? ""}\n${block.code.trimEnd()}\n\`\`\``;
case "image":
return `![${block.alt ?? ""}](${block.url})`;
case "html":
return turndownService.turndown(block.html).trim();
case "markdown":
return block.markdown.trim();
}
}
function isDefinedValue(value: unknown): boolean {
return value !== undefined && value !== null && value !== "";
}
function renderFrontmatterValue(value: unknown): string {
if (typeof value === "string") {
if (value.includes("\n")) {
return `|-\n${value
.replace(/\r\n/g, "\n")
.split("\n")
.map((line) => ` ${line}`)
.join("\n")}`;
}
return JSON.stringify(value);
}
if (typeof value === "number" || typeof value === "boolean") {
return String(value);
}
return JSON.stringify(value);
}
function renderFrontmatter(document: ExtractedDocument): string {
const fields = new Map<string, unknown>();
const preferredOrder = [
"title",
"url",
"requestedUrl",
"author",
"authorName",
"authorUsername",
"authorUrl",
"coverImage",
"siteName",
"publishedAt",
"summary",
"adapter",
];
fields.set("title", document.title);
fields.set("url", document.canonicalUrl ?? document.url);
fields.set("requestedUrl", document.requestedUrl ?? document.url);
fields.set("author", document.author);
fields.set("siteName", document.siteName);
fields.set("publishedAt", document.publishedAt);
fields.set("summary", document.summary);
fields.set("adapter", document.adapter);
for (const [key, value] of Object.entries(document.metadata ?? {})) {
if (!fields.has(key)) {
fields.set(key, value);
}
}
const orderedKeys = [
...preferredOrder.filter((key) => fields.has(key)),
...Array.from(fields.keys()).filter((key) => !preferredOrder.includes(key)).sort(),
];
const lines = orderedKeys
.map((key) => [key, fields.get(key)] as const)
.filter(([, value]) => isDefinedValue(value))
.map(([key, value]) => `${key}: ${renderFrontmatterValue(value)}`);
if (lines.length === 0) {
return "";
}
return `---\n${lines.join("\n")}\n---`;
}
function cleanMarkdown(markdown: string): string {
return normalizeMarkdownMediaLinks(markdown.replace(/\n{3,}/g, "\n\n").trim());
}
function normalizeComparableTitle(value: string): string {
return value
.trim()
.toLowerCase()
.replace(/^>\s*/, "")
.replace(/^#+\s+/, "")
.replace(/(?:\.{3}|…)\s*$/, "");
}
function bodyStartsWithTitle(body: string, title: string): boolean {
const firstMeaningfulLine = body
.replace(/\r\n/g, "\n")
.split("\n")
.map((line) => line.trim())
.find((line) => line && !/^!?\[[^\]]*\]\([^)]+\)$/.test(line));
if (!firstMeaningfulLine) {
return false;
}
const comparableTitle = normalizeComparableTitle(title);
const comparableFirstLine = normalizeComparableTitle(firstMeaningfulLine);
if (!comparableTitle || !comparableFirstLine) {
return false;
}
return (
comparableFirstLine === comparableTitle ||
comparableFirstLine.startsWith(comparableTitle) ||
comparableTitle.startsWith(comparableFirstLine)
);
}
export function renderMarkdown(document: ExtractedDocument): string {
const sections: string[] = [];
const frontmatter = renderFrontmatter(document);
if (frontmatter) {
sections.push(frontmatter);
}
const body = document.content
.map((block) => renderBlock(block))
.filter(Boolean)
.join("\n\n");
if (document.title && !bodyStartsWithTitle(body, document.title)) {
sections.push(`# ${document.title}`);
}
if (body) {
sections.push(body);
}
return cleanMarkdown(sections.join("\n\n"));
}
@@ -0,0 +1,161 @@
import path from "node:path";
import { mkdir, writeFile } from "node:fs/promises";
import {
buildFileName,
isDataUri,
normalizeContentType,
normalizeMediaUrl,
resolveExtensionFromContentType,
resolveExtensionFromUrl,
resolveMediaKind,
resolveOutputExtension,
toPosixPath,
} from "./media-utils";
import type { MediaAsset, MediaDownloadRequest, MediaDownloadResult, MediaKind } from "./types";
const DOWNLOAD_USER_AGENT =
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36";
function parseBase64DataUri(rawUrl: string): { contentType: string; bytes: Buffer } | null {
const match = rawUrl.match(/^data:([^;,]+);base64,([A-Za-z0-9+/=\s]+)$/i);
if (!match?.[1] || !match[2]) {
return null;
}
const contentType = normalizeContentType(match[1]);
if (!contentType) {
return null;
}
try {
const bytes = Buffer.from(match[2].replace(/\s+/g, ""), "base64");
if (bytes.length === 0) {
return null;
}
return { contentType, bytes };
} catch {
return null;
}
}
function dedupeMedia(media: MediaAsset[]): MediaAsset[] {
const deduped: MediaAsset[] = [];
const seen = new Set<string>();
for (const item of media) {
const normalizedUrl = normalizeMediaUrl(item.url);
if (!normalizedUrl || seen.has(normalizedUrl)) {
continue;
}
seen.add(normalizedUrl);
deduped.push({
...item,
url: normalizedUrl,
});
}
return deduped;
}
function toRelativePath(fromDir: string, absoluteTarget: string): string {
const relative = path.relative(fromDir, absoluteTarget) || path.basename(absoluteTarget);
return toPosixPath(relative);
}
export async function downloadMediaAssets(
request: MediaDownloadRequest,
): Promise<MediaDownloadResult> {
const dedupedMedia = dedupeMedia(request.media);
const absoluteOutputPath = path.resolve(request.outputPath);
const markdownDir = path.dirname(absoluteOutputPath);
const baseDir = request.mediaDir ? path.resolve(request.mediaDir) : markdownDir;
const replacements: MediaDownloadResult["replacements"] = [];
let downloadedImages = 0;
let downloadedVideos = 0;
for (const asset of dedupedMedia) {
try {
let sourceUrl = normalizeMediaUrl(asset.url);
let contentType = "";
let extension: string | undefined;
let kind: MediaKind | undefined;
let bytes: Buffer | null = null;
if (isDataUri(asset.url)) {
const parsed = parseBase64DataUri(asset.url);
if (!parsed) {
request.log.warn(`Skipping unsupported embedded media: ${asset.url.slice(0, 32)}...`);
continue;
}
contentType = parsed.contentType;
extension =
resolveExtensionFromContentType(contentType) ??
resolveExtensionFromUrl(asset.fileNameHint ?? "");
kind = resolveMediaKind(sourceUrl, contentType, extension, asset.kind);
bytes = parsed.bytes;
} else {
const response = await fetch(sourceUrl, {
method: "GET",
redirect: "follow",
headers: {
"user-agent": DOWNLOAD_USER_AGENT,
...(asset.headers ?? {}),
},
});
if (!response.ok) {
request.log.warn(`Skipping media (${response.status}): ${asset.url}`);
continue;
}
sourceUrl = normalizeMediaUrl(response.url || sourceUrl);
contentType = normalizeContentType(response.headers.get("content-type"));
extension =
resolveExtensionFromUrl(sourceUrl) ??
resolveExtensionFromUrl(asset.url) ??
resolveExtensionFromUrl(asset.fileNameHint ?? "");
kind = resolveMediaKind(sourceUrl, contentType, extension, asset.kind);
bytes = Buffer.from(await response.arrayBuffer());
}
if (!kind || !bytes) {
request.log.debug(`Skipping media with unresolved kind: ${asset.url}`);
continue;
}
const outputExtension = resolveOutputExtension(contentType, extension, kind);
const nextIndex = kind === "image" ? downloadedImages + 1 : downloadedVideos + 1;
const dirName = kind === "image" ? "imgs" : "videos";
const targetDir = path.join(baseDir, dirName);
await mkdir(targetDir, { recursive: true });
const fileName = buildFileName(kind, nextIndex, sourceUrl, outputExtension, asset.fileNameHint);
const absolutePath = path.join(targetDir, fileName);
await writeFile(absolutePath, bytes);
replacements.push({
url: asset.url,
localPath: toRelativePath(markdownDir, absolutePath),
absolutePath,
kind,
});
if (kind === "image") {
downloadedImages = nextIndex;
} else {
downloadedVideos = nextIndex;
}
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
request.log.warn(`Failed to download media ${asset.url}: ${message}`);
}
}
return {
replacements,
downloadedImages,
downloadedVideos,
imageDir: downloadedImages > 0 ? path.join(baseDir, "imgs") : null,
videoDir: downloadedVideos > 0 ? path.join(baseDir, "videos") : null,
};
}
@@ -0,0 +1,458 @@
import remarkGfm from "remark-gfm";
import remarkParse from "remark-parse";
import { unified } from "unified";
import type { ContentBlock, ExtractedDocument } from "../extract/document";
import {
isDataUri,
normalizeContentType,
normalizeMediaUrl,
resolveExtensionFromContentType,
resolveExtensionFromUrl,
resolveKindFromExtension,
} from "./media-utils";
import type { MediaAsset, MediaReplacement } from "./types";
const MARKDOWN_LINK_RE =
/(!?\[[^\]\n]*\])\((<)?((?:https?:\/\/[^)\s>]+)|(?:data:[^)>\s]+))(>)?\)/g;
const FRONTMATTER_COVER_RE = /^(coverImage:\s*")((?:https?:\/\/[^"]+)|(?:data:[^"]+))(")/m;
const RAW_URL_RE = /(?:https?:\/\/[^\s<>"')\]]+|data:[^\s<>"')\]]+)/g;
interface MarkdownAstNode {
type: string;
url?: string | null;
alt?: string | null;
title?: string | null;
value?: string | null;
children?: MarkdownAstNode[];
position?: {
start?: { offset?: number | null };
end?: { offset?: number | null };
};
}
interface MarkdownReplacementRange {
start: number;
end: number;
value: string;
}
function inferMediaKindFromLabel(label: string, rawUrl: string): "image" | "video" | undefined {
if (label.startsWith("![")) {
return "image";
}
const normalizedLabel = label.replace(/[!\[\]]/g, "").trim().toLowerCase();
if (/\b(video|animated[_ -]?gif|gif)\b/.test(normalizedLabel)) {
return "video";
}
if (isDataUri(rawUrl)) {
const contentType = normalizeContentType(rawUrl.slice(5, rawUrl.indexOf(";")));
return contentType.startsWith("image/") ? "image" : contentType.startsWith("video/") ? "video" : undefined;
}
return resolveKindFromExtension(resolveExtensionFromUrl(rawUrl));
}
function inferMediaKindFromRawUrl(rawUrl: string): "image" | "video" | undefined {
if (isDataUri(rawUrl)) {
const contentType = normalizeContentType(rawUrl.slice(5, rawUrl.indexOf(";")));
return contentType.startsWith("image/") ? "image" : contentType.startsWith("video/") ? "video" : undefined;
}
return resolveKindFromExtension(resolveExtensionFromUrl(rawUrl));
}
function pushMedia(assets: MediaAsset[], seen: Set<string>, media: MediaAsset): void {
const normalizedUrl = normalizeMediaUrl(media.url);
if (!normalizedUrl || seen.has(normalizedUrl)) {
return;
}
seen.add(normalizedUrl);
assets.push({
...media,
url: normalizedUrl,
});
}
function getNodeOffsets(node: MarkdownAstNode): { start: number; end: number } | null {
const start = node.position?.start?.offset;
const end = node.position?.end?.offset;
if (typeof start !== "number" || typeof end !== "number" || start < 0 || end < start) {
return null;
}
return { start, end };
}
function escapeMarkdownLabel(value: string): string {
return value.replace(/\\/g, "\\\\").replace(/\[/g, "\\[").replace(/\]/g, "\\]");
}
function escapeMarkdownTitle(value: string): string {
return value.replace(/\\/g, "\\\\").replace(/"/g, '\\"');
}
function formatMarkdownDestination(url: string): string {
return /[\s()<>]/.test(url) ? `<${url}>` : url;
}
function serializeImageNode(node: MarkdownAstNode): string {
const rawUrl = node.url ?? "";
const normalizedUrl = normalizeMediaUrl(rawUrl);
const alt = escapeMarkdownLabel(node.alt ?? "");
const title = node.title ? ` "${escapeMarkdownTitle(node.title)}"` : "";
return `![${alt}](${formatMarkdownDestination(normalizedUrl)}${title})`;
}
function serializeLinkedImageNode(linkNode: MarkdownAstNode, imageNode: MarkdownAstNode): string {
const imageMarkdown = serializeImageNode(imageNode);
const imageUrl = normalizeMediaUrl(imageNode.url ?? "");
const linkUrl = normalizeMediaUrl(linkNode.url ?? "");
if (!linkUrl || linkUrl === imageUrl) {
return imageMarkdown;
}
const title = linkNode.title ? ` "${escapeMarkdownTitle(linkNode.title)}"` : "";
return `[${imageMarkdown}](${formatMarkdownDestination(linkUrl)}${title})`;
}
function isParagraphWithSingleText(node: MarkdownAstNode | undefined, expectedValue: string): boolean {
if (node?.type !== "paragraph" || node.children?.length !== 1) {
return false;
}
const child = node.children[0];
return child?.type === "text" && child.value?.trim() === expectedValue;
}
function getSingleImageFromParagraph(node: MarkdownAstNode | undefined): MarkdownAstNode | null {
if (node?.type !== "paragraph" || node.children?.length !== 1) {
return null;
}
return node.children[0]?.type === "image" ? node.children[0] : null;
}
function extractBrokenLinkedImageDestination(node: MarkdownAstNode | undefined): string | null {
if (node?.type !== "paragraph") {
return null;
}
const children = node.children ?? [];
if (children.length !== 3) {
return null;
}
const [prefix, linkNode, suffix] = children;
if (prefix?.type !== "text" || prefix.value?.trim() !== "](") {
return null;
}
if (linkNode?.type !== "link" || !linkNode.url) {
return null;
}
if (suffix?.type !== "text" || suffix.value?.trim() !== ")") {
return null;
}
return linkNode.url;
}
function collectLinkedImageReplacements(
node: MarkdownAstNode,
replacements: MarkdownReplacementRange[],
): void {
const children = node.children ?? [];
if (node.type === "link" && children.length === 1 && children[0]?.type === "image") {
const offsets = getNodeOffsets(node);
if (offsets) {
replacements.push({
start: offsets.start,
end: offsets.end,
value: serializeLinkedImageNode(node, children[0]),
});
}
return;
}
for (const child of children) {
collectLinkedImageReplacements(child, replacements);
}
}
function collectBrokenLinkedImageReplacements(
node: MarkdownAstNode,
replacements: MarkdownReplacementRange[],
): void {
const children = node.children ?? [];
for (let index = 0; index <= children.length - 3; index += 1) {
const openParagraph = children[index];
const imageParagraph = children[index + 1];
const closeParagraph = children[index + 2];
if (!isParagraphWithSingleText(openParagraph, "[")) {
continue;
}
const imageNode = getSingleImageFromParagraph(imageParagraph);
if (!imageNode) {
continue;
}
const linkUrl = extractBrokenLinkedImageDestination(closeParagraph);
if (!linkUrl) {
continue;
}
const start = openParagraph.position?.start?.offset;
const end = closeParagraph.position?.end?.offset;
if (typeof start !== "number" || typeof end !== "number" || end < start) {
continue;
}
replacements.push({
start,
end,
value: serializeLinkedImageNode({ type: "link", url: linkUrl }, imageNode),
});
index += 2;
}
for (const child of children) {
collectBrokenLinkedImageReplacements(child, replacements);
}
}
function applyReplacements(source: string, replacements: MarkdownReplacementRange[]): string {
if (replacements.length === 0) {
return source;
}
let result = source;
const sorted = [...replacements].sort((left, right) => right.start - left.start);
for (const replacement of sorted) {
result = `${result.slice(0, replacement.start)}${replacement.value}${result.slice(replacement.end)}`;
}
return result;
}
function normalizeLinkedImageMarkdown(markdown: string): string {
let tree: MarkdownAstNode;
try {
tree = unified().use(remarkParse).use(remarkGfm).parse(markdown) as MarkdownAstNode;
} catch {
return markdown;
}
const replacements: MarkdownReplacementRange[] = [];
collectLinkedImageReplacements(tree, replacements);
collectBrokenLinkedImageReplacements(tree, replacements);
return applyReplacements(markdown, replacements);
}
export function normalizeMarkdownMediaLinks(markdown: string): string {
MARKDOWN_LINK_RE.lastIndex = 0;
let result = markdown.replace(MARKDOWN_LINK_RE, (full, label, openAngle, rawUrl, closeAngle) => {
const normalizedUrl = normalizeMediaUrl(rawUrl);
if (normalizedUrl === rawUrl) {
return full;
}
return `${label}(${openAngle ?? ""}${normalizedUrl}${closeAngle ?? ""})`;
});
result = result.replace(FRONTMATTER_COVER_RE, (full, prefix, rawUrl, suffix) => {
const normalizedUrl = normalizeMediaUrl(rawUrl);
if (normalizedUrl === rawUrl) {
return full;
}
return `${prefix}${normalizedUrl}${suffix}`;
});
RAW_URL_RE.lastIndex = 0;
result = result.replace(RAW_URL_RE, (rawUrl) => normalizeMediaUrl(rawUrl));
return normalizeLinkedImageMarkdown(result);
}
export function collectMediaFromText(
text: string,
options: {
role?: MediaAsset["role"];
defaultKind?: MediaAsset["kind"];
seen?: Set<string>;
into?: MediaAsset[];
} = {},
): MediaAsset[] {
const assets = options.into ?? [];
const seen = options.seen ?? new Set<string>();
MARKDOWN_LINK_RE.lastIndex = 0;
let linkMatch: RegExpExecArray | null;
while ((linkMatch = MARKDOWN_LINK_RE.exec(text))) {
const label = linkMatch[1] ?? "";
const rawUrl = linkMatch[3] ?? "";
const kind = inferMediaKindFromLabel(label, rawUrl) ?? options.defaultKind;
if (!kind) {
continue;
}
pushMedia(assets, seen, {
url: rawUrl,
kind,
role: options.role ?? "inline",
});
}
RAW_URL_RE.lastIndex = 0;
let rawMatch: RegExpExecArray | null;
while ((rawMatch = RAW_URL_RE.exec(text))) {
const rawUrl = rawMatch[0] ?? "";
const kind = inferMediaKindFromRawUrl(rawUrl) ?? options.defaultKind;
if (!kind) {
continue;
}
pushMedia(assets, seen, {
url: rawUrl,
kind,
role: options.role ?? "inline",
});
}
return assets;
}
function collectMediaFromBlock(
block: ContentBlock,
assets: MediaAsset[],
seen: Set<string>,
): void {
switch (block.type) {
case "image":
pushMedia(assets, seen, {
url: block.url,
kind: "image",
role: "inline",
alt: block.alt,
});
return;
case "html":
case "markdown":
collectMediaFromText(block.type === "html" ? block.html : block.markdown, {
role: "inline",
seen,
into: assets,
});
return;
case "paragraph":
case "quote":
collectMediaFromText(block.text, {
role: "inline",
seen,
into: assets,
});
return;
case "list":
for (const item of block.items) {
collectMediaFromText(item, {
role: "attachment",
seen,
into: assets,
});
}
return;
case "heading":
case "code":
return;
}
}
export function collectMediaFromDocument(document: ExtractedDocument): MediaAsset[] {
const assets: MediaAsset[] = [];
const seen = new Set<string>();
const coverImage =
typeof document.metadata?.coverImage === "string" ? document.metadata.coverImage : undefined;
if (coverImage) {
pushMedia(assets, seen, {
url: coverImage,
kind: "image",
role: "cover",
});
}
for (const block of document.content) {
collectMediaFromBlock(block, assets, seen);
}
return assets;
}
export function collectMediaFromMarkdown(markdown: string): MediaAsset[] {
const assets: MediaAsset[] = [];
const seen = new Set<string>();
const fmMatch = markdown.match(/^---\n([\s\S]*?)\n---/);
if (fmMatch) {
const coverMatch = fmMatch[1]?.match(FRONTMATTER_COVER_RE);
if (coverMatch?.[2]) {
pushMedia(assets, seen, {
url: coverMatch[2],
kind: "image",
role: "cover",
});
}
}
collectMediaFromText(markdown, { seen, into: assets });
return assets;
}
export function rewriteMarkdownMediaLinks(
markdown: string,
replacements: MediaReplacement[],
): string {
if (replacements.length === 0) {
return markdown;
}
const replacementMap = new Map<string, string>();
for (const item of replacements) {
replacementMap.set(item.url, item.localPath);
replacementMap.set(normalizeMediaUrl(item.url), item.localPath);
}
MARKDOWN_LINK_RE.lastIndex = 0;
let result = markdown.replace(MARKDOWN_LINK_RE, (full, label, _openAngle, rawUrl) => {
const replacement = replacementMap.get(rawUrl) ?? replacementMap.get(normalizeMediaUrl(rawUrl));
if (!replacement) {
return full;
}
return `${label}(${replacement})`;
});
result = result.replace(FRONTMATTER_COVER_RE, (full, prefix, rawUrl, suffix) => {
const replacement = replacementMap.get(rawUrl) ?? replacementMap.get(normalizeMediaUrl(rawUrl));
if (!replacement) {
return full;
}
return `${prefix}${replacement}${suffix}`;
});
for (const { url, localPath } of replacements) {
result = result.split(url).join(localPath);
const normalizedUrl = normalizeMediaUrl(url);
if (normalizedUrl !== url) {
result = result.split(normalizedUrl).join(localPath);
}
}
return result;
}
export function resolveDataUriExtension(rawUrl: string): string | undefined {
if (!isDataUri(rawUrl)) {
return undefined;
}
const separatorIndex = rawUrl.indexOf(";");
const contentType = normalizeContentType(rawUrl.slice(5, separatorIndex === -1 ? undefined : separatorIndex));
return resolveExtensionFromContentType(contentType);
}
@@ -0,0 +1,261 @@
import path from "node:path";
import type { MediaKind } from "./types";
const IMAGE_EXTENSIONS = new Set([
"jpg",
"jpeg",
"png",
"webp",
"gif",
"bmp",
"avif",
"heic",
"heif",
"svg",
]);
const VIDEO_EXTENSIONS = new Set(["mp4", "m4v", "mov", "webm", "mkv"]);
const MIME_EXTENSION_MAP: Record<string, string> = {
"image/jpeg": "jpg",
"image/jpg": "jpg",
"image/png": "png",
"image/webp": "webp",
"image/gif": "gif",
"image/bmp": "bmp",
"image/avif": "avif",
"image/heic": "heic",
"image/heif": "heif",
"image/svg+xml": "svg",
"video/mp4": "mp4",
"video/webm": "webm",
"video/quicktime": "mov",
"video/x-m4v": "m4v",
};
export function normalizeContentType(raw: string | null): string {
return raw?.split(";")[0]?.trim().toLowerCase() ?? "";
}
export function normalizeExtension(raw: string | undefined | null): string | undefined {
if (!raw) {
return undefined;
}
const trimmed = raw.replace(/^\./, "").trim().toLowerCase();
if (!trimmed) {
return undefined;
}
if (trimmed === "jpeg" || trimmed === "jpg") {
return "jpg";
}
return trimmed;
}
export function resolveExtensionFromUrl(rawUrl: string): string | undefined {
try {
const parsed = new URL(rawUrl);
const extFromPath = normalizeExtension(path.posix.extname(parsed.pathname));
if (extFromPath) {
return extFromPath;
}
const extFromFormat = normalizeExtension(parsed.searchParams.get("format"));
if (extFromFormat) {
return extFromFormat;
}
} catch {
return undefined;
}
return undefined;
}
export function resolveExtensionFromContentType(contentType: string): string | undefined {
return normalizeExtension(MIME_EXTENSION_MAP[contentType]);
}
export function resolveKindFromContentType(contentType: string): MediaKind | undefined {
if (!contentType) {
return undefined;
}
if (contentType.startsWith("image/")) {
return "image";
}
if (contentType.startsWith("video/")) {
return "video";
}
return undefined;
}
export function resolveKindFromExtension(extension: string | undefined): MediaKind | undefined {
if (!extension) {
return undefined;
}
if (IMAGE_EXTENSIONS.has(extension)) {
return "image";
}
if (VIDEO_EXTENSIONS.has(extension)) {
return "video";
}
return undefined;
}
export function resolveMediaKind(
rawUrl: string,
contentType: string,
extension: string | undefined,
hint?: MediaKind,
): MediaKind | undefined {
const kindFromType = resolveKindFromContentType(contentType);
if (kindFromType) {
return kindFromType;
}
const kindFromExtension = resolveKindFromExtension(extension);
if (kindFromExtension) {
return kindFromExtension;
}
if (contentType && contentType !== "application/octet-stream") {
return undefined;
}
if (hint) {
return hint;
}
if (rawUrl.startsWith("data:image/")) {
return "image";
}
if (rawUrl.startsWith("data:video/")) {
return "video";
}
return undefined;
}
export function resolveOutputExtension(
contentType: string,
extension: string | undefined,
kind: MediaKind,
): string {
const fromMime = resolveExtensionFromContentType(contentType);
if (fromMime) {
return fromMime;
}
const normalized = normalizeExtension(extension);
if (normalized) {
return normalized;
}
return kind === "video" ? "mp4" : "jpg";
}
export function isDataUri(value: string): boolean {
return value.startsWith("data:");
}
export function safeDecodeURIComponent(value: string): string {
try {
return decodeURIComponent(value);
} catch {
return value;
}
}
function extractEmbeddedUrl(value: string): string | undefined {
const encodedMatch = value.match(/https?%3A%2F%2F.+$/i)?.[0];
if (encodedMatch) {
const decoded = safeDecodeURIComponent(encodedMatch);
try {
return new URL(decoded).href;
} catch {
return undefined;
}
}
const literalMatch = value.match(/https?:\/\/.+$/i)?.[0];
if (!literalMatch) {
return undefined;
}
try {
return new URL(literalMatch).href;
} catch {
return undefined;
}
}
export function normalizeMediaUrl(rawUrl: string): string {
if (isDataUri(rawUrl)) {
return rawUrl;
}
try {
const parsed = new URL(rawUrl);
const hostname = parsed.hostname.toLowerCase();
if (hostname === "substackcdn.com" || hostname.endsWith(".substackcdn.com")) {
const embeddedUrl = extractEmbeddedUrl(`${parsed.pathname}${parsed.search}`);
if (embeddedUrl) {
return embeddedUrl;
}
}
return parsed.href;
} catch {
return rawUrl;
}
}
export function sanitizeFileSegment(input: string): string {
return input
.replace(/[^a-zA-Z0-9_-]+/g, "-")
.replace(/-+/g, "-")
.replace(/^[-_]+|[-_]+$/g, "")
.slice(0, 48);
}
export function resolveFileStem(rawUrl: string, extension: string, fileNameHint?: string): string {
const hintBase = fileNameHint?.trim();
if (hintBase) {
const parsed = path.posix.parse(hintBase);
const stem = parsed.name || parsed.base;
return sanitizeFileSegment(stem);
}
if (isDataUri(rawUrl)) {
return "";
}
try {
const parsed = new URL(rawUrl);
const base = path.posix.basename(parsed.pathname);
if (!base) {
return "";
}
const decodedBase = safeDecodeURIComponent(base);
const normalizedExtension = normalizeExtension(extension);
const stripExtension = normalizedExtension ? new RegExp(`\\.${normalizedExtension}$`, "i") : null;
const rawStem = stripExtension ? decodedBase.replace(stripExtension, "") : decodedBase;
return sanitizeFileSegment(rawStem);
} catch {
return "";
}
}
export function buildFileName(
kind: MediaKind,
index: number,
sourceUrl: string,
extension: string,
fileNameHint?: string,
): string {
const stem = resolveFileStem(sourceUrl, extension, fileNameHint);
const prefix = kind === "image" ? "img" : "video";
const serial = String(index).padStart(3, "0");
const suffix = stem ? `-${stem}` : "";
return `${prefix}-${serial}${suffix}.${extension}`;
}
export function toPosixPath(value: string): string {
return value.split(path.sep).join(path.posix.sep);
}
@@ -0,0 +1,34 @@
import type { Logger } from "../utils/logger";
export type MediaKind = "image" | "video";
export interface MediaAsset {
url: string;
kind?: MediaKind;
role?: "cover" | "inline" | "attachment";
alt?: string;
fileNameHint?: string;
headers?: Record<string, string>;
}
export interface MediaReplacement {
url: string;
localPath: string;
absolutePath: string;
kind: MediaKind;
}
export interface MediaDownloadRequest {
media: MediaAsset[];
outputPath: string;
mediaDir?: string;
log: Logger;
}
export interface MediaDownloadResult {
replacements: MediaReplacement[];
downloadedImages: number;
downloadedVideos: number;
imageDir: string | null;
videoDir: string | null;
}
@@ -0,0 +1,31 @@
declare module "defuddle/node" {
export interface DefuddleResponse {
content?: string;
title?: string;
description?: string;
author?: string;
published?: string;
image?: string;
language?: string;
}
export interface DefuddleOptions {
markdown?: boolean;
}
export function Defuddle(
input:
| Document
| string
| {
window: {
document: Document;
location: {
href: string;
};
};
},
url?: string,
options?: DefuddleOptions,
): Promise<DefuddleResponse>;
}
@@ -0,0 +1,17 @@
declare module "turndown" {
export interface TurndownOptions {
codeBlockStyle?: "indented" | "fenced";
headingStyle?: "setext" | "atx";
bulletListMarker?: "-" | "*" | "+";
}
export default class TurndownService {
constructor(options?: TurndownOptions);
use(plugin: unknown): void;
turndown(input: string): string;
}
}
declare module "turndown-plugin-gfm" {
export const gfm: unknown;
}
@@ -0,0 +1,30 @@
export interface Logger {
info(message: string): void;
warn(message: string): void;
error(message: string): void;
debug(message: string): void;
}
export function createLogger(debugEnabled = false): Logger {
const print = (level: string, message: string): void => {
console.error(`[${level}] ${message}`);
};
return {
info(message: string) {
print("info", message);
},
warn(message: string) {
print("warn", message);
},
error(message: string) {
print("error", message);
},
debug(message: string) {
if (debugEnabled) {
print("debug", message);
}
},
};
}
@@ -0,0 +1,12 @@
export function normalizeUrl(input: string): URL {
try {
return new URL(input);
} catch {
throw new Error(`Invalid URL: ${input}`);
}
}
export function sanitizeFilename(input: string): string {
return input.replace(/[^a-zA-Z0-9._-]+/g, "-").replace(/^-+|-+$/g, "") || "document";
}