mirror of
https://github.com/JimLiu/baoyu-skills.git
synced 2026-08-08 09:53:02 +08:00
feat(baoyu-url-to-markdown): add URL-specific parser layer for X/Twitter and archive.ph
- New parsers/ module with pluggable rule system for site-specific HTML extraction - X status parser: extract tweet text, media, quotes, author from data-testid elements - X article parser: extract long-form article content with inline media - archive.ph parser: restore original URL and prefer #CONTENT container - Improved slug generation with stop words and content-aware slugs - Output path uses subdirectory structure (domain/slug/slug.md) - Fix: preserve anchor elements containing media in legacy converter - Fix: smarter title deduplication in markdown document builder
This commit is contained in:
@@ -0,0 +1,10 @@
|
||||
import { archivePhRuleParser } from "./archive-ph.js";
|
||||
import { xArticleRuleParser } from "./x-article.js";
|
||||
import { xStatusRuleParser } from "./x-status.js";
|
||||
import type { UrlRuleParser } from "../types.js";
|
||||
|
||||
export const URL_RULE_PARSERS: UrlRuleParser[] = [
|
||||
archivePhRuleParser,
|
||||
xArticleRuleParser,
|
||||
xStatusRuleParser,
|
||||
];
|
||||
Reference in New Issue
Block a user