scrape-codegen-generate
SkillGenerate web-poet page object code from per-page extraction analyses
Install
git clone https://github.com/zytedata/claude-skills.git ~/.claude/skills/scrape-codegen-generateWhat is scrape-codegen-generate?
Generate web-poet page object code from per-page extraction analyses
What this can do
Capabilities declared in this component's own frontmatter — not inferred.
Run shell commands
Declares Bash
Create and modify files
Declares Write or Edit
~17 tokens of context used while enabled, before you invoke anything
All declared tools (4)
BashReadSkillWriteDocumentation
README · ~4 min readYou are generating web-poet page object code. You receive per-page extraction analyses (from Stage 1) that describe WHERE and HOW each field can be extracted from pages on a given domain. Your job is to synthesize these analyses into a single page object class that works across the entire domain.
Input
The raw argument string is $ARGUMENTS. Split it into 3 whitespace-separated positional arguments:
- work_path: directory containing Stage 1 analysis files, e.g.
.scrape/.work/spec - output_path: where to save the generated page object, e.g.
.scrape/spec/page_object.py - spec_path: path to spec.json file, e.g.
.scrape/spec/spec.json
Plus, taken from the surrounding prompt text (not from the argument string):
- fields: optional, specific fields to generate (provided in the prompt as "Only generate these fields: ..."). When set, only generate
@fieldmethods for those fields. When not set, generate all fields found in the analyses.
Process
1. Read inputs
Reviews
Log in to leave a review.
No reviews yet — be the first.
Explore related
Other things in this space — across every part of the ecosystem, not just skills.
Skillssimilar to this one
All skills →link-rot-scanner
Crawl and validate all internal and external links across an AEM Edge Delivery Services site. Uses the query index or sitemap to discover pages, extracts links from .plain.html renditions, checks HTTP status codes, and produces a prioritized report of broken, redirecting, and insecure links. Use whe
151 stars
extract
Crawl an existing website (capped, multi-page) and seed stardust/current/ with PRODUCT.md, DESIGN.md, DESIGN.json, a per-page inventory, and the consolidated brand surface — the captured design system, palette, typography, motifs, and voice of the live site. Use when the user wants to analyze an exi
151 stars
scrape-webpage
Use this when the page-import pipeline needs to fetch a source webpage and prepare it for import/migration to AEM Edge Delivery Services. Covers scraping content, extracting metadata, downloading images, and returning analysis JSON with paths, metadata, cleaned HTML, and local images. Do not invoke
151 stars
MCP Servers
All mcp servers →Plugins
All plugins →exa
Exa AI web search, deep research, and content extraction. Provides MCP tools and research skills for comprehensive web search, people discovery, company research, academic papers, and more.
4.8K stars
cli
CLI and Agent Skill for Firecrawl - Add scrape, search, and browsing capabilities to your AI agents
549 stars
tavily
Build AI applications with real-time web data using Tavily's search, extract, crawl, and research APIs.
436 stars