Skip to content

jwu/pi-webfetch

Repository files navigation

pi-webfetch

A pi package that adds a webfetch tool for reading web URLs as clean Markdown, HTML, text, or YouTube metadata JSON.

webfetch is optimized for agent use:

  • Direct image responses are downloaded as visual input for image-capable models, with inline terminal preview where supported.
  • General web pages are cleaned into readable content.
  • GitHub URLs are fetched with gh for better repository, issue, PR, file, and directory results.
  • YouTube URLs are fetched with yt-dlp for video metadata, transcripts, playlists, and channel listings.

Install

pi install npm:@johnnywu/pi-webfetch

Or via local path in ~/.pi/agent/settings.json while developing:

{
  "packages": ["~/dev/jwu/pi-webfetch"]
}

Requirements

Install the optional CLI tools for the URL types you want to support:

URL type Required executable Notes
Direct images none Detected by response Content-Type; JPEG, PNG, WebP, GIF, and AVIF up to 20 MiB
General web pages scrapling Used for non-GitHub, non-YouTube, non-image URLs
GitHub / Gist gh Must be installed and authenticated for private or rate-limited content
YouTube yt-dlp Used for videos, transcripts, playlists, Shorts, and channels

Defuddle is bundled as an npm dependency and is used by default to improve Markdown output for general web pages.

Tool

webfetch

Fetch and clean an HTTP(S) URL.

Parameter Type Default Description
url string required HTTP(S) URL to inspect and fetch
mode markdown | html | text | json markdown Output mode. json is supported for YouTube results.

Examples:

{ "url": "https://example.com/article" }
{ "url": "https://github.com/jwu/pi-webfetch" }
{ "url": "https://www.youtube.com/watch?v=PIdETjcXNIk" }
{ "url": "https://www.youtube.com/@Brandon-Melville", "mode": "json" }

What you get

Direct images

For a general URL whose HTTP response declares image/avif, image/gif, image/jpeg, image/png, or image/webp, webfetch downloads the image directly instead of sending binary data to Scrapling. It returns the image as a Pi image content block for visual analysis and, in terminals that support inline images, displays a preview. Downloads are limited to 20 MiB.

General web pages

Default output is readable Markdown. You can request raw-ish cleaned HTML or plain text with mode: "html" or mode: "text".

GitHub URLs

GitHub URLs are routed through gh, so common GitHub pages return useful CLI/API content instead of noisy browser HTML. Supported URL shapes include:

  • repositories
  • users
  • issues
  • pull requests
  • releases
  • Actions runs
  • gists
  • files and directories
  • commits and other API-backed paths

YouTube URLs

YouTube URLs are routed through yt-dlp.

Supported inputs include:

  • video URLs
  • youtu.be short links
  • playlists
  • channel handles, for example https://www.youtube.com/@name
  • channel Videos / Shorts / Streams tabs

For videos, webfetch returns metadata and tries to include a transcript. Missing transcripts do not fail the request.

For playlists, webfetch returns a flat list of entries.

For channel root URLs, webfetch expands available Videos, Shorts, and Streams tabs and merges their entries into one channel result.

mode: "json" returns a curated stable JSON shape for YouTube video, playlist, or channel data.

Configuration

Add webfetch settings to .pi/settings.json (project) or ~/.pi/agent/settings.json (global):

{
  "webfetch": {
    "useDefuddle": true,
    "qualityJudge": false,
    "qualityJudgeModel": "google/gemini-2.5-flash",
    "qualityJudgeThinkLevel": "off"
  }
}

Project settings override global settings. The dotted key form also works:

{
  "webfetch.useDefuddle": true,
  "webfetch.qualityJudge": true,
  "webfetch.qualityJudgeModel": "google/gemini-2.5-flash",
  "webfetch.qualityJudgeThinkLevel": "off"
}

Settings

Setting Default Description
webfetch.useDefuddle true Use Defuddle to convert cleaned HTML to Markdown for general web pages. Set false to use Scrapling Markdown directly.
webfetch.qualityJudge false Ask a model to reject unusable fetched Markdown, such as boilerplate, captcha/challenge pages, or unrelated content.
webfetch.qualityJudgeModel current pi model Optional judge model in provider/model form.
webfetch.qualityJudgeThinkLevel off Optional judge thinking level: off, minimal, low, medium, high, or xhigh.

Output behavior

  • Only http:// and https:// URLs are accepted.
  • Missing CLI executables return a friendly failed tool result.
  • Output is truncated with pi's standard limits: 2000 lines or 50 KiB, whichever is hit first.
  • If output is truncated, the full extracted content is saved to a temp file and the path is included in the result.

Internal docs

Implementation details are documented in:

Development

# Install dependencies
bun install

# Run tests
bun test

# Type check
bun run typecheck

# Format
bun run format

# Release (local, requires GH_TOKEN and NPM_TOKEN)
bun run release

This project uses semantic-release with conventional commits.

License

MIT

About

Simple webfetch extension for Pi

Resources

License

Stars

1 star

Watchers

0 watching

Forks

Packages

 
 
 

Contributors