[ Case study · Data Infrastructure ]

Launching an AI Scraping SDK on PyPI

Internal Product · AI Product Engineering

[ Data Infrastructure · AI Product Engineering ]

Internal Product

Inscrape package page on the Python Package Index

The problem

Web scraping remains one of the most common data engineering tasks, yet the developer experience is still poor — CSS selectors break, browser automation is slow and heavy, and output is unstructured. We wanted to build a scraping SDK that makes extraction as simple as an API call, with AI powering the extraction logic so it works reliably across sites without site-specific configuration.

What we built

Inscrape turns web scraping from a fragile, maintenance-heavy process into a simple API call. The AI-powered extraction layer handles the complexity — developers don't write selectors, don't manage browsers and don't maintain site-specific code. They get structured data back in the format they need.

  1. API & SDK DesignDesigned an API-first architecture where the AI-powered extraction runs server-side, and the Python SDK is a thin, well-typed client. Three-line usage pattern: init client, call scrape, get structured data. Typed exceptions for every failure mode (auth, rate limits, quota).
  2. Extractor DevelopmentBuilt specialised extractors for social media profiles (Instagram, X/Twitter) that return structured JSON with follower counts, bios, engagement metrics and post data. General URL extractor returns structured content, Markdown and screenshot options.
  3. Async & PublishingAdded full async support via AsyncInscrape for high-throughput pipelines. Comprehensive test suite with pytest and pytest-asyncio. Linting with Ruff. Published on PyPI with Hatchling build system and full documentation.
v0.1.0Published on PyPI (beta)
3 linesTo scrape any URL
AsyncFull AsyncInscrape support
TypedComplete type hints and error handling
PythonhttpxAsyncIOHatchlingpytestRuff

[ Your turn ]

Have a hard problem?
Let’s build the answer.