eCommerce data

Every product page on the web, as data.

Structured eCommerce data across retailers and marketplaces. Product details, prices and promotions, availability by store, search rankings, reviews, Q&A and sellers from any retailer or marketplace, matched to your catalog and refreshed daily. Delivered as managed feeds, in Aperture, or through the API and MCP so your applications and AI agents can pull it themselves.

Search TermResult RankPriceProduct NameRatingAvailabilityAdditional Info
product page → record walmart.com · store 45202 · rendered 1.3 s extracting
product.recordschema v3 · structured record
{
"gtin": "0 84190 53207 4", "model": "SM150-ER",
"title": "5 Quart Tilt-Head Stand Mixer, Empire Red",
"category": "Kitchen › Mixers › Stand mixers",
"images": 7, "bullets": 6, "description_chars": 1240,
"rating": 4.7, "review_count": 12480, "velocity_7d": 86,
"price": 399.00, "list_price": 449.99,
"promo": { "type": "rollback", "depth_pct": 11.3 },
"availability": "in_stock", "store": "45202",
"pickup": "today", "delivery": "tomorrow",
"variants": ["Empire Red", "Onyx Black", "Silver"],
"size": "5 qt", "unit_price": 79.80,
"seller": "Walmart", "buy_box": true,
"other_sellers": 3, "min_other_price": 409.00,
"search_rank": { "stand mixer": 2, "sponsored": false },
"source": "walmart.com/ip/…/812394011",
"retrieved_at": "2026-09-16T06:12:02Z"
}
matched to catalog by GTIN · 32 fields · 188 tokens1.84M records today
One capture, every eCommerce data type. Sample values. same record via feed, Aperture, API or MCP ·

Structured commerce data for retailers, brands and data platformsProduct, price, stock, rank, reviews and Q&A from one capture.

Every eCommerce data type, from one capture.

A product page is fetched once and yields every field below, typed and matched to your catalog. Add search results and category pages and you have the whole shelf.

catalog

Product Details

The Product Details data feed provides all information from a product page over time, supporting brand consistency and competitive positioning.

price

Prices and promotions

Shelf price, list price, unit price, promotion type and depth, coupons, loyalty and bundle prices, with history.

stock

Availability

Availability and Inventory data feeds help teams track stock status and adjust inventory strategies as availability changes.

search

Product Rankings

The Product Rankings data feed shows how visible products are to customers and how highly they rank in search results.

voice of customer

Reviews

Product Reviews data helps teams monitor customer sentiment, ratings and feedback over time.

voice of customer

Product Q&A

Product Q&A data shows what customers are asking so teams can improve listings and remove friction from purchase decisions.

Choose how you use the data.

Same capture, same matching, same QA underneath. Pick the surface that fits your team: a delivered feed, an application, or an API your systems and agents call directly.

managed

Managed Services

A named delivery team captures, matches, QA's and delivers on your schedule to S3, SFTP, API or webhook, under an SLA.

  • Any retailer, marketplace or distributor
  • Store and zip-level pricing and stock
  • Daily AI QA against your contracted schema
  • Monitoring dashboard and monthly report
See Managed Services
application

Aperture

Pricing intelligence and digital shelf views for retail and brand teams, with the matched catalog maintained for you.

  • Competitor price index and assortment gaps
  • Shelf scorecards by retailer and key term
  • Alerts and evidence pages
  • Exports to CSV, BI and pricing engines
Explore Aperture
api and agents

Platform, API and MCP

Build on the same infrastructure directly, or give your AI agents eCommerce data as tools through the Web Scraper MCP.

extract(url="walmart.com/ip/…", schema=Product)
search("stand mixer", site="target.com", store="45202")
monitor(sites=[…], fields=["price","availability"])
  • REST API and SDKs
  • MCP server for Claude, Cursor and any agent runtime
  • Rendering, anti-bot and egress handled
See the Web Scraper MCP

Built for eCommerce sites as they actually are.

Retail sites are the hardest on the web: aggressive anti-bot, store-specific pricing, JavaScript everywhere, variants nested in variants, and prices that change while you read them. This infrastructure has run them since 2012.

Marketplace scale

Sustained capture on the largest marketplaces in the world without degrading.

450M queries a year, one feed

Store and zip context

Prices, stock, pickup and delivery as a shopper in that location sees them, not a national default.

store 45202 · pickup today

Variants and pack sizes

Every colour, size and count as its own row with parent linkage, so unit prices compare correctly.

per-unit price normalised

Promotions as data

Rollback, clearance, coupon, BOGO, loyalty and bundle prices captured as typed fields with depth and dates.

promo rollback · depth 11.3%

Anti-bot handled

Fingerprinting, rate limits, challenges and geo-blocks handled inside the run, not by your team.

challenge handled inside SLA

JavaScript storefronts

Client-rendered prices, lazy-loaded variants and dynamic promotions resolved before extraction.

rendered 23 requests in 1.3 s

Matched to your catalog

GTIN and model first, then AI attribute and image matching with a confidence score, persisting across runs.

match 0.99 · retailer renames survive

Change events

Price cuts, stock-outs, new sellers and delistings emitted as typed events, not snapshot diffs.

price_cut · oos · new_seller

International storefronts

Any locale, currency and script, fetched from the right region with the right store context.

de fr ja pt-BR 190 locales

Reviews and Q&A at scale

Full text, ratings, dates, votes and verified flags across retailers, deduplicated for sentiment analysis.

Review text · ratings · dates · velocity

Images and content

Image URLs, counts, video presence and enhanced content extracted so listings can be diffed against approved assets.

7 images · A+ content present

Daily QA on every drop

Row counts, fill rates, price sanity and match stability checked against baseline before delivery.

7 checks released 06:41

Built for teams working with eCommerce data.

The same records, six jobs. Each has a dedicated page where the work is deeper.

Pricing and category teams

Competitor prices, assortment gaps and availability matched SKU by SKU into pricing engines and category reviews.

priceassortmentstock

Brand and e-commerce teams

Content compliance, share of search, MAP, availability, reviews and unauthorized sellers across every retailer that carries you.

contentmapsellers

Digital shelf analysts

The six shelf metrics by retailer, store and key term, from one daily capture.

searchrankscorecard

Analytics and data platforms

The capture layer under your product, at retailer scale, under a pass-through SLA.

feedschemasla

Supply and inventory

Competitor stock-outs, lead times and availability by store, so you price into gaps and plan replenishment against the market.

availabilitystore

Data science and AI

Clean, labelled, matched product data for models, agents and search, pulled through the API and MCP on demand.

apimcptraining data

eCommerce data extraction, explained.

eCommerce data extraction is the automated collection of product, price, availability, ranking, review and seller information from retailer websites and marketplaces, structured into records you can analyse, load into systems or feed to models. It powers pricing intelligence, digital shelf analytics, assortment planning, brand protection and, increasingly, AI agents that need to know what is for sale, where, and for how much.

The work has three hard parts. Access: retail sites use fingerprinting, challenges, rate limits and geo-blocks, and render prices in JavaScript. Context: price, stock and rank differ by store, zip, device and login, so a national capture is not a measurement. Matching: the same product is titled, imaged and categorised differently on every site, so records must be matched to a catalog by identifier and attribute before any comparison is valid.

Import.io handles all three as infrastructure, in production since 2012 for the platforms that sell commerce data to the rest of the market. The output is the same record whether you receive it as a managed feed, view it in Aperture, or pull it through the API and MCP.

  • SitesNational retailers, marketplaces, grocery, DIY, electronics, fashion, auto, distributors, DTC, international.
  • FieldsProduct, price, promotion, availability, rank, reviews, Q&A, sellers, categories, images and content.
  • FrequencyDaily as standard. Hourly for key value items, sellers and launches. On demand via API.
  • MatchingGTIN and model first, then AI attribute and image matching with a confidence score.
  • DeliveryS3, SFTP, API, webhook, MCP; CSV, JSON, Parquet; Aperture dashboards and exports.
PROVEN IN PRODUCTION

One capture. Every product-data type your team needs.

Retail and marketplace data operated in production for commerce-intelligence platforms, brands and retailers.

Ask for a reference
2012

In production continuously.

450M

Queries a year on one production feed.

72

Sources under one enterprise SLA.

3

Leading platforms run their capture here.

6

Core product-data types, one capture.

24/7

Managed production operations.

PROOF POINT

72. Sources under one enterprise SLA for a single customer, with query rollover.

eCommerce data questions.

Short answers. Longer ones on the call.

Which retailers and marketplaces can you extract data from?

National retailers, marketplaces, grocery, DIY, electronics, fashion, auto and industrial distributors, DTC stores and international sites. Feasibility per site is confirmed in days, including sites with aggressive anti-bot measures.

What eCommerce data fields do you provide?

Product details, prices and promotions, availability and inventory, search rankings, reviews and ratings, Q&A, sellers and Buy Box, categories and assortment, images and content. A single product-page capture yields 30-plus typed fields.

Can you capture store-level or zip-level prices and stock?

Yes. Prices, promotions, stock, pickup and delivery options are captured in the store or zip context you specify, so records reflect what a shopper in that location sees.

How do you match products across retailers?

GTIN, UPC, EAN, MPN and model first, then AI attribute and image matching with a confidence score on every pair. Low-confidence matches go to a review queue and confirmed matches persist across runs.

Can our AI agents use this data?

Yes. The Web Scraper MCP exposes extraction, search and monitoring as tools for Claude, Cursor and any MCP-compatible agent runtime, with rendering, anti-bot and egress handled by Import.io.

How often is the data refreshed?

Daily as standard. Hourly for key value items, sellers and launches. On demand through the API. Cadence is set per site and per product set.

Is eCommerce web scraping legal?

Collecting publicly available product and pricing information is standard practice across retail. Import.io operates rate-aware collection, respects robots and terms, and works under data processing agreements. Specific requirements are handled in scoping.

How is the data delivered?

As managed feeds to S3, SFTP, API or webhook in CSV, JSON or Parquet; in Aperture as dashboards, alerts and exports; or on demand through the API and MCP.

Send us 50 URLs.
Get 50 records back, matched and typed.