I will scrape ecommerce data in bulk and build a reliable web crawler
Business Web Apps and Ecommerce Data Collection
Over deze dienst
Get structured ecommerce data at scale, not another manual spreadsheet task.
I build crawlers for public or authorized marketplace data, including category discovery, pagination, product-to-seller mapping, deduplication, checkpoints and CSV/JSON exports.
Documented work includes a finalized Gmarket dataset of 32,820 distinct products and 4,203 sellers, plus a Coupang run across 13 categories with 11,101 product entries and 819 deduplicated sellers. These are past project results, not promised volume or speed for your site.
My experience includes Akamai-protected Coupang environments: detecting access challenges, recording failures and recovering interrupted runs. Access conditions must be tested; I do not promise universal bypass.
BASIC: One-source pilot, up to 500 records and extraction code.
STANDARD: Up to 2 agreed sources, up to 10,000 records, deduplication, resume support and run notes.
PREMIUM: Up to 3 agreed source structures, 50,000 records, resumable batch collection, validation and handover.
Source code is included. Proxy, data-service and hosting fees are separate. Send target URLs, fields and intended volume before ordering so I can confirm feasibility.
Technologie:
Python
•
Toneelschrijver
Techniek:
Geautomatiseerd
Mijn portfolio
Veelgestelde vragen
Have you worked with Akamai-protected ecommerce sites?
Yes. Project records document Coupang collection in an Akamai-protected environment, with challenge classification, diagnostics and recovery. Results depended on the tested environment and access conditions; this is not a guarantee for every protected site.
What do the portfolio collection numbers mean?
Gmarket: 32,820 distinct products and 4,203 sellers in a recovered/finalized dataset, not all planned categories. Coupang: 11,101 category-level product entries across 13 categories and 819 deduplicated sellers. These are separate past runs, not speed promises.
Are proxy and third-party service costs included?
No. Any necessary proxy, data-provider or hosting costs are agreed separately before ordering. Feasibility depends on the target site, permitted access and available fields.
Can you collect any requested number of records?
Scope caps: Basic 1 source/100 pages/500 records; Standard 2 sources/2000 pages/10000 records; Premium 3 sources/10000 pages/50000 records. Collection stops at the agreed cap. Available records and source feasibility are confirmed before ordering.
What data and files will I receive?
The agreed accessible fields, a consistent CSV/JSON schema, deduplicated results, source code and run notes. Public product identifiers, category links and business seller records are examples; availability is checked per source.
What happens if collection is interrupted?
Standard and Premium include resume/checkpoint support for the agreed workflow and logs for incomplete work. A site redesign or new access restriction may require a separately scoped update; ongoing maintenance is not included.

