HomeBlogScraping
Scraping

Build vs Buy: What Running a Google Maps Scraper Actually Costs

The first version takes a weekend, which is what makes the decision deceptive. The token layer, the 120-result ceiling, silent parser failures, and what maintenance really costs.

Livescraper TeamSep 28, 202614 min read
build vs buy google maps scraper

The first version of a Google Maps scraper takes a competent engineer about a weekend. That is the fact that makes this decision so much harder than it looks. The weekend build works, it returns real data, and it produces a demo that makes the whole buy-versus-build conversation feel settled. Then it runs for three months and someone discovers that the output has been quietly wrong for five weeks.

This is not an argument that you should never build one. Some teams should, and the last section says exactly which. It is an argument that the thing you are choosing to own is not the scraper. It is the ongoing obligation to keep a scraper correct against a target that changes without telling you, and the cost of that obligation is almost never estimated at the point the decision gets made — because at that point you are looking at a working demo.

Here is what is actually in the gap between the two.

What the weekend build doesn't include

The demo scraper reads pages and extracts fields. Everything below is what turns that into something you can depend on, and each item is a real engineering problem rather than a line of configuration.

Getting to the data at all

Google Maps is not a set of static documents. The listing content you want is loaded by the page's own JavaScript through internal endpoints, and those endpoints are not open — requests carry a short-lived, cryptographically generated token produced by an anti-automation system running in the browser. You cannot construct that token; it has to be produced by a real browser environment, then used before it expires.

That single fact shapes the entire architecture. You now need a browser in your stack, not just an HTTP client. You need to drive that browser through a path that produces a valid token, keep the token warm, and detect the moment it stops being accepted. The token acquisition step becomes a component with its own failure modes, its own retry policy, and its own monitoring — and when it fails, it fails for every job at once rather than for one record, so it is a single point of failure sitting underneath everything else.

Teams almost never budget for this because it is invisible from the outside. It is also the component most likely to break, because it is the one the target actively maintains against automation.

The 120-result ceiling and what it forces

A single Google Maps search does not return an unlimited result set. In practice you get somewhere around 120 results per query regardless of how many businesses actually match. Search "dentist in Berlin" and you will not get Berlin's dentists; you will get 120 of them, and no error telling you the rest exist.

Complete coverage therefore requires decomposing one logical query into many real ones — by district, by postcode, by grid cell, by narrower category — until each sub-query comes back under the ceiling. That means writing a subdivision strategy, detecting when a sub-query is still capped and splitting further, and then dealing with the consequence: massive overlap between adjacent searches.

Which means deduplication, which has its own trap. Matching on name and address fails constantly: the same practice appears as "Dr. Schmidt Zahnarztpraxis" and "Zahnarztpraxis Dr. med. dent. Schmidt," with addresses formatted three different ways. You need a stable identifier — Google's own place ID or CID — and you need to have kept it on every record from the start. Teams that deduplicate on name and address end up with inflated counts and no way to tell which of two near-identical rows is current. Our guide to place IDs and CIDs covers why these identifiers, not the business details, are the right key for anything you intend to re-collect.

Silent failure: the expensive one

If you take one thing from this article, take this. The failure mode that costs real money is not the scraper crashing. A crash is cheap, because you find out immediately.

The expensive failure is this: the page structure shifts slightly, your parser no longer finds the field it expects, it returns an empty string instead, and the job completes successfully with a green status. Your database fills with records where phone is blank. Nothing alerts, because nothing failed. The pipeline reports success. Six weeks later someone in sales asks why so many rows have no phone number.

There is a nastier variant. A job can return zero records and report success, which is indistinguishable from a legitimate "this search genuinely had no matches" unless you designed for the distinction. We have seen production systems where a collection run that hit a hard blocking wall and got nothing reported as "completed, 0 records" — technically accurate, operationally a lie, and completely invisible until someone compared it against last month's numbers by hand.

Guarding against this is not a monitoring dashboard. It is a set of correctness assertions running on every job: expected field fill rates, record counts within a tolerance of the previous run for the same query, a canary query with a known-good answer, and an explicit distinction in your data model between "collected nothing" and "found nothing." That is real work, nobody builds it in the first version, and it is the difference between a scraper you can trust and a scraper you merely operate.

Partial completion and resumability

Small jobs either work or fail. Large jobs do something worse: they get 78% of the way through 40,000 records and then stop.

Now you need to answer questions your weekend build has no way to answer. What completed? What was in flight when it died? If you re-run it, do you duplicate the first 78%? If you resume, how do you know where the boundary was? If two workers picked up the same chunk, did you charge yourself twice and write both?

The answer is a job model rather than a script: work split into chunks with stable identifiers, per-chunk status, idempotent writes so re-processing a chunk is harmless, and a finalizer that assembles the finished chunks into one output and knows the difference between "done" and "done except for chunk 47." This is ordinary distributed systems work, it is well-understood, and it is several weeks of it — for the plumbing around a scraper that already worked.

Proxies, blocking and the retry subtlety

Sustained collection from one IP address degrades and then stops. So you need a proxy pool, and with it: health checking, removing dead proxies, distributing requests so no single exit is hammered, and handling the fact that proxy quality varies enormously and cheap pools are frequently already burned by someone else's traffic.

One non-obvious detail worth the paragraph, because it catches nearly everyone. When a request is refused because you have been rate limited, the instinctive response — retry immediately, maybe three times — is close to useless. You are hitting the same limiter within the same window, so all three attempts fail, and you have converted one failure into a slower failure. Retries need genuine backoff spread across enough time for the limiter to release, which means a job that recovers takes minutes rather than milliseconds, which means your timeouts and your job model have to accommodate that. Getting this wrong produces a scraper that appears to have exhausted every option while never really having tried twice.

Related trap: a fallback path that hits the same wall is not resilience. If your primary method is blocked and your fallback method is blocked for the same underlying reason, all the fallback adds is latency before the same failure. A fallback is only worth having if it fails independently — and verifying that is itself work.

The rest of it

  • Consent and regional variation. Requests from different countries hit different consent flows and return different page structures. A scraper working from a US IP may not work from a German one.
  • Field normalisation. Phone formats, address components, opening hours across timezones and locales, unicode in business names. Tedious, endless, and where data quality actually lives.
  • Encoding. Downstream consumers are unforgiving about it, and a byte-order mark in the wrong place breaks every JSON parser that touches your output.
  • Reviews are a second scraper. Paginated, separately gated, with their own sort and filter semantics that behave differently from the business listings.
  • Cost control. Without a pre-flight estimate, a mis-specified job can run for eleven hours before anyone notices what it is doing.

What it costs, with the arithmetic shown

Illustrative figures, and your numbers will differ — but the shape holds, and the shape is the point.

Initial build. A demo is a weekend. Something you would put in front of a paying customer, with the token layer, subdivision, deduplication, a job model and correctness checks, is realistically three to six weeks of senior engineering time. At a blended cost of around $80/hour that is roughly $10,000 to $19,000 — spent once, on work that produces no differentiation for your product because your competitors' scrapers return the same public data.

Running infrastructure. Proxies from $50 to several hundred a month depending on quality and volume, plus compute for browser instances, which are memory-hungry, plus storage and log retention. Call it $100 to $700 a month.

Maintenance. This is the number people leave out, and it is the one that dominates. Budget two to four engineering hours a week for parser fixes, proxy churn, failed-job triage and verifying that output still matches expectations. At the same rate that is $700 to $1,400 a month, or $8,400 to $16,800 a year — recurring, forever, and disproportionately landing in weeks when something more important is happening.

Now the comparison. At $0.002 per record, a 100,000-record pull costs $200. Set aside the build cost entirely and compare only against one year of DIY maintenance:

Annual volumeManaged costvs DIY maintenance alone ($8.4k–16.8k/yr)
100,000 records$200Managed is 40–80× cheaper
1,000,000 records$2,000Managed is 4–8× cheaper
5,000,000 records$10,000Roughly comparable
10,000,000 records$20,000DIY may be cheaper on cash

So the honest breakeven sits somewhere around four to eight million records a year — and only if you value the engineering time at zero for the build, and only if nothing goes wrong in a way that costs you a customer. Below a few million records annually, building is not a cost saving. It is a decision to spend engineering capacity on infrastructure instead of on your product, and it should be made for a reason other than money.

When building is genuinely the right call

There are real ones, and a vendor pretending otherwise is not worth listening to. Build if:

  • The extraction itself is your product. If you are selling data or data infrastructure, this is your core competency and outsourcing it outsources your moat.
  • You need something no tool exposes. An unusual field, a bespoke traversal pattern, a source nobody supports. Real constraints, and worth owning.
  • Data cannot leave your network. Regulatory or contractual isolation requirements that a hosted service cannot satisfy, whatever its compliance posture.
  • Your volume is genuinely enormous. Above the breakeven and sustained, with the engineering capacity already in-house.
  • Scraping is already a maintained competency. If you run twelve scrapers, the thirteenth is marginal cost against a platform you already staff. That is a materially different calculation from a first one.

And the reverse, stated plainly because it belongs in a fair comparison. Buying has real costs: you depend on a vendor's continuity and roadmap, you get the fields they expose rather than any field you can imagine, you work within their rate and volume limits, and you accept their API shape instead of designing your own. If any of those is disqualifying for you, that is a legitimate reason to build, and it has nothing to do with cost.

The middle option people forget

This is not binary, and the third option is frequently the right one. You can own your pipeline and not own the extraction layer.

The pattern: a managed service handles token acquisition, subdivision, deduplication, proxies, retries and output correctness, and hands you clean records through an API. Everything downstream — your schema, your enrichment, your scoring, your storage, your scheduling, your business logic — stays yours and stays in your repository. You keep all the parts that are specific to your product and none of the parts that are identical across every company doing this.

Concretely, that means the Google Maps Data Scraper for business records with coordinates and place IDs intact, the Reviews Scraper for review histories, and the Email & Contact Scraper for published contact details, with the big jobs split and reassembled server-side and a cost estimate before anything runs. Your code consumes results and does the part that is actually yours. Most teams who think they need to build a scraper actually need this, and discover it after the build.

A worked decision

A twelve-person property-tech company wants commercial premises and local amenity data across 40 UK towns, refreshed quarterly. Roughly 250,000 records a year.

The build case: two engineers, six weeks, so about $19,000 — except it displaces the mapping feature those engineers were hired for, which is the actual cost and it is not measured in dollars. Then $300 a month of infrastructure and a standing three hours a week of maintenance, which in practice means one of two engineers is permanently the scraper person and gets interrupted at inconvenient times. Annual running cost around $15,000, and a quarterly refresh that is now a risk event rather than a routine task.

The buy case: 250,000 records at $0.002 is $500 a year. Two days of integration work. The quarterly refresh is a scheduled job. Nobody is the scraper person.

The decision is not close, and the money is not why. It is that the company's differentiation is what it does with location data, not its ability to obtain location data, and six weeks of its only two engineers is a substantial fraction of a year's capacity for building the actual product. The version of this where building is right is the one where they intend to sell the dataset itself — and then the extraction layer stops being overhead and becomes the thing.

Conclusion

The weekend build is real and it works, which is exactly why this decision goes wrong so often. What you are choosing is not whether to write a scraper. It is whether to take on permanent ownership of token acquisition against an adversarial anti-automation system, a subdivision strategy to get past the per-search result ceiling, deduplication on stable identifiers, a resumable job model for partial failures, a proxy pool, backoff semantics that actually work, and — most importantly — correctness monitoring capable of detecting that your output went quietly wrong while every status light stayed green.

Below a few million records a year, that ownership is not cheaper than buying; it is a reallocation of your scarcest resource away from your product. Above it, or when extraction is your product, or when the data genuinely cannot leave your network, building is the right answer and the cost is worth it. The mistake is not choosing either one. The mistake is choosing based on how easy the first version looked.

If what you want is the data rather than the pipeline, Livescraper runs that extraction layer as a service at $0.002 a record with 500 records free to test the output against whatever you were about to build — which is a cheaper way to find out where the gaps are than discovering them in week five.

Related reading: How to Scrape Google Maps: A Step-by-Step Guide, How to Scrape Google Maps Places in Python, Free Google Maps Scraper or API: Which Is Better for Business Data?.

Frequently asked questions

Should I build my own Google Maps scraper?

A working demo takes a competent engineer about a weekend, which is what makes the decision deceptive. What you are choosing to own is not the scraper but the permanent obligation to keep it correct against a target that changes without notice. Below roughly four to eight million records a year, building is not a cost saving, it is a reallocation of engineering capacity away from your own product.

What does a DIY Google Maps scraper actually cost?

Illustratively: three to six weeks of senior engineering for a production-grade build, around $10,000 to $19,000 once. Then $100 to $700 a month of proxies and compute, plus two to four engineering hours a week of maintenance, which is $8,400 to $16,800 a year recurring. Maintenance is the number teams leave out of the estimate and it dominates the total.

Why can't I just make HTTP requests to Google Maps?

The listing content is loaded by the page's own JavaScript through internal endpoints, and those requests carry a short-lived token generated by an anti-automation system running in the browser. You cannot construct it, so a real browser environment has to produce it. That makes token acquisition a component with its own failure modes, and because it fails for every job at once it is a single point of failure beneath everything else.

Why does my scraper only return about 120 results per search?

A single Google Maps search caps the result set at roughly 120 regardless of how many businesses match, with no error telling you the rest exist. Complete coverage requires decomposing one logical query into many narrower ones by district, postcode, grid cell or category until each comes back under the ceiling, then deduplicating the heavy overlap between adjacent searches.

What is the most expensive scraper failure mode?

Silent incorrectness, not crashes. A crash is cheap because you find out immediately. The expensive case is a page structure shifting slightly so the parser returns empty strings, the job completing with a green status, and nobody noticing for weeks. A job can even return zero records and report success, which is indistinguishable from a genuinely empty search unless you designed for that distinction.

How should I deduplicate Google Maps records?

On a stable identifier such as the place ID or CID, kept on every record from the start. Matching on name and address fails constantly because the same business appears under several name variants with differently formatted addresses, which produces inflated counts and no way to tell which of two near-identical rows is current.

Why don't immediate retries fix rate limiting?

Retrying instantly hits the same limiter inside the same window, so every attempt fails and you have only made the failure slower. Retries need genuine backoff spread over enough time for the limiter to release, which means recovery takes minutes and your timeouts and job model must accommodate it. A fallback path blocked for the same underlying reason adds latency, not resilience.

When is building genuinely the right choice?

When extraction is your product and outsourcing it outsources your moat, when you need a field or traversal no tool exposes, when data cannot leave your network for regulatory reasons, when volume is sustained above the breakeven, or when you already run several scrapers so the next one is marginal cost on a platform you already staff.

Is there an option between building and buying?

Yes, and it is often the right one: own your pipeline but not the extraction layer. A managed service handles token acquisition, subdivision, deduplication, proxies, retries and output correctness, and returns clean records through an API. Your schema, enrichment, scoring, storage and business logic all stay in your own repository.

Livescraper Team
Practical writing on Google Maps data, scraping techniques and lead generation - from the Livescraper team.