GraphQL API Scraping: APQ, Request Limits and TLS Fingerprints

GraphQL API scraping means collecting open data through a single GraphQL endpoint while controlling requests, limits and network signals. It is used to test your own integrations, monitor public catalogs, collect analytics data and check pipeline stability. In real scenarios, you need to account for access rules, rate limits, disabled introspection, Automatic Persisted Queries and the client's TLS fingerprint.
Unlike classic HTML parsing, GraphQL often hides all logic behind one endpoint such as /graphql. That is convenient for the frontend, but harder for a team building legal data collection: one URL can represent dozens of real operations. If you have already worked with WAF and browser fingerprinting, the logic will feel familiar from Cloudflare web scraping protection, but here query shape, operationName, variables and persisted query hash matter too.
GraphQL API scraping should not be reduced to copying a request from DevTools. At scale, quotas, 429 errors, schema changes, APQ cache behavior, CDN behavior and differences between a browser-like client and a plain HTTP library start to matter. For tasks such as price monitoring at scale, this quickly becomes an engineering problem, not a one-script question.
One layer is often underestimated: the session environment. A platform can see not only tokens and cookies, but also the TLS handshake, header order, IP stability, browser behavior and automation signals. That is why technical teams compare API scraping with LinkedIn automation and controlled sessions: in complex systems, one correct request does not mean the whole session looks consistent.
How GraphQL Differs From REST for Data Collection
GraphQL differs from REST because the client describes the response shape instead of calling many separate URLs. For data collection, this allows a more precise request, but it also makes analysis harder: the endpoint may be the same, while server behavior depends on query text, variables, fragments and the current user's permissions.
In REST, the structure is often visible from URLs: /products, /users, /orders. In GraphQL, you more often see a POST request to one endpoint and a JSON body. Inside it can be operationName, query, variables and extensions for persisted queries. So the first beginner mistake is simple: looking only at the URL and ignoring the request body.
To test your own GraphQL API integration correctly, collect a minimum set of request signals:
- operation name
operationName - query text or persisted query hash
variablesstructure- response size and fields that are actually needed
- HTTP statuses and GraphQL errors in the
errorsfield - authorization headers, cookies and CSRF tokens
- caching conditions at the CDN or application level
In REST, an error is often readable from the HTTP status. In GraphQL, the server can return 200 OK, while the JSON body contains an errors array, partial data or a message about missing permissions. That is why monitoring has to check both status and response content.
| Signal | REST API | GraphQL API | What It Means for Data Collection |
|---|---|---|---|
| address | many endpoints for resources | one or several endpoints | URL explains almost nothing without the request body |
| response shape | defined by server | defined by query | you can request only needed fields |
| errors | more often via HTTP status | HTTP status plus errors in JSON | double checking is required |
| caching | simpler by URL and method | harder because of body and variables | query normalization is needed |
| limits | often per endpoint | often per operation, cost or user | request weight must be counted |
After that inventory, it becomes clear where data collection really relies on an open API and where a script accidentally repeats a private internal frontend call. Those are different risk levels.
How Automatic Persisted Queries APQ Work
Automatic Persisted Queries are a mode where the client sends a request hash instead of the full GraphQL query text. The server finds the saved query in cache or asks the client to repeat the request with the full query if the hash is not known yet. For scraping, this matters because copying only the hash without understanding the related operation often breaks after a frontend update.
APQ appeared as an optimization, not as an anti-bot mechanism. Long GraphQL queries can be heavy, especially when the frontend uses fragments. A hash reduces payload, simplifies caching and gives the server a stable key for known operations. In practice, though, APQ also raises the bar for accidental parsing.
A typical APQ flow looks like this:
- open the page in a browser and find the GraphQL request in Network
- check whether the body contains
extensions.persistedQuery - find
sha256Hashandversion - see whether the full
queryis sent together with the hash - repeat the request in a test environment and record the server response
- check whether the hash changes after a frontend release or build asset change
This chain is easy to see in the diagram: the problem is usually not the hash itself, but the team's lack of visibility into the full path from query to cache and server response.

If the server returns PersistedQueryNotFound, that is not always blocking. Often it is a normal first step of the protocol: the client has to repeat the request with the full query so the server can store the mapping between hash and query text. If you see PersistedQueryNotSupported, the server either does not support APQ or the specific endpoint is configured differently.
The practical takeaway is simple. Do not build an integration only around a hash you saw once in DevTools. Store the mapping between operationName, variables, hash, frontend version and expected response shape. Otherwise, after a small release, you may get a quiet failure: HTTP succeeds, but the data becomes empty or incomplete.
What Disabled Introspection Means
Disabled introspection means the server does not allow the client to fetch the full GraphQL schema through a standard introspection query. This does not make the API invisible, but it removes the easiest way to see types, fields, arguments and relationships between objects.
Some teams disable introspection in production APIs. Open introspection helps developers, but it can also expose extra API structure details to external clients: type names, mutations, deprecated fields, enum values and internal edge cases. That is why many platforms keep it for staging or internal environments.

When introspection is disabled, legal work with your own integration relies on documentation, a contract with the API owner, network traces in your own account and change control. Trying to reconstruct the whole schema from request fragments can violate service terms if you do it on someone else's platform without permission.
For internal and partner projects, this process works better:
- request an official schema or documentation from the API team
- save examples of allowed operationName and variables
- add contract tests for key response fields
- use a separate technical account with minimal permissions
- check schema changes before the frontend release
- log GraphQL errors separately from HTTP errors
If data is business-critical, DevTools should not be your only source of documentation. It shows how the current frontend works, but it does not guarantee that the contract will stay stable tomorrow.
How to Read Request Limits and Error Codes
GraphQL request limits may be counted not only by the number of HTTP calls, but also by operation weight. One short request with deep nesting can be more expensive for the server than ten simple product list requests.
GraphQL often uses cost analysis. The server evaluates query depth, the number of requested fields, pagination arguments, nested connections and user permissions. If a request is too heavy, the response may contain a GraphQL error even with HTTP 200. If there are too many requests, you may see 429 Too Many Requests, a custom GraphQL error code or temporary response truncation.
For stable data collection, check the whole response profile, not just one code:
- HTTP status and retry-after headers
- GraphQL
errors[].messageanderrors[].extensions.code - presence of
datatogether with partial errors - payload size and response time
- pagination cursor change after each page
- repeatability of the error with the same
variablesset
A bad reaction to limits looks like this: a script receives 429, immediately repeats the request ten times, creates extra load and may increase the risk of further request restrictions.
| Signal | What It May Mean | How to Respond |
|---|---|---|
429 | rate limit exceeded | reduce concurrency, add backoff and jitter |
PersistedQueryNotFound | APQ hash not found | repeat with the full query if this is an allowed flow |
GRAPHQL_VALIDATION_FAILED | query does not match the schema | update the contract and check the frontend build |
empty data without HTTP error | no permissions or filter changed | check auth, variables and account scope |
| slow responses | expensive or throttled request | reduce nesting and split the query |
A correct strategy includes exponential backoff, jitter, concurrency limits, caching repeated entity lookups and reducing query depth.

In a production pipeline, these signals should go into a separate log instead of being scattered across access logs, application logs and developer notes. A minimal log should include timestamp, operationName, hash, normalized variables, HTTP status, GraphQL error code, latency, retry count and final result. Without that, the team only sees "the API sometimes fails", which is almost impossible to work with.
What Role the TLS Fingerprint Plays
A TLS fingerprint describes the parameters a client uses to establish a secure connection before the server reads the GraphQL body. Fingerprinting systems may consider TLS version, cipher suites, extensions, ALPN, parameter order and HTTP/2 behavior. If an HTTP client sends the correct query, but its TLS handshake characteristics do not match a typical real-browser network profile, the session may look unusual.
This does not mean everything has to pretend to be a browser at any cost. For your own API, an honest server-to-server client with an access key, stable limits and a transparent user agent is often enough. Problems start when a team takes a private frontend request, runs it through a random library and expects the server to treat it as a normal browser session.
Several layers intersect here:
- TLS handshake and JA3-like signals
- HTTP/2 settings and pseudo-header order
- browser headers, including
sec-ch-ua - cookies, local storage and CSRF tokens
- IP reputation, ASN, geo and proxy type
- session behavior after data retrieval
If data collection runs through browser automation, WebDriver signals, CDP and headless indicators also matter. It is useful to separately review Puppeteer CDP leaks, because the GraphQL request can be correct while the execution environment still exposes automation.

In tests, a TLS fingerprint is better used as a consistency check. Compare how the same allowed request looks from a real browser, Playwright, curl, Node fetch and Python requests. Differences are not a problem by themselves. A problem appears when a session imitates Chrome on macOS, while its network profile looks more like a server library from a data center.
How to Set Up Correct Testing of Your Own GraphQL Integrations
Correct GraphQL integration testing starts with permission, a contract and a controlled environment. If you test your own product, a partner API or open data with allowed access, the main task is not to "push through" the endpoint, but to retrieve the needed fields consistently without adding unnecessary load to the server.
A practical test plan can look like this:
- define the list of allowed operationName values and the business task of each operation
- normalize variables to remove random fields and duplicates
- record APQ mapping between query, hash and client version
- add a request rate limiter at the operation level, not only the endpoint level
- measure latency, payload size and the share of partial errors
- configure backoff for
429, timeout and temporary GraphQL errors - test the session from different environments: browser, Playwright, server-to-server client
- document boundaries: what data is collected, how often and on what legal basis
For browser automation, it is better not to mix all tasks in one profile. A separate profile for each test account, stable cookies, a predictable proxy route and a controlled extension set produce cleaner results. If you test behavior through Playwright, Puppeteer or Selenium, keep the Playwright, Puppeteer and Selenium material nearby so you do not confuse API errors with browser automation signals.

A small detail matters: do not change several parameters at once. If errors start after moving from a real browser to an HTTP client, check auth, cookies, headers, APQ hash, variables, TLS profile, IP and concurrency one by one. That is how you can identify which exact change caused the access error.
How Afina Helps With GraphQL API Scraping
Afina fits this process not as a "magic bypass", but as an environment for controlled browser sessions. When a team checks its own GraphQL integrations through a web interface, it needs isolated profiles, stable cookies, proxies, team access and repeatable scenario execution. This helps separate API errors from environment problems: one profile, one test account, one traffic route, one change log.
For a server-side API client, Afina does not replace a normal contract, rate limiter or documentation. But for browser checks, QA sessions and scenarios where GraphQL requests are born inside the frontend, controlled profiles reduce chaos in tests. This material is provided for informational and educational purposes only.
DownloadFAQ — Frequently Asked Questions
What is GraphQL API scraping?
GraphQL API scraping is collecting open or permitted data through a GraphQL endpoint. It requires control over query, variables, limits, access rights and network signals.
Why is GraphQL more complex than REST for data collection?
GraphQL is more complex because one endpoint can execute many different operations. You need to analyze the request body, operationName, variables and GraphQL errors.
What are Automatic Persisted Queries APQ?
APQ is a mechanism where the client sends a GraphQL request hash instead of the full query. If the server does not know the hash, it may ask the client to repeat the request with the full text.
Why do GraphQL APIs disable introspection?
Introspection is disabled to avoid exposing the full production API schema to external clients. It does not hide every request, but it removes the easiest way to browse types and fields.
How can I tell that a GraphQL request hit a rate limit?
Signals may include 429, a custom GraphQL error code, partial data, growing latency or an empty response. Check both HTTP status and the JSON errors field.
What does a TLS fingerprint mean in GraphQL?
A TLS fingerprint describes how the client establishes a secure connection to the server. It can distinguish a browser session from a server-to-server library before the GraphQL body is analyzed.
Can GraphQL APIs be scraped without permission?
It depends on the data source, service terms and applicable law. For work tasks, use your own APIs, open data or access explicitly allowed by the resource owner.
Why does GraphQL return 200 OK with an error inside JSON?
GraphQL can communicate transport success through HTTP 200 and business errors through the errors field. Checking only HTTP status does not show the real operation result.
