WEBROBOT

Web Scraping Security and Compliance: How WebRobot Handles Your Data

Traffic runs over TLS. Stored data and any target-site credentials you give a robot are encrypted at rest and used only to run your robots. Internal access follows least privilege. You control retention and can delete your data whenever you want, and we never sell, resell or share the rows your robots extract. Crawls honor robots.txt and polite rate limits.

This page is written for the person whose signature the purchase needs: security, legal, or the ops lead who has to explain the tool to both. It describes practices and policy commitments in plain language, and it is explicit about what we do not claim.

Run the robot

Last updated July 2026

01 / POSTURE FIG. 1 · DATA HANDLING

What we do with your data, in six lines

Scraping tools sit in an uncomfortable place: they hold your credentials, they touch other people's websites, and they produce data your business then acts on. Here is exactly how we treat each part.

IN TRANSIT

Encrypted connections

Traffic between your browser, the WebRobot application and your delivery targets runs over TLS. Nothing about your robots, your rows or your credentials moves over a plaintext channel.

AT REST

Encrypted storage

Stored data is encrypted at rest, including the target-site credentials a login robot needs. Those credentials are used for one purpose only: signing that robot in to the site you pointed it at.

ACCESS

Least privilege inside

Internal access is limited to the people who need it to operate the service or to answer a support request you opened. Inside your workspace, seats are yours to grant and revoke.

RETENTION

You hold the delete key

Extracted datasets and stored credentials can be deleted from the workspace whenever you decide, including before you cancel. Retention is a setting you control, not a policy you inherit.

USE OF DATA

No resale, no sharing

The rows your robots extract are yours. We do not sell them, share them with other customers, or use them to build a dataset we license to anyone else. That is a commitment, in writing, on this page.

CONDUCT

Polite crawling by default

Robots honor robots.txt and pace their requests. We do not build features whose purpose is to defeat a site access control, and we do not add them on request.

02 / SHARED RESPONSIBILITY FIG. 2 · WHO OWNS WHAT

What WebRobot handles, and what stays with you

No scraping vendor can take on the legal weight of a target you chose. This table is the honest split. Read it before you buy, not after your first login robot goes live.
Area WebRobot handles You are responsible for
Platform security TLS in transit, encryption at rest, least-privilege internal access, patching the service Keeping your own workspace accounts secure and using strong, unique passwords
Target selection Honoring robots.txt and rate limits on whatever you point us at Choosing lawful targets and reading their terms of service before you scrape them
Login credentials Storing them encrypted and using them only to run your robots on the site you configured Supplying credentials you are entitled to use, ideally a dedicated least-privilege account
Personal data Processing what your robots collect on your instruction, and deleting it when you ask Your lawful basis, purpose limitation, retention policy and deletion requests as controller
Crawl conduct Human-paced requests, no tooling built to defeat access controls Not asking us to bypass a site that has told you no
Data retention Deleting datasets and credentials when you delete them, and on request after cancellation Deciding how long you keep extracted data, in WebRobot and in your own systems
Team access Seat-based workspace access on every plan Granting and revoking seats when people join or leave your team
Downstream use Delivering clean rows to CSV, Sheets, Slack, Zapier, webhooks or the API What you do with the rows: republication, resale and enrichment are your call and your risk

The pattern is simple. We are responsible for the machine and how it behaves. You are responsible for where you aim it and what you do with what comes back. If you want the mechanics of a run before you sign off on it, the how it works page shows what the robot does on each pass, and the web scraping tool overview covers the self-healing behavior your team will ask about.

03 / CREDENTIALS FIG. 3 · LOGIN ROBOTS

How we treat the credentials you hand a robot

Agent actions (logins, forms, multi-step flows) are the reason a scraping tool ever touches a password. Here is what happens to it.

When you set up a robot that signs in to a target site, you store a username and password for that site in your workspace. Those credentials are encrypted at rest and used for exactly one thing: authenticating that robot, on that site, during your runs. They are not shared with other customers, not reused for other purposes, and not exported anywhere.

You can delete them at any time. Deleting them stops the robot from logging in on the next run, which is the behavior you want when someone leaves your team or a vendor relationship ends.

Our advice, which costs us nothing to give and saves both of us trouble: create a dedicated account on the target site for the robot, give it read-only or the lowest role that still sees the data, and rotate it on the same schedule as your other service accounts. If the target site supports API keys or a data feed, use that instead of a login. Scraping behind a login is a legitimate technique, and it is also the place where terms of service matter most, so read them first. Our general take on the law is in the piece on whether web scraping is legal.

CREDENTIAL CHECKLIST

  • Use a dedicated account, not a personal one
  • Grant the lowest role that can see the data
  • Rotate on your normal service-account schedule
  • Delete from the workspace when the robot retires
  • Prefer an official API or feed where one exists
See what agent actions can do
04 / CONDUCT FIG. 4 · CRAWL ETHICS

How our robots behave on someone else's website

A scraper is a guest on infrastructure it does not pay for. That posture drives three rules we do not trade away.

RULE 1

robots.txt is honored

Robots read and respect robots.txt directives on the sites they visit. If a path is disallowed, the robot does not crawl it, whatever the sentence you wrote asked for.

RULE 2

Requests are paced

Runs are rate limited by default so a crawl looks like a visitor rather than a flood. Slower and finished beats fast and blocked, and it keeps the target site healthy for its actual users.

RULE 3

No bypass tooling

We do not build features whose purpose is to defeat a site's access controls, and we will not add them because a deal depends on it. If a target has said no, the answer is a different source.

This is why the compliance conversation with WebRobot is usually short. The robot is polite, the boundaries are stated, and the interesting question moves to where it belongs: is the data you want lawful to collect and to use? That one is answered by your counsel and your use case, not by a vendor page. The common cases (competitor prices, public catalogs, job listings, market research) rarely touch personal data at all, and the plans that run them are on pricing.

05 / PERSONAL DATA FIG. 5 · GDPR AND CCPA

Personal data, GDPR and CCPA: who carries what

General information, not legal advice. Your counsel decides what is right for your sources and your use case. What follows is how the roles usually fall.

If your robots collect personal data, you are the controller of it. You decide the purpose, you need a lawful basis, and you own the retention limit and the deletion requests that follow. WebRobot processes what you instruct your robots to collect and deletes it when you say so. That division is standard, and pretending otherwise would not help you at review time.

The practical advice most legal teams land on: scrape data that is not personal. Prices, catalogs, stock levels, listings, public company information, job postings without applicant detail. These carry the commercial value in most projects and almost none of the regulatory weight. When a project genuinely needs personal data, treat it like any other regulated dataset: minimize the fields, set a retention window, and be able to delete on request.

HOW THE ROLES USUALLY FALL

Deciding what to collectYou
Lawful basis and purposeYou
Running the extractionWebRobot
Storing it encryptedWebRobot
Retention windowYou set it
Deleting on requestYou ask, we delete

Region requirements? Ask before you buy and we will tell you plainly what we can do today.

06 / HONESTY FIG. 6 · WHAT WE DO NOT CLAIM

What this page is, and what it is not

Everything above is a description of practice and a set of policy commitments. It is not an audit report, and we are not going to dress it up as one. We do not currently hold a SOC 2 attestation or an ISO 27001 certification, and you will not find a badge on this page implying that we do.

We say this because security buyers can tell the difference, and because the vendors who blur it are the ones who cause problems later. If your procurement process requires a formal attestation before purchase, tell us during the conversation. If your policy requires a specific processing region, ask before you sign rather than after. You will get a straight answer either way, including a no.

Enterprise agreements can include a data processing agreement and a security review as part of the terms. Everything else on this page applies to every plan, from Launch at $79 per month upward. Security posture is not a tier.

READ IT LIKE THIS

Commitment
Something we promise and will hold to: no resale of your data, robots.txt honored, deletion on request.
Practice
How the system is built: TLS in transit, encryption at rest, least-privilege internal access.
Not claimed
Third-party attestations and certifications. We do not have them, so we do not imply them.
07 / COMPLIANCE FAQ FIG. 7 · REVIEW QUESTIONS

The questions security reviews actually send us

Copy any of these into your vendor questionnaire. The answers are the same ones we would give on a call.

Traffic between your browser, our application and your delivery targets runs over TLS. Data we store, including any credentials you give a robot, is encrypted at rest. This is a description of how the service is built, not a third-party audit finding.

Encrypted, inside your workspace, and used for exactly one purpose: signing your robots in to the target site you configured them for. They are not shared with other customers, not used for anything else, and you can delete them from the workspace at any time. Give the robot a dedicated account with the least privilege the job actually needs.

Access is limited to the staff who need it to operate the service or to answer a support request you opened. Internal access follows least privilege. That is a policy commitment and an engineering practice, and we would rather state it plainly than dress it up as a certification.

No. Rows your robots extract belong to you. We do not sell them, share them with other customers, or fold them into a dataset we license to anyone else. Your extracted data exists to be delivered to you and then deleted when you say so.

No. We do not currently hold a SOC 2 attestation or an ISO 27001 certification, and we would rather tell you that than imply otherwise. What we can describe is practice: TLS in transit, encryption at rest, least-privilege internal access, and retention you control. If your procurement process requires a formal attestation, tell us before you buy.

Ask us before you buy. If your policy requires processing or storage in a specific region, we will tell you plainly what we can and cannot do today rather than promise a residency guarantee we have not built. Region requirements are worth settling before a contract, not after.

That depends on you, not on the tool. If your robot collects personal data you are the controller: you need a lawful basis, a purpose, a retention limit, and a way to honor deletion requests. Most commercial scraping (prices, catalogs, listings, public company data) involves no personal data at all.

Yes. Robots honor robots.txt directives and apply polite pacing by default, so a crawl reads like an ordinary visitor rather than a flood. We do not build features whose purpose is to defeat a site access control, and we will not add them on request.

You control retention throughout. Extracted datasets and stored credentials can be deleted from the workspace at any time, including before you cancel, and anything you already exported to Sheets, Excel or your own systems stays yours. Ask us to remove everything after you leave and we will.

Enterprise agreements can include a data processing agreement and a security review as part of the terms. Raise it during the conversation rather than after signature, and bring the specific questions your legal and security teams need answered. Enterprise terms are listed on the WebRobot pricing page.

Product questions rather than security ones (cost, logins, why scrapers break, how the data is delivered) are answered on the web scraping FAQ. If you want to watch a robot work before the review starts, the interactive demo runs one in your browser, and the WebRobot AI web scraper home page shows the full picture.

FINAL ASSEMBLY

Bring it to your security review

Encrypted in transit and at rest, credentials used only to run your robots, retention you control, and no resale of your data.