Most generators hand you a grid of independent random values: an email that has nothing to do with the name beside it, a state code that contradicts the state, a credit card number your own validation rejects. That data is fine for filling space and useless for finding bugs. Here, each row is built with a small amount of context — the person's name drives the email, the username, the initials and the social handle; the company drives the work email, the domain and the website; the locale drives the city, the street format, the postcode, the phone number and the IBAN country. What comes out looks like a real export, so it breaks your code in the same places a real export would.
Seeding a fresh environment so it looks real. Mocking API responses before the backend exists. Filling a dashboard for a screenshot without exposing a single customer. Load-testing an import with fifty thousand rows. Proving that your CSV parser survives a comma inside a company name, that your table layout survives a Belgian address, and that your report survives a null. All of it without the legal and operational risk of copying production data into a system with weaker controls — the single most common finding in a data-protection review.
Generate up to 50,000 rows and export in ten formats. CSV is RFC 4180-escaped; CSV Excel adds a BOM and semicolons so European Excel opens it in columns with accents intact; SQL can emit a matching CREATE TABLE with inferred column types, UNIQUE and NOT NULL constraints, batched 500 rows per statement; NDJSON feeds log pipelines and bulk loaders; Markdown and HTML are for documentation and mock-ups. Switching format re-renders the same rows rather than drawing new ones, so you can export one dataset five ways.
Almost nobody wants “fake data” in the abstract. They want sample insurance claims, test patient records, dummy invoice lines, example meter readings — a file that looks like the one their system actually ingests. The sector library holds 18 ready-made industry schemas on top of the twelve generic templates, each with the columns, code lists and status mixes that domain really uses: a claims file with reserves and an excess column, a CDR file with cell IDs and call duration, a lab file with analytes, units and a reference range. Pick one, change what you need, generate.
Each link loads that schema and generates it immediately, so you can bookmark or share the exact starting point. Add &rows=1000 or &fmt=json to the URL to fix the row count and the export format too — handy in documentation, a ticket or a README.
One table of random rows is easy. What breaks most test fixtures is the second table: an orders file whose customer_id points at customers that do not exist, so every JOIN returns nothing and every report looks empty. The Related tables tab generates a small relational data set — parents first, then children whose foreign keys are drawn only from parent keys that were actually created. Four sets are built in:
The distribution is deliberately skewed rather than uniform: a few customers have many orders, some have none. That is what production looks like, and it is what exposes the report that silently drops customers with no rows and the pagination that assumes an even spread. Export the whole set as one SQL file with CREATE TABLE, PRIMARY KEY, REFERENCES and batched inserts — it loads into an empty Postgres, MySQL or SQLite database as-is — or as a single JSON object with one array per table, or one CSV per table.
The other half of this problem is the file you are not allowed to share. You have a real export and you need it in a ticket, a bug report, a demo or a shared staging database. The Anonymise a CSV tab reads the file in your browser, works out what each column holds, and replaces every identifying value with a fake one of the same shape:
Nothing is uploaded: the file never leaves the tab, which is the only way a masking step is also a compliance step. It is a pragmatic data-masking and pseudonymisation tool for test and development environments, not a formal anonymisation certificate — a file with one Belgian rheumatologist in it can still be re-identifiable from the columns you kept, so drop what you do not need.
You rarely need only data. You need the table to put it in, the model that reads it, the type that validates it. The Code & schema tab turns the schema you just built into working code in thirteen shapes: PostgreSQL and MySQL DDL (with UNIQUE, NOT NULL, CHECK or ENUM from your value lists), Prisma, TypeScript, Zod, Mongoose, Pydantic, SQLAlchemy, a Go struct with json and db tags, JSON Schema 2020-12, an OpenAPI 3.1 component, and a ready-to-run generator script for Faker.js and Python Faker — so you can move the same schema into CI, where a browser tab cannot go.
Generated data is only useful if it is what you think it is. The Profile & validate tab reads the dataset in the output panel and reports, per column: how many distinct values it holds, the real null rate, the numeric range or the string-length range, the three most common values with their share, and — where the type has one — a checksum verdict. Cards are checked with Luhn, IBANs with mod-97, EAN-13 and ISBN-13 with the GS1 weighted sum, UUIDs against the v4 layout, private IPs against RFC 1918. It also counts duplicate rows and tells you which columns are unique across every row, so you know what can serve as a key before you import.
This is the part that answers “is this data actually valid?” without a round-trip through your own validator, and it is how you catch a null rate you set to 5% that landed at 40% because a column was drawn from a short list.
Real test data rarely sits inside one theme. A few combinations that come up constantly, and how to build them here:
Mockaroo is the tool most people compare this to. Its free tier is capped at 1,000 rows per download and its multi-table relational export sits behind a paid plan; it is a hosted service, so your schema and the generated file go through a server. This page is client-side: no account, no row-count paywall and no upload, with relational sets, schema/code export, CSV anonymisation and profiling included. What a hosted tool still does better is the part a browser tab cannot reach: a REST API you can call from CI, regex-driven custom types, and datasets far beyond what a tab can hold in memory.
Free-tier limits and paid features change; check the vendor's current page before you rely on a row in this table. The honest summary: for a file you need now, in a browser, a page like this one wins on limits and privacy. For generating data inside a test suite that runs on every commit, a library wins — which is why the Code & schema tab hands you the Faker.js and Python Faker version of whatever you built here.
Fake data, test data, anonymisation and synthetic data, explained
Building a schema, starting from an industry instead of a blank page, generating tables whose foreign keys join, masking a CSV you already have, exporting the schema as SQL, Zod or Pydantic, verifying checksums, and why generated data is the safe answer to copying production.
Getting started
What is a fake data generator and how does this one work?
A fake data generator creates realistic but entirely fictional records — names, emails, addresses, IBANs, IP addresses, timestamps, UUIDs — so you can build and test software without touching real customer information. Here you build a schema (a list of columns, each with a type), choose how many rows you want and an export format, then press Generate. Everything runs in your browser; nothing is uploaded.
How do I get started quickly?
Click a template — Person, Address, Company, Tech / API, E-commerce, Full Record, Healthcare, Event Log, Social Media, Network / FW, Transactions or IoT Sensors. Each loads a complete, sensibly configured schema and generates immediately. Change a type, adjust the row count, pick a format and you are done.
How many rows can I generate?
Up to 50,000 rows per run. The slider covers the common range up to 5,000 and the number box next to it accepts any value up to the maximum. Above roughly 20,000 rows the page shows a note, because the browser will pause for a moment while it builds the file.
Do I need an account? Is there a paywall?
No account, no email, no paywall and no row limit behind a subscription. That is the main practical difference from hosted generators: because everything happens in your browser, there is no server cost to recover and nothing to sign up for.
Does it work offline?
Oncethe page has loaded. All the data pools and generators are part of the page itself, so you can disconnect and keep generating. It is not installed as an offline app, so the first load needs a connection.
Which keyboard shortcuts exist?
Ctrl / ⌘ + Enter generates, and Ctrl / ⌘ + S downloads the current output as a file instead of opening the browser's save-page dialog.
How do I build a schema by hand?
Press Add Field, type a column name, and choose a type from the grouped dropdown. Use ▲▼ to reorder columns, ∅% to make a share of the values null, U to force unique values, and ✕ to delete a column. Numeric and date types get inline min/max boxes.
Can I save a schema and use it again?
Three ways. 💾 Save stores it in this browser; 📂 Load brings it back. ⎘ Copy JSON puts the schema on your clipboard so you can paste it into a repo or a ticket, and ⇤ Import JSON reads it back. 🔗 Copy share link packs the schema, seed, row count, locale and format into a single URL.
What is the Table view for?
⊞ Table renders the first 200 rows as a real table with sticky headers, so you can scan the data and spot nulls, wrong ranges or a column that is not varied enough. </> Raw shows the exact text you will copy or download.
Field types and realistic values
How many field types are there?
123 types in 11 groups: Identity, Contact, Location, Business, Finance, Commerce, Network, Date & Time, Technical, Text and Special. The exact number is shown next to the Add Field button, since it changes as types are added.
Are the generated credit card numbers valid?
They pass the Luhn check and use the correct issuer prefixes and lengths — Visa 16 digits starting with 4, Mastercard 16 starting 51–55, American Express 15 starting 34 or 37, Discover 16 starting 6011. That means your validation code will accept them, which is exactly what you need for testing. They are not linked to any real account and cannot be charged.
Are the IBANs real?
They are structurally valid: the correct country length, and a mod-97 check digit calculated per ISO 13616, so an IBAN validator will accept them. The bank and account portions are random, so they do not correspond to any real account.
What about EAN-13 and ISBN-13?
Both carry a correct GS1 modulo-10 check digit, and ISBN-13 uses a real 978 or 979 prefix. Barcode software and ISBN validators will accept them; they will not match a real product or book.
Do the values inside one row belong together?
This is the part most generators get wrong. Within a row, Email, Work Email, Username, Initials and Social Handle are derived from the First Name and Last Name in the same row, Work Email and Domain and URL come from the Company in the same row, and State and State Code always match. Pick the Person template and look at the table: emma.peeters@… sits next to Emma Peeters, not next to a stranger.
Are the email addresses safe to send to?
They are invented and use ordinary public mail domains, so an address could in theory belong to someone. Never point a test mail server at generated addresses — use a catch-all mailbox, a tool like MailHog, or replace the domain with example.com, which RFC 2606 reserves for exactly this purpose.
What is the difference between Email and Work Email?
Email uses consumer domains such as gmail.com or proton.me. Work Email builds first.last@company-domain from the Company field in the same row, so a table of employees looks like a real staff directory.
Which types are useful for logs and observability data?
Log Level, Severity, ISO 8601, Duration ms, HTTP Status, HTTP Method, URL Path, Referrer, IP Address, Private IP, CIDR Block, ASN, User Agent, Browser, OS, Device Type, K8s Pod, Docker Image, Cloud Region and File Path. The Event Log and Network / FW templates wire several of these together.
Which types produce identifiers?
UUID (v4), Nano ID, Mongo ObjectId, Git SHA and Short SHA, MD5, SHA-256, API Key, JWT, Base64, Transaction ID, Employee ID, SKU and Number Sequence. Turn on U for any of them if the column must be a primary key.
What is the Weighted List type?
It draws from a list where you control the odds. Write active:70,inactive:20,banned:10 and roughly 70% of rows will be active. Real data is almost never uniformly distributed, and a status column that is one-third banned users will not exercise your UI the way production does.
What is the difference between Custom List and Constant?
Custom List picks one value at random from your comma-separated list. Constant writes the same value into every row — useful for a tenant id, an environment tag or a version marker that has to be identical across the whole import.
What does Foreign Key do?
It is an integer within a range you set, meant to reference another table. Generate a parents table with a Number Sequence id from 1 to 500, then give the child table a Foreign Key column with min 1 and max 500 — the references will resolve.
Can I generate nested JSON?
Only in a limited way: the JSON Object type puts a small object inside a single column. Full nested documents with arrays and sub-objects are not supported — every row here is a flat record, which is what CSV, SQL and spreadsheets need anyway.
Why is there a Diagnosis field?
For healthcare and insurance test data, where you need a plausible clinical column to build screens and reports around. The values are common condition names attached to entirely fictional patients — nothing in the output relates to a real person.
Can I add my own field type?
Not from the interface. The page is a single self-contained HTML file, so if you save it you can add an entry to the GEN object and to TYPE_GROUPS — a type is just a name and a function that returns a value. For one-off needs, Custom List or Weighted List usually covers it without touching code.
Locales, seeds and reproducibility
What does the locale setting change?
It switches the whole row to one country's conventions: first and last names, cities, street naming and address order, postcode format, phone and mobile number format, country and country code, and the IBAN country. Nine options are available — Global mix, United States, United Kingdom, Belgium, Netherlands, Germany, France, Spain and the Nordics.
What does Global mix do exactly?
It picks a locale per row rather than per dataset, so one row is Belgian, the next German, the next American — but each row stays internally consistent. That is the closest thing to an international customer base, and it is the setting that finds encoding and formatting bugs fastest.
Why does locale matter for testing?
Because most real bugs live in the assumptions. A five-digit postcode field breaks on 1000 in Belgium and on SW1A 1AA in London. A name column that assumes ASCII breaks on Müller, Léa and Álvarez. A phone validator built for (555) 123-4567 rejects +32 470 12 34 56. Generating with a non-US locale surfaces all of that before a customer does.
What is the seed and why should I set one?
The generator is driven by a seeded pseudo-random number generator, so the seed decides every value. Type order-fixtures as your seed and the same schema will produce byte-identical data every time, on any machine. Leave the field empty and a new random seed is used per run.
How do I keep a dataset I just generated?
Press 📌 next to the seed field. The seed that produced what you are looking at is copied into the box, so the next Generate reproduces it exactly. 🎲 does the opposite: a new random seed and a fresh dataset.
Why does reproducible test data matter?
Because a fixture that changes on every run makes tests lie. A snapshot test, a visual-regression screenshot or a report with a fixed expected total all break when the data underneath them shifts. A fixed seed turns generated data into something you can commit and diff — the only changes you see are the ones you made.
Does the same seed give the same data if I change the schema?
And that is intentional. The values are drawn in schema order, so adding, removing or reordering a column changes what every later column receives. Keep the schema and the seed together — the share link stores both, which is the safest way to hand a dataset to someone else.
What exactly is in the share link?
The full schema (names, types and every per-field option), the seed, the row count, the locale, the export format and the table name — encoded into the part of the URL after the #. Browsers never send that fragment to a server, so the link is self-contained and private. Open it and the page rebuilds the dataset immediately.
Can I put the share link in a repository?
That is a good use for it. Drop it in a comment above your fixture file or in the README of a test suite, and anyone can regenerate or extend exactly the same data instead of guessing how the file was made.
Which random number generator is used?
A small deterministic generator (mulberry32) seeded with an FNV-1a hash of your seed string. It is fast and identical across browsers, which is the point. It is not cryptographically secure: the Password and API Key types are fine as placeholders but must never be used as real secrets.
Nulls, uniqueness and realistic edge cases
What does the ∅% column do?
It sets the probability, per field, that a value comes out null. Set 20 and roughly one row in five will be empty in that column. In CSV and TSV that is an empty cell, in JSON and YAML a real null, in SQL NULL, and in XML a self-closing element.
Why would I deliberately generate nulls?
Because missing data is where software falls over. A dashboard that averages a column, a template that prints user.middle_name.toUpperCase(), a CSV import that assumes every row is complete — all of them work perfectly until the first empty cell. Ten or twenty per cent nulls in optional columns is a realistic and revealing setting.
What does the U button do?
It forces unique values in that column. The generator retries a value that has already appeared, up to fifty times per row. Use it for primary keys, emails, usernames and any column with a unique index — otherwise a 10,000-row import will hit a duplicate-key error somewhere in the middle.
What if uniqueness cannot be satisfied?
If a type has fewer distinct possible values than the number of rows you asked for — 10,000 rows of a five-item Custom List, say — the generator appends a numeric suffix so the column stays unique, and tells you how many values needed one. That is a signal to switch to a type with a bigger space, such as Number Sequence, UUID or Nano ID.
Does U guarantee a unique primary key?
For Number Sequence, UUID, Nano ID and Mongo ObjectId, effectively yes. For short or list-based types it is best-effort with the suffix fallback described above. If the column is a real primary key, use Number Sequence — it is unique by construction and it sorts.
How do I generate deliberately awkward data?
Combine the tools: a Custom List containing an apostrophe, a comma, a quote and an emoji; Paragraph in a column your UI expects to be short; ∅% at 50 to flood the null path; the Nordics or Belgium locale for non-ASCII names; and Amount with a negative minimum to test refunds and credits. That mix finds escaping, layout and validation bugs in one pass.
Can I control decimal places?
Float has a decimals option in its defaults (two by default), and Price and Amount always produce two. For currency, two is what you want; storing money as a float at all is a separate argument your database will eventually win.
Can values be negative?
Set a negative minimum on Integer, Float or Amount. Amount defaults to a range that includes negatives precisely so a transactions table contains refunds as well as payments.
Do date ranges handle leap years and month lengths correctly?
Dates are drawn as real timestamps inside the year range you set, so you get 29, 30 and 31-day months and a genuine 29 February in leap years. That matters: the classic date bug is code that only ever saw days 1–28 in its test data.
What timezone are the dates in?
UTC. Date, DateTime, ISO 8601 and both Timestamp types are all derived from the same UTC instant, so a row's date and timestamp columns agree with each other. The ISO 8601 type is the one to use when your importer expects a timezone marker.
Export formats and getting data into your tools
Which export formats are available?
Ten: CSV, CSV Excel, TSV, JSON, NDJSON, SQL, YAML, XML, Markdown and HTML. Switching format re-renders the data you already generated — it does not draw new values.
What is the difference between CSV and CSV Excel?
Plain CSV is comma-separated UTF-8, which is what most importers and pandas expect. CSV Excel is semicolon-separated and starts with a UTF-8 byte-order mark, which is what Excel needs in most European locales to split the columns properly and show é, ü and ø correctly instead of mojibake. If your spreadsheet dumps everything into column A, this is the option you want.
Is the CSV properly escaped?
Any value containing the separator, a double quote, a line break or leading or trailing spaces is wrapped in quotes, and internal quotes are doubled — RFC 4180 behaviour. That matters as soon as your data contains a Paragraph field or a company name with a comma in it.
When should I use NDJSON instead of JSON?
JSON gives one array you can drop into a fixture or a mock API. NDJSON gives one object per line, which is what log pipelines, BigQuery, Elasticsearch bulk loaders and most streaming tools ingest — and it lets you read a huge file line by line without holding all of it in memory.
Can it write the CREATE TABLE statement too?
Choose SQL and tick Include CREATE TABLE. Column types are inferred from the field types — BIGINT for counters and timestamps, DECIMAL(12,4) for money and coordinates, DATE and TIMESTAMP for dates, BOOLEAN, TEXT for long prose and VARCHAR(255) otherwise. Columns marked unique get a UNIQUE constraint, and columns with no nulls get NOT NULL.
Which SQL dialect is the output?
The inserts use backtick-quoted identifiers, which is MySQL and MariaDB syntax. For PostgreSQL or SQLite, replace the backticks with double quotes or delete them — a single find-and-replace. Values are escaped with doubled single quotes, which is standard everywhere.
Why is the SQL split into several INSERT statements?
Rows are batched 500 at a time. One statement with 50,000 value tuples exceeds the maximum packet size on a default MySQL install and is painful to debug when one row fails. Batches of 500 import fast and fail readably.
How do I import the result into Excel or Google Sheets?
For Excel, use CSV Excel and open the file directly. For Google Sheets, plain CSV works better — File → Import → Upload. If you only need a few rows, copy the Markdown or plain CSV output and paste it in.
How do I load the data into pandas or R?
Download plain CSV and read it with pandas.read_csv('fake-data-abc123.csv') or readr::read_csv(). For a JSON workflow use NDJSON with pandas.read_json(path, lines=True). The filename includes the seed, so you can always trace a file back to the exact configuration that produced it.
What is the Markdown format for?
A ready-made table you can paste into a README, a pull request, a ticket or documentation — useful when you want to show the shape of a dataset rather than ship it. Pipes inside values are escaped so the table does not break.
What does the HTML export give me?
A clean <table> with a thead and tbody and everything HTML-escaped, ready to paste into a page, a component or a design mock-up that needs a realistic-looking data table.
Can I use it to mock an API?
Generate JSON and serve it from a static file, or paste it into a mocking tool such as Mockoon, WireMock, MSW or json-server. Set a seed so the mock is stable, and use Work Email, UUID and ISO 8601 to make the payload look like something a real backend would return.
Why is the preview cut off on very large exports?
Only the on-screen preview is truncated, at around 300,000 characters, because rendering a multi-megabyte string in the browser is slow and pointless. Copy and Download always give you the complete dataset, and the row counter in the toolbar shows the real total.
Privacy, GDPR and using fake data safely
Is anything I generate sent to a server?
There is no back end. The schema, the seed and every generated row exist only in your browser tab, and the share link keeps its payload in the URL fragment, which browsers never transmit. The site uses privacy-first analytics that count page views without cookies or personal data.
Is the generated data GDPR safe?
The data is invented, so it is not personal data and the GDPR does not apply to it — which is exactly why generated data is the right answer to “can we copy production into staging?”. Two caveats worth stating plainly: a random name can coincide with a real person's, and an email address on a real domain could exist. Treat the output as fictional, never as anonymised production data.
Can I use this instead of copying the production database?
That is the best reason to use it. Copying production into a test environment spreads personal data into systems with weaker access control, longer retention and more people looking at it — one of the most common findings in a data-protection audit. A generated dataset with the same shape carries none of that risk.
Is generated data the same as anonymised or pseudonymised data?
And the difference is legally important. Pseudonymised data is still personal data — it is derived from real records and can be re-identified. Synthetic data is invented from scratch and has no data subject behind it. Only the second one takes you out of scope.
Can I use the output commercially?
There is no licence on the output, no attribution requirement and no watermark. Use it in products, demos, courses, screenshots and test suites.
Are the company names real?
They are a mix of invented names and well-known fictional companies — Acme, Initech, Contoso, Fabrikam, Umbrella. They are placeholders. If you are producing screenshots for publication, prefer your own invented names to avoid any trademark question.
Can I use generated data for a public demo?
And it is the safest option. Set a seed so the demo shows the same records every time you present it, avoid anything that looks like a real person's contact details, and use the example.com domain if the screenshots will be published.
Is it safe to use the generated passwords and API keys?
Only as placeholders. They come from a seeded, non-cryptographic generator, so anyone who knows the seed can reproduce them exactly. Never use them as real credentials — for that, use your operating system's or language's secure random source.
Does using fake data help with HIPAA or medical testing?
It removes the protected health information problem, which is most of it: no real patient record ever enters your test environment. It does not make a system compliant by itself — access control, audit logging and retention are still yours to get right.
Limits, troubleshooting and alternatives
Is this a free Mockaroo alternative?
It covers the same core ground — custom schemas, well over a hundred field types, per-field null rates and uniqueness, ten export formats and a live table preview — with no account, no row-count paywall and no upload. What hosted tools still do better is very large datasets, related multi-table exports, regex-driven custom types and API access.
What can this tool not do?
- Multiple related tables in one export (generate them separately and join on a Foreign Key range)
- Deeply nested JSON documents
- Custom values driven by a regular expression
- Server-side generation or an API you can call from CI
- Datasets larger than 50,000 rows in a single run
For millions of rows, a library in your own language — Faker, Bogus, factory_boy — belongs in the pipeline instead.
The page freezes for a moment on large exports.
Generation happens on the browser's main thread, so a 50,000-row export with twenty columns will block the tab for a second or two while it builds the string. That is expected. If it happens often, generate in a few smaller runs or reduce the number of long text columns, which dominate the output size.
Copy does not work.
The clipboard API needs a secure context and a user gesture; if it is blocked, the tool falls back to an older copy method automatically. For very large outputs, Download is more reliable than Copy — some browsers refuse to put multi-megabyte strings on the clipboard.
The download does nothing.
The file is built in memory and handed to the browser as a blob, which some strict privacy extensions and managed corporate profiles block. Copying the output and pasting it into a new file gives the identical result. On iOS the file may open in a viewer instead of saving — use the share sheet from there.
My import failed on a duplicate key.
Turn on U for that column. Without it, values are drawn independently and a collision in a 5,000-row set is not just possible but likely — the birthday problem makes duplicates far more common than intuition suggests.
Two columns have the same name — what happens?
The schema panel flags them in red, and on export they are automatically numbered (email, email_2) so the file stays valid. Empty column names are replaced with field_1, field_2 and so on. It is worth fixing them properly before you export.
Why did all my values change when I only added a column?
Values are drawn in schema order from a single seeded stream, so inserting a column shifts everything after it. If you need a dataset to stay stable, freeze the schema first and only then pin the seed — or keep the share link, which stores both together.
Does it work on a phone?
Below roughly 1000 pixels the schema panel moves above the output instead of beside it. Building a wide schema is more comfortable on a desktop, but generating, previewing and downloading all work on mobile.
Which browsers are supported?
Any current version of Chrome, Edge, Firefox, Safari, Brave, Opera or Vivaldi, on desktop and mobile. There are no external libraries and no framework — the page is self-contained, which is also why it keeps working offline.
Is the tool free, and will it stay free?
It is free with no account and no paid tier. It is one of a set of client-side browser tools at jasperbernaers.com/apps — no server means no running cost, no sign-up and nothing to leak.
Sector data sets and cross-industry schemas
Which industries have a ready-made schema?
Eighteen: insurance claims, shipments and tracking, HR and payroll, students and enrolment, property listings, SaaS subscriptions, support tickets, telecom call and data records, energy meter readings, clinical and lab results, retail POS lines, stock and warehouse, CRM leads and deals, bank transactions, vehicles and fleet, travel bookings, permits and public-sector cases, and campaign and web analytics. They sit alongside the twelve generic templates (person, address, company, tech/API, e-commerce, full record, healthcare, event log, social media, network, transactions, IoT).
What makes a sector schema different from just picking field types?
The code lists and the mix. An insurance schema is not "a name and a number" — it is a cause-of-loss column drawn from the causes insurers actually record, a status column weighted the way a claims book really sits (most claims approved or paid, a few in fraud review), a reserve that is in the same order of magnitude as the claimed amount, and an excess column that is often zero. That is what makes the file recognisable to someone who works with the real thing.
Can I generate sample data for two sectors at once?
Load a sector, then add columns from any other category. A fraud fixture is bank transactions plus IP address, device type and country code. A cold-chain file is shipments plus a temperature Float. Nothing stops you combining them, and the cross-sector recipes above list the combinations that come up most.
Is there a URL I can bookmark for a specific sector?
?sector=insurance, ?sector=telecom, ?sector=clinical and so on load that schema and generate immediately. You can add &rows=1000, &fmt=json, &locale=be and &seed=abc123 to pin the row count, format, locale and seed, which makes the link a complete, reproducible instruction you can paste into a ticket.
Can I get a sample CSV file to download without building anything?
That is what the sector links are for: open one and press Download. Because the seed is in the URL when you share it, everyone who opens that link gets the same file, which is what you want in documentation or a test fixture.
The schema is close but not exactly my system’s columns.
Rename any column in place — the name is just a label, so claim_id becomes SinisterNr without changing the generator behind it. Add, remove and reorder columns freely, then use Copy JSON to keep your version, or the share link to keep schema, seed, locale and format together.
Related tables and relational test data
How do I generate test data for two tables that have to join?
Open the Related tables tab and pick a set. The parent table is generated first, its keys are collected, and every child row draws its foreign key from that list — so every join finds a match and no insert fails on a missing reference. A scale factor multiplies all three tables at once.
Why are the children unevenly spread across parents?
Because real data is. The foreign key is drawn from a skewed distribution, so a few parents get many children and some get none. Uniform fixtures hide a whole class of bug: the report that silently drops customers with no orders, the average that assumes every parent has at least one child, the pagination that breaks on a long tail.
Can I load the result straight into a database?
The SQL export writes a CREATE TABLE per table with a PRIMARY KEY, a REFERENCES clause on every foreign key, and the inserts in parent-before-child order. It loads into an empty PostgreSQL, MySQL or SQLite database as-is. Column types are inferred from the first row, so check them before you use the file as a migration.
Can I define my own multi-table set?
Not in the UI yet — the four built-in sets (shop, clinic, HR, fleet) cover the common shapes. What you can do is build each table separately with a Foreign Key column whose min/max matches the parent’s key range: the keys will be in range, though the distribution will be flat rather than skewed.
Is the relational set reproducible too?
Each table is seeded from the master seed plus its own name, so the same seed and the same scale factor rebuild the whole set byte-for-byte, including which child belongs to which parent.
How large can a relational set get?
The scale factor goes up to 20×, which on the shop set is roughly 800 customers, 2,400 orders and 6,000 items. For bigger volumes generate each table on its own in the main panel — it handles 50,000 rows per table — and keep the foreign key ranges aligned.
Anonymising and masking a CSV you already have
Can I anonymise a real CSV instead of generating a new one?
Paste it into the Anonymise a CSV tab. Columns are classified from both the header and the values, identifying values are replaced with fake ones of the same shape, and everything else passes through unchanged. The file is read and rewritten inside the page — it is never uploaded.
Will the relationships in my file survive?
That is the point. The mapping is consistent: one original value always becomes the same replacement, so a customer who appears in forty rows is still one customer, duplicate detection still finds the same duplicates, and a column you join on still joins. Row count and column order are untouched.
Do the replaced values stay valid?
Where a checksum exists, yes. A masked card number is repaired to pass Luhn, a masked IBAN keeps its country code and gets a recomputed mod-97 check digit, and phone numbers keep their length and separators. Your validation and your import will accept the anonymised file wherever they accepted the original.
What about dates and amounts?
Dates can be shifted by a single constant offset, which blurs birth dates while preserving every interval — the days between an order and its refund stay exactly the same. Amounts are left alone by default, because numbers are rarely identifying on their own; if you need them blurred, set a jitter percentage and each value is nudged within that band.
What does it detect automatically?
E-mail addresses, names (first, last, full), usernames, phone numbers, streets, cities, postcodes, countries, companies, job titles, IBANs, card numbers, VAT numbers, national ID numbers, licence plates, IP addresses, URLs, UUIDs and dates. Detection combines a value-pattern test with header keywords in English, Dutch, French, German and Spanish, and every decision is shown in a table you can override per column.
It classified a column wrongly.
Change it in the dropdown next to that column and the file is rewritten instantly. You can also set a column to keep as-is, drop column, or pseudonym (stable hash) — the last replaces each distinct value with a short stable token, which is what you want for an internal ID you must keep joinable but not readable.
Is this GDPR anonymisation?
It is pseudonymisation and masking, which is the right tool for building test and development data. Under the GDPR, truly anonymised data is data that can no longer be attributed to a person by any reasonable means — and a file can still be re-identifiable from the columns you kept, even with every name replaced, if the combination is rare enough. Drop columns you do not need, generalise the ones you do (a year instead of a date of birth, a region instead of a postcode), and treat the result as safer, not anonymous. For a decision with legal consequences, ask the person who owns that risk in your organisation.
How big a file can I paste?
It reads the first 5,000 data rows and tells you when it has truncated. That is a browser-textarea limit, not an algorithm limit — for larger files, mask a representative sample here to agree the column mapping, then implement the same mapping in your pipeline.
Which delimiters does it handle?
Comma, semicolon, tab and pipe, detected from the first line, with RFC 4180 quoting — so quoted fields containing the delimiter, escaped double quotes and multi-line values survive the round trip. The output uses the same delimiter it found.
Can I use it on a file with no header row?
Untick first row is a header and columns are named col_1, col_2 and so on. Detection then relies on the values alone, which works well for e-mails, IBANs, cards, IPs, UUIDs and dates, and less well for names and cities. Set those by hand.
Exporting the schema as code
Can I get a CREATE TABLE statement for my schema?
Two, in fact: PostgreSQL and MySQL, in the Code & schema tab. Nullable columns follow the ∅% you set, unique columns get UNIQUE, weighted and custom lists become a CHECK … IN (…) on Postgres or an ENUM on MySQL, and the first unique column becomes the primary key. The SQL export in the main panel also offers a CREATE TABLE alongside the inserts.
Which languages and frameworks are covered?
Thirteen targets: PostgreSQL, MySQL, Prisma, TypeScript interfaces, Zod, Mongoose, Pydantic v2, SQLAlchemy, a Go struct with json and db tags, JSON Schema 2020-12, an OpenAPI 3.1 component, plus runnable generator scripts for Faker.js and Python Faker.
Why would I want the Faker.js or Python Faker version?
Because a browser tab cannot run in CI. Build and refine the schema here where you can see the data, then take the generated script into your test suite so every commit generates the same shape of data. The script is seeded from the seed you are using, so the two stay comparable.
Are the generated types safe to use as a migration?
Treat them as a first draft. Lengths are generic (VARCHAR(255)), the primary key is guessed from your unique columns, and no indexes, defaults or foreign keys are inferred. It saves the typing; it does not replace the review.
Can I change the table or model name?
The name box above the code output feeds the table name, the class name and the file name of the download. Column names are converted to snake_case for SQL and Python, camelCase for Prisma fields, and PascalCase for Go struct fields, with the original kept in a mapping or tag where the target supports it.
Profiling and validating generated data
How do I check that what I generated is valid?
Open the Profile & validate tab after generating. Every column is listed with its distinct count, its real null rate, its numeric or length range, its three most common values, and — where the type carries a checksum — a pass count. Cards are run through Luhn, IBANs through mod-97, EAN-13 and ISBN-13 through the GS1 weighted sum.
Why does the null rate not match the ∅% I set?
Because ∅% is a probability per row, not a quota. At 5% over 100 rows, anything from 1 to 11 nulls is unremarkable; the profile shows what you actually got. Over 10,000 rows it converges on the number you set.
How do I know a column is safe to use as a primary key?
The profile marks any column whose values are distinct across every row and never null as usable as a key, and the header counts them. If a column you meant to be unique is not listed, turn on U for it and regenerate.
What does the duplicate row count tell me?
That two rows are identical in every column. With a unique ID column it should always be zero; without one, a narrow schema of low-cardinality columns will legitimately repeat, which is useful to know before you import into a table with a composite unique constraint.
Can it profile a file I did not generate here?
Not directly — the profiler reads the dataset in the output panel. For an external file, paste it into the anonymiser, which reports its column kinds, delimiter, row count and how many distinct values each identifying column holds.
Using it with your stack
How do I seed a local PostgreSQL or MySQL database?
Generate, choose the SQL format, tick Include CREATE TABLE, download, and pipe the file in — psql -f fake-data-xxxx.sql mydb or mysql mydb < fake-data-xxxx.sql. For more than one table use the Related tables tab, whose SQL export carries the foreign keys as well.
How do I load it into MongoDB or Elasticsearch?
Use NDJSON: one object per line is exactly what mongoimport and the Elasticsearch bulk API expect. The Mongoose model in the Code & schema tab gives you the matching document shape.
Can I use it as a fixture for Jest, pytest or Playwright?
Generate with a fixed seed, export JSON, and commit the file — the same seed and schema reproduce it byte-for-byte, so a snapshot test stays stable. If you would rather generate at test time, take the Faker.js or Python Faker script instead.
How do I fill a dashboard or a design mock-up?
Pick the sector that matches what the screen shows, set the row count to whatever the view needs, generate, and use the table preview directly or export CSV into your charting tool. It is the one way to produce a realistic screenshot with no customer data in it.
How do I test an importer properly?
Deliberately make the file awkward: raise ∅% on two columns, force a company name with a comma and a quote in it, include non-ASCII surnames, set an Amount range that crosses zero so negatives appear, and use the Excel CSV format to check your parser copes with a BOM and semicolons. A clean file proves very little.
Can I generate data for a load test?
Up to 50,000 rows per run, and the seed makes the runs comparable: change the seed for a different dataset of the same shape, keep it to repeat the same one. For millions of rows, take the Python Faker script and loop it — that is what a library is for.