Skip to main content
October 1, 2026 · dev · 9 min read

Fake Data for Prototyping — Picking the Right Generator

Which fake-data generator for which job. Mock JSON vs faker vs lorem vs UUIDs — stop seeding your demos with "John Doe" at Acme Corp.

Last updated October 1, 2026 · 9 min read

Every prototype ships with the same seven users: John Doe, Jane Doe, Bob Smith, Alice Smith, Charlie Brown, and two slots filled with lorem ipsum. The product lead screenshots the prototype, uploads it to Slack, and someone always asks the same question: "Can we use real names?"

The strong opinion worth defending: the fake data in your prototype is the prototype. A demo with real-shaped names, plausible job titles, and realistic activity timestamps lands differently than one with "User123" and "Lorem ipsum dolor sit amet." Spend an extra five minutes on data, save your week of "the UI feels off" feedback.

Here's the taxonomy of fake-data generators and when each one earns its place.

Category 1 — identity and user data

What you need: names, emails, avatars, job titles, locales. The thing populating most of your empty states.

Default to full user-profile generators, not just name generators. A real user has a first name, last name, email that matches the name, avatar initial that's consistent, job title, company, and timezone. If you're pulling five of those from five different generators, they won't be consistent and the demo will feel off.

For a complete profile-shaped output, the Mock User Profile Generator returns coherent user objects where the name, email, and avatar actually relate. Faster than wiring four generators together and reconciling by hand.

For when you need just the email addresses, formatted to a specific domain (e.g., all users at @acme.co), the Fake Company Email Generator does the one job without pulling in the rest of the profile machinery.

For international variety — users whose names span cultures without you manually curating — the Mock User Data Generator biases toward diverse name sets, which is a real feature if your product markets globally and your demos keep showing six white Anglo-Saxon names.

What not to do

Don't type the names into your seed script by hand. You'll end up with "Alice, Bob, Charlie, Dave, Eve" — the cryptography-textbook user list — and your product sense will drift.

Don't use faker.js output as-is for a demo video. The default faker names ("Jennyfer Durgan," "Stuart Hoeger") read as faker names to anyone who's seen a demo before. Shuffle, filter, or generate from a tool with less recognisable output distributions.

Category 2 — structured JSON / API payloads

What you need: realistic-shaped API responses for building against before the backend exists.

The failure mode is a JSON with the shape right but the values obviously wrong. Three nested users, each with the same timestamp, same email format, same firstName: "John" default. Your frontend works. Your demo convinces nobody.

For API response shapes, the Mock JSON Data Generator is the general-purpose tool — give it a schema sketch, get back realistic payload data. For specifically REST-shaped endpoints with pagination metadata, the Mock REST Endpoint Generator returns the full envelope (data, metadata, links) rather than just the inner records.

For GraphQL work where the shape is nested and the field types matter, the Mock GraphQL Response Generator outputs the right nested-object structure with plausible field values at every level. Saves the "oh I forgot to populate the deeply-nested author.badges field" discovery during the demo.

The "fake-but-believable" standard

A good mock payload passes three tests:

1. Values are not obviously synthetic. No "name": "User 1". 2. Values are internally consistent. The created_at is before the updated_at. The email matches the name. The avatar_url returns a 200 (or at least a plausible 404 for example.com). 3. Values are diverse. Not all country: "US". Not all price: 99.99.

Hand-rolled mocks usually fail test 3. Generators usually fail test 2. The fastest fix is to generate the base payload, then write a 10-line script that walks the output and enforces consistency (sort timestamps, derive emails from names, etc.).

Category 3 — placeholder text

Lorem ipsum, hipster ipsum, bacon ipsum, whatever. What you need: text of the right length and texture to see how your layout handles real content.

The strong opinion: lorem ipsum is wrong 80% of the time. Latin-looking text does not stress-test typography for actual words. English content has shorter average sentence lengths, more consistent lowercase letter frequencies, and more punctuation per paragraph than Latin. Your layout that looked balanced in lorem will look stuffy in English.

Use English placeholder text whenever the content will be English. Keep lorem ipsum for the rare case where you want to signal "this is placeholder" so the client doesn't ask "wait what is this product about?"

For English-register placeholders, the Business Placeholder Copy Generator is the right default for product-page prototypes. The Tech Placeholder Text Generator fits dev-tool prototypes better. The Random Product Description Generator handles e-commerce specifically.

If you want to signal "obviously placeholder, don't read this" without resorting to lorem, the Hipster Ipsum Generator or Pirate Ipsum Generator carry the "this isn't the real copy" message more clearly than Latin, while still being English-length-and-rhythm.

A pattern worth stealing

For dashboard prototypes, generate a mix:

  • 70% English placeholder text at realistic lengths.
  • 20% actual real-sounding content (five hand-written strings in your voice).
  • 10% deliberately long strings (worst-case content) to catch overflow bugs.

The mix catches layout problems the uniform "all 200-character strings" default misses.

Category 4 — IDs, tokens, secrets

What you need: UUIDs, API keys, JWTs, snowflake IDs. The things that make a mock payload look real.

For UUIDs the native crypto.randomUUID() is fine; a generator exists for the moments you need a batch and don't want to open a console. The UUID v4 Generator outputs batches cleanly.

For API-key-shaped strings that look authentic (correct prefix, right entropy, no "APIKEY123"), the Fake API Key Generator returns realistic-looking tokens — critical for screenshots where a visible "APIKEY123" undermines the demo.

For JWTs with realistic payload claims (iat, exp, sub, aud), the Mock JWT Token Generator returns valid-decoded JWTs you can paste into Postman and have them parse correctly.

The security note

Never ship fake API keys that match the real pattern of a specific provider (Stripe starts sk_live_, OpenAI starts sk-proj- or sk-, GitHub starts ghp_). Someone will see them in a screenshot and try to use them, or worse, you'll accidentally commit the "fake" key that is actually real. Use obvious fake prefixes: demo_ or mock_.

Category 5 — time-series and events

What you need: logs, metrics, event streams, time-bucketed data.

The lazy mistake: generate events with timestamps spaced evenly 15 minutes apart. Looks nothing like real event data, which clusters by business hours, has drop-offs on weekends, and spikes around deploys.

For syslog-shaped events, the Mock Syslog Line Generator outputs the right format and realistic distribution of levels. For log lines with request IDs and status codes that correlate, the Fake Log Entry Generator is tuned for request-log shape. For metrics in the Prometheus format, the Mock Prometheus Metric Generator gives you the right label structure and value plausibility.

For webhooks that look like real third-party deliveries (Stripe, GitHub, SendGrid shape), the Mock Webhook Payload Generator is useful for building webhook-receiver prototypes without actually hooking up to the service.

The pattern that fools people

Generate your event timestamps with:

  • A diurnal cycle (busy 9am–5pm local time).
  • A weekly cycle (weekends ~30% of weekday volume).
  • Random spikes (one or two days with 3x normal volume — simulates launches).

Even a 20-line Python script that imposes this shape makes mock event data feel like real event data. Uniform distributions always look wrong.

Category 6 — one-off field types

The miscellany. What you need: a batch of ISBNs, a bunch of credit card numbers (that pass Luhn), fake IP addresses, fake user agents.

Each has a tool. Pulling from them keeps your mock data from containing obvious tells.

The common theme: a generated field looks more realistic than a hand-rolled one, both because it's drawn from a plausible distribution and because you won't accidentally pick the same ten values every time.

The anti-pattern cheat sheet

Patterns to recognise and fix on sight in a prototype:

| Smell | What to do | |---|---| | All users have the same timestamp | Generate timestamps across a 30-day window with diurnal variation | | All users at @example.com | Mix 3–5 realistic domains | | 10 users, same gender/culture | Use a diverse-name generator or filter | | Status always "Active" | 10% "Pending," 5% "Inactive," 5% "Suspended" | | Prices all round to .99 | Mix round and non-round values | | Lorem ipsum anywhere | Replace with English placeholder, matched to context | | Tags: ["tag1", "tag2", "tag3"] | Real tag values from your domain | | Numbers all 1-1000 | Match the distribution of real data (long-tail, power-law, etc.) |

The one workflow

For a weekend prototype that will be demoed Monday:

1. Generate 50 user profiles. Save as users.json. 2. Generate 200 events across 30 days with diurnal cycle. Save as events.json. 3. Generate 20 realistic-length placeholder descriptions. Save as descriptions.json. 4. Hand-write five hero items (the ones that appear top-of-screen). Mix them into the generated data. 5. Load into local Supabase / local JSON API / local SQLite. Build the UI. 6. Right before demo: scan the UI for any "User 1," "Lorem," or empty-string artifacts. Fix the five that remain.

Step 4 is the one almost nobody does and the one that matters most. The top-of-screen items get looked at by the viewer first. Everything else is scanned. The hand-written hero items set the tone; the generated fill-in-the-blanks reinforce without disrupting.

A prototype with 50 generated users and 5 hand-written ones at the top reads as thoughtful. A prototype with 55 generated users reads as generated. The difference is 20 minutes of hand-crafting the hero and never-looked-at by anyone, except the viewer who absorbs the overall quality unconsciously.

When to graduate from fake data to recorded production shapes

Fake data is for before you have real users. Once you have 50+ real customers, your fake data starts to lag behind the actual distributions in your system. Prototypes built against stale fake data make product decisions based on shapes that don't match reality.

The upgrade: capture a snapshot of production, anonymise it, and use that as your prototype data instead of generation.

The anonymisation pattern that works:

1. Pull a full snapshot of a representative subset (e.g., last month's active users). 2. Replace every PII field (name, email, phone, address) with a generator output — Mock User Profile Generator for the primary identity, Fake IP Address Generator for IPs, Fake Browser Fingerprint Generator for anything device-related. 3. Keep all the numeric / temporal / behavioural fields untouched — these are what makes the shape real. 4. Verify no cross-field correlation leaks identity (if your anonymised data has a user who signed up on 2024-03-14, has 847 transactions, and lives in Reykjavik, there's exactly one person that is in production — anonymise the city or round the signup date).

The result: shape-accurate prototype data with no PII risk. Design decisions made against it reflect real distributions. The "fake data makes the UI look wrong because real data doesn't cluster this evenly" problem goes away.

The signal that you've waited too long to graduate: your designers keep asking "but what does this look like with real data?" during reviews. That question means your fake data has stopped reflecting the product.