我创建了一个真实的虚假数据源。

1 分•作者: przeslijmi•大约 2 个月前•原帖
大家好。 我想向大家介绍一个我在过去几个月里一直在开发的工具——真实虚假数据生成器。 网址是:https://real-fake-data.com 这个工具的基本需求是提供测试数据,这些数据(1)与真实数据相同但并不真实,(2)确保您的数据管道能够捕捉到现实生活中的问题,而不仅仅是开发者在开发时会记得覆盖的问题。 首先,现实是不断变化的。您可能会有一个提交表单,要求填写车辆注册号码,但您会遇到问题(不同国家,不同规则)。您要么花一个月的时间挖掘各种场景,要么就随便设置一个“最大长度7”的限制,然后离开。您的工具已经上线并正常工作——但是有一天,某个国家或州决定将车辆牌照的长度设置为8……您的系统就会失败(用户无法输入新的车辆牌照),而您的测试管道仍在正常工作,并给您发送虚假的绿色微笑。 其次,您真的无法预测所有情况。您面向全球用户?很好,但他们使用不同的字母表。您美丽的用户体验/用户界面可能会因为用户的姓名缩写变成4或5个字母而崩溃。或者姓氏字段可能太短。或者它是否会接受即使在欧洲也使用的非拉丁字符? 当您创建新软件时,您还需要一些种子数据——以便在开发时能够看到您的工具与一些模拟用户、模拟帖子、模拟产品、模拟交易等的表现。同样,您要么花费大量时间来编写这些数据——即使有AI的帮助——要么您就只能留下user1、user2、user3,失去了从真实用户的角度来看待您的工具的能力。 我的工具——real-fake-data.com——解决了所有这些问题。 - 您想接受德国身份证、美国车辆号码、西班牙18岁以上的个人、来自波兰的真实地址——我们有超过300个真实数据生成器(真正存在,保证校验和,遵循所有规则)。 - 您想创建一个包含30个用户的种子数据库,每个用户都有几个订单、支付详情和日志——带有有意义的真实时间戳——只需定义数据模式,数据就会生成。 - 您想在一个恶劣的环境中运行测试——我们可以做到——您可以为每个生成器切换到正常模式、边缘模式(正确的数据但在正确性边缘,例如:出生日期是昨天,最长的姓氏,波兰最短的车辆号码)、极端模式(正确的数据但故意制造问题,例如未修剪的空格、新行、隐藏的UTF字符等)和无效模式(不正确的数据,用于检查您的表单是否会拒绝或正确处理)。 - 您想使用相同的数据重新运行测试——这可以做到,使用种子号,您将始终获得相同的随机数据——因此您可以选择何时需要随机数据,何时需要重复的安全性。 - 您希望您的Claude以低成本为您提供数据?太好了,MCP服务器已准备好供您使用。 - 您想轻松编写Playwright测试?只需`const person = await fakeData.plPerson({ sex: 'f' });`,您就可以完成。 - 您想使用VSCode插件在浏览器中直接获取测试数据,而无需离开?很好——按Ctrl+Shift+P“生成虚假UUID”——您就可以在不离开VS Code界面的情况下准备好数据。 - 您想确保任何数据都不会让您的软件面临风险?很好——确保种子是随机的,开启边缘模式,您将确信一旦现实世界中出现新的数据格式,您的管道将会针对它们进行测试。 这要花多少钱?对于简单使用,它是免费的。没有隐藏费用,不需要信用卡,也没有任何形式的月订阅。每月2000个免费代币。 我非常希望能收到更多的反馈,并听听您对这个工具的看法。 它还具有MCP插件,VS Code扩展中的Playwright插件,允许您在不离开IDE的情况下获取任何数据。 祝好, Karol Nowakowski
查看原文
Hello guys.<p>I want to interest You guys in a tool I&#x27;ve been working on for a last few months - real-fake-data generator. https:&#x2F;&#x2F;real-fake-data.com<p>The basic need it covers is to deliver test data that is (1) identical to the real data but not real, (2) make sure Your pipelines will catch real life problems - not only those which developer will remember to cover when developing.<p>First of all - reality changes. You can have a submission form with vehicle registration number and You have a problem (different countries, different rules). You will either spend month on digging various scenarios or You will just let it go and put a &quot;max length 7&quot; and go away. Your tool is up and working - BUT someday some country or state decides to have lenght-8 for vehicle plates and .... your system fails (users cant input new vehicle plates) while your test pipelines work and are sending You fake-green smile.<p>Second of all - You really cant predict everythin. You are open for users from around the world? Nice but they use different alphabets. Your beautiful UX&#x2F;UI will crash with user initials becoming 4 or 5 letters. Or the field for surname will be just too short. Or will it accept non-latin characters used even in Europe?<p>When You create new software You also need a seed starting data - so You can see (while developing) Your tool with some mock users, mock posts, mock products, mock transactions etc. Again - You either spend a lot of time writing it - even with help of AI - either You just leave user1, user2, user3 and You loose ability to look at Your tool in a way real life user will be looking.<p>My tool - real-fake-data.com - fixes all of it. - you want to accept german ID document, USA vehicle number, spanish person over 18 y.o., real existing address from Poland - we have over 300 generators of real data (really existing, checksum guaranteed, following all the rules) - you want to create a seed database of 30 users, each with few orders, each with payment details, and logs - with REAL timestamps that make sense - just define data schema - and You are covered - data is generated - you want to run Your tests in a hostile environment - we got this - You can switch for every generator from normal mode into: EDGE (correct data but on edge of correctness - eg. born date yesterday, longest possible surname, shortest vehicle number from Poland), EXTREME (correct data but deliberately made problematic with spaces untrimmed, new lines, hidden UTF characters, etc.) and INVALID (incorrect data that is useful to check if Your form with reject it or behave correctly) - you want to rerun the test with THE SAME data - it&#x27;s there with a SEED number You will always receive the same random data - so You choose whenever You want random things and whenever You want repeated security - you want Your Claude to deliver You data on low cost? Great to hear - MCP server is ready for You to use - you want to write playwright tests easily nice `const person = await fakeData.plPerson({ sex: &#x27;f&#x27; });` and You are covered - you want to use VSCode Addon to have test data directly on the browser without leaving it? Cool - Ctrl+Shift+P &quot;generate fake UUID&quot; - and You have it ready without leaving the VS code screen - you want to be sure that none data will ever put Your software at risk - great - make sure seeds are random, turn on edge mode and You will be sure as soon as some new formats of data will be there in real world Your pipeline will be tested against them.<p>How much does it cost? For simple using it costs Nothing. No hidden fees, no credit card required, no monthly subscription of any kind. 2000 free tokens&#x2F;month.<p>I would really love to receive more insight and hear Your opinion on the tool.<p>It also have MCP addon, Playwright addon on VS Code extension that allows You to grab any data without ever leaving the IDE.<p>with regards, Karol Nowakowski