请问HN:在审查之前,您将如何增强对一百万行遗留SaaS的AI更改的安全性?

2 分•作者: thegreatkahuna•2 个月前•原帖
我不是软件工程师,但我一直在进行一个实验,看看代理开发是否能够在现有的SaaS代码库上生成一个有用的原型。<p>这个代码库有超过100万行,已有15年历史,托管在Azure上,主要使用C#和React编写。<p>原型需要在九月份提供给客户进行测试。虽然没有开发人员能够全职参与,但我偶尔可以获得一些特定技术问题的帮助。一名工程师将在八月份评估实现情况,并决定我们对AI生成代码的信心程度,以便我们决定如何将代码“转换”为生产级别。<p>我的问题是,在八月份之前,我可以使用AI工具和手动测试做些什么,以使代码尽可能稳健和可审查。我希望增加AI生成的代码质量,以便将其推向生产级别的路径更接近“复制粘贴”,而不是从头开始重新构建一切。<p>环境 原型正在一个独立的分支上开发,部署到一个独立的内部环境,并连接到自己的数据库和架构。<p>规划过程 我们在六月份采访了客户,并将结果转化为MVP规格。<p>这个过程大致如下: 1. 编写PRD(产品需求文档)。 2. 使用大型语言模型(LLM)将PRD转换为架构文档,并由架构师进行审查。 3. 创建产品设计,包括屏幕图像和包含交互细节及边缘案例的md文件,使用Claude Design。 4. 使用代理将工作分解为史诗任务,使用PRD和架构文档作为指导。<p>这些史诗任务是最细粒度的规划文档,经过我和架构师的审查。<p>开发过程 开发流程旨在尽量减少干预: 1. 规划代理将史诗任务转换为md故事文件和Jira故事。 2. 编码代理实现故事,包括测试,并打开PR(拉取请求)。 3. 审查代理审查PR,要求更改,并将其合并到原型分支中。<p>编码代理会轮询PR以获取审查意见,并可以将问题升级回规划者。<p>代理还有一个“停止并询问”的列表,用于他们不被允许自主做出的决策。我(以及几次一名工程师)通过解决这些升级问题并每天手动测试累积的更改来参与其中。<p>大多数实现和初步审查都是通过基于Claude的代理完成的。对于风险较高的PR,我还使用Codex作为审查者。第二个模型的审查在Claude生成的代码中发现了更多相关问题,但令牌配额限制了使用。<p>我还对代码进行了单独的重构和加固运行。<p>结果 规划、设置环境和代理流程大约花了两周时间,然后代理在大约两周内构建了整个MVP。从规模上看,它包含了13000行功能代码,以及相同数量的测试代码。<p>我希望得到的建议 假设在审查之前我无法获得实质性的开发人员参与,我该如何增加代码尽可能接近生产级别的可能性?<p>以下是我一直在思考的一些问题: 1. 在工程师审查代码之前,哪些检查或开发循环会带来最大的信心提升? 2. 你会如何使用独立的代理或模型来降低编码者和审查者做出相同错误假设的风险? 3. 测试是否应该由与实现代码不同的代理生成? 4. 什么文档或证据会使最终的工程审查更快且更可靠? 5. 如果你只有几周时间来改善这个原型,然后交给工程师,你会优先考虑什么?
查看原文
I’m not a software engineer, but I’ve been running an experiment to see whether agentic development could produce a useful prototype on top of an existing SaaS codebase.<p>The codebase is 1M+ lines, 15 years old, hosted on Azure, and primarily written in C# and React.<p>The prototype needs to be available for customer testing in September. No developers were available to work on it full-time, although I could occasionally get help with specific technical issues. An engineer will evaluate the implementation in August and decide how much confidence we can have in the AI-generated code so that we can decide how to “convert” the code to production-grade.<p>My question is what I can do before August, primarily using AI tools and manual testing, to make the code as robust and reviewable as possible. I want to increase the likelihood that the AI-generated code would be so good that the path to production-grade would be closer to “copy-paste” than building everything again from scratch.<p>Environment The prototype is being developed in a separate branch, deployed to a separate internal environment, and connected to its own database and schema.<p>Planning process We interviewed customers in June and turned the resuls into an MVP spec.<p>The process was approximately: 1. Write a PRD. 2. Use an LLM to convert the PRD into an architecture document, which was reviewed by an architect. 3. Create product designs consisting of screen images and md files containing interaction details ans edge cases with Claude Design. 4. Use an agent to break the work into epics using the PRD and architecture document as guardrails.<p>The epics were the most granular planning artifacts that received review by me and the architect.<p>Development process The development flow was intended to run with little intervention: 1. A planner agent converted epics into md story files and Jira stories 2. A coding agent implemented the stories including tests and opened PRs 3. A reviewer agent reviewed the PRs, requested changes, and merged them into the prototype branch<p>The coding agent polled PRs for review comments and could escalate issues back to the planner.<p>Agents also had a “stop and ask” list for decisions they were not allowed to make autonomously. I (and a few times an engineer) were involved by resolving those escalations and by manually testing the accumulated changes end to end each day.<p>Most implementation and initial review were done with Claude-based agents. For riskier PRs, I also used Codex as a reviewer. The second-model review found substantially more relevant issues in the Claude-generated code, but token quotas limited the use.<p>I had separate refactoring and harden runs for the code as well.<p>Results The planning and setting up the environment and agentic flow took about two weeks and then the agents built the whole MVP in about two weeks. Size-wise it was 13k lines of functional code + the same amount for tests.<p>What I would like advice on Assuming that I cannot get substantial developer involvement before the review, how can I increase the likelihood that the code is as close to production-grade as possible?<p>Here are some of the questions I have been thinking about: 1. What checks or development loops would give the largest improvement in confidence before an engineer reviews the code? 2. How would you use independent agents or models to reduce the risk that the coder and reviewer make the same incorrect assumptions? 3. Should tests be generated by a separate agent from the one that wrote the implementation? 4. What documentation or evidence would make the eventual engineering review faster and more reliable? 5. If you had only a few weeks to improve this prototype before handing it to an engineer, what would you prioritize?