启动 HN:Screenpipe(YC S26)——通过 24/7 屏幕录制为您的代理提供动力
嗨,Hacker News,我是Louis。我创建了Screenpipe([https://screenpipe.com](https://screenpipe.com))。Screenpipe是一款应用程序,能够本地(仅限本地)录制您的屏幕和音频,并为AI代理提供可搜索的记忆,记录您所看到、听到和说过的内容。这使得自动化重复任务变得更加容易,并将其转化为标准操作程序(SOP)等。
我制作了一个HN风格的演示视频,可以在这里观看:[https://www.tella.tv/video/build-your-ai-second-brain-with-screenpipe-e1j7](https://www.tella.tv/video/build-your-ai-second-brain-with-screenpipe-e1j7),还有一个营销视频:[https://www.youtube.com/watch?v=c1jV6E9pyug](https://www.youtube.com/watch?v=c1jV6E9pyug)。
我对这个项目一直非常痴迷。从2020年开始,我就一直在维护一个“第二大脑”,在其中存储日记、手写笔记、我听的音乐、我正在进行的项目、与人交谈的内容、个人客户关系管理等。我在早期与ParlAI、数百个微调的GPT2模型和GPT3进行了大量的RAG实验([https://forum.obsidian.md/t/fine-tuning-openai-api-gpt3-on-your-second-brain-obsidian/21849/4](https://forum.obsidian.md/t/fine-tuning-openai-api-gpt3-on-your-second-brain-obsidian/21849/4))。后来,我构建了Ava,这是第一个Obsidian AI插件,迅速吸引了几千名用户。它后来演变为Embedbase,一个API,旨在简化构建基于RAG的AI应用程序。
我从中学到的最重要的一点是,模型需要了解您在计算机上所做的事情的上下文,以便能够完成您想要的操作。
在早期,虽然有微调的过程,但这太麻烦了;接着出现了工具调用,使得AI能够访问您使用的软件,但仍然不够自主,需要微管理。然后出现了MCP,但它感觉太静态,非技术用户在构建和使用MCP时遇到困难。随后我们得到了技能。最近,我们看到了Karpethy的LLM维护的维基、Garry的GBrain等,其中一个代理逐步维护一个持久的Markdown页面集合。新的来源更新实体页面,增强或反驳现有的主张,并改善随时间推移而累积的综合。我喜欢这种模式,但它仍然始于有人选择和导入来源。AI仍然无法了解您和您的公司每天在各个应用程序中的活动,而不仅仅是在应用程序内部。
当然,并不是每个人都想要这个。但我想要!我希望AI知道我在做什么,并且永远不会失去记忆,我希望它使用与人类相同的软件,而不需要痛苦的上下文切换。
我在2024年开始为自己构建Screenpipe——一个命令行工具,用于录制您的屏幕并将这些上下文信息传递给AI。2024年,一位HN用户在这里发布了它([https://news.ycombinator.com/item?id=41695840](https://news.ycombinator.com/item?id=41695840)),那次讨论对产品产生了影响。最有用的批评集中在录制同意、本地安全性、CPU使用、信噪比以及代理是否能够在数据之上进行操作等方面。
最初的简单实现是不断录制视频并对每一帧进行OCR处理。但这会产生重复数据,消耗大量资源(基本上会把您的计算机变成一个取暖器!),并且丢弃了操作系统已经知道的结构。现在,Screenpipe改为监听应用程序切换、点击、输入暂停、滚动和空闲回退等事件。当某些有意义的变化发生时,它会在相同的时间戳下将屏幕截图与操作系统的可访问性树配对。当结构化的可访问性数据不可用时,会使用OCR。我们还会持续捕获音频,识别说话者并通过Parakeet/Whisper或使用云模型进行本地转录。
所有数据都在本地SQLite数据库、mp4文件和有时的md文件中进行索引。一个友好的AI API在3030端口开放给代理,带有身份验证、MCP和技能。
一旦Screenpipe运行了一段时间,您可以通过我们内置的聊天工具Claude、ChatGPT、Hermes、Openclaw或任何代理来执行以下操作:
- 为您当前的聊天添加上下文,例如“收集关于任务X的所有上下文”,从而减少实现目标所需的提示
- 检索信息,例如“检索我从上午8点到下午4点所做的任务,列出已完成和未完成的事项”
- 为您的代理创建和维护个人维基/第二大脑:“每小时将我在Obsidian库中进行的项目、人员、任务、会议等内容整理为Markdown文件和文件夹”
- 创建自动化:每当我访问某人的LinkedIn资料时,更新我的CRM
- 寻找自动化机会:查看我团队本周所做的一切,并将其转化为自动化机会列表
Screenpipe的数据是本地存储的,尽管我们也提供企业计划,以发现自动化机会,具体由公司决定数据存放的位置。我们构建了自己的AI PII模型来删除敏感信息,它在Apple MLX或Windows DirectML上本地运行,我们还支持低端设备的云保密推理,尽管我们的本地模型旨在使用<1% CPU和<400 MB RAM。用户可以设置应用程序、窗口和URL进行过滤,此外还有浏览器隐身模式。
我们还支持录制计划和其他隐私功能。
我们的代码库大部分是用Rust、MLX、Onnx编写的,我们喜欢使用cidre或直接C调用Apple API,以及windows-rs用于Windows API。我们还实验性地支持Linux。
我们有一个桌面应用程序([https://screenpipe.com/how-to-install](https://screenpipe.com/how-to-install))和一个命令行工具:
```
npx screenpipe record
```
您可以在不创建账户的情况下运行它。所有代码都可以在这里找到:[https://github.com/screenpipe/screenpipe](https://github.com/screenpipe/screenpipe)。我们采取了令人畏惧的步骤,制定了自己的Screenpipe商业许可证。我知道HN强烈偏好OSI开源(MIT/Apache等),但找不到可持续的方式来继续开发Screenpipe,同时让公司免费使用它。因此,现在个人非商业、非营利、教育和研究用途是免费的,但商业用途需要许可证。
在许可证变更之前发布的版本仍然以MIT协议提供。我们有一个免费层和其他计划,包括企业计划,帮助公司发现自动化机会。
期待听到您对Screenpipe的反馈、您使用Screenpipe所做的事情或您希望的功能。
查看原文
Hi Hacker News, I'm Louis. I built Screenpipe (<a href="https://screenpipe.com">https://screenpipe.com</a>). Screenpipe is an app that records your screen and audio locally (only!), and gives AI agents a searchable memory of what you've seen, said, and heard. This makes it easier to automate your repetitive tasks, turn them into SOPs (Standard Operating Procedure) and so on.<p>I made a HN-style demo video at <a href="https://www.tella.tv/video/build-your-ai-second-brain-with-screenpipe-e1j7" rel="nofollow">https://www.tella.tv/video/build-your-ai-second-brain-with-s...</a> and there’s a marketing video at <a href="https://www.youtube.com/watch?v=c1jV6E9pyug" rel="nofollow">https://www.youtube.com/watch?v=c1jV6E9pyug</a>.<p>I’ve been obsessed with this for a long time. I’ve been maintaining a “second brain” since 2020, in which I would store journals, handwritten notes, music I listen to, projects I'm working on, conversations I have with people, personal CRM etc. I experimented a lot of RAG in the early days with ParlAI, hundreds of fine-tuned GPT2 models, and GPT3 (<a href="https://forum.obsidian.md/t/fine-tuning-openai-api-gpt3-on-your-second-brain-obsidian/21849/4" rel="nofollow">https://forum.obsidian.md/t/fine-tuning-openai-api-gpt3-on-y...</a>). Later I built Ava, the first Obsidian AI plugin, which grew to a few thousands of users quickly. It then became Embedbase, an API to make it easier to build AI apps powered by RAG.<p>What I learned from all this is how important it is for the models to have context about what you’re doing on your computer, in order to get them to do what you want.<p>In the early days there was fine tuning but it was too much pain, then there was tool calling so that AI can access software you use but still kinda not autonomous enough. needing micro management. Then MCP came, but it felt too static, and non technical users struggled to build and use MCP. Then we got skills. Most recently we’ve seen Karpethy’s LLM-maintained wiki, Garry's GBrain, etc., where an agent incrementally maintains a persistent collection of Markdown pages. New sources update entity pages, strengthen or contradict existing claims, and improve a synthesis that compounds over time. I like this pattern, but it still begins with someone selecting and importing the sources. There is still no way AI can know what you and your company are doing every day, across apps, not just inside of apps.<p>Of course, not everyone wants this. But I do! I want AI to know what I'm doing and never lose memory ever again, and I want it to use the same software that humans do, without painful context switches.<p>I started building Screenpipe for myself in 2024 - a CLI to record your screen and plug this context into AI. An HN user posted it in 2024 (<a href="https://news.ycombinator.com/item?id=41695840">https://news.ycombinator.com/item?id=41695840</a>) and that discussion influenced the product. The most useful criticism concerned recording consent, local security, CPU usage, signal-to-noise, and whether agents could act on top of the data.<p>The naive implementation started from continuously recording video and running OCR over every frame. But that creates duplicate data, consumes substantial resources (it basically turns your computer into a space heater!), and discards structure the operating system already knows. Screenpipe now instead listens for events such as app switches, clicks, typing pauses, scrolling, and idle fallbacks. When something meaningful changes, it pairs a screenshot with the operating system’s accessibility tree at the same timestamp. OCR is used when structured accessibility data is unavailable. We also capture audio continuously, identify speakers and transcribe locally through Parakeet/Whisper or using cloud models.<p>Everything is indexed in a local SQLite database, mp4 files, and sometimes md files. An AI friendly API on port 3030 is open for agents, with authentication and a MCP and skills.<p>Once Screenpipe has been up and running for a while, you can use it through our built-in chat, Claude, ChatGPT, Hermes, Openclaw, or any agent, to do things like:<p>- adding context to your current chat, e.g. "gather all context about task X", then requiring less prompts to achieve your goal<p>- retrieve information, e.g. "retrieve the tasks i was working on from 8 am to 4 pm, make a list of what got done and what's left"<p>- create and maintain a personal wiki / second brain for your agents: "every 1h organize everything i do in projects, people, tasks, meetings in my Obsidian vault as markdown files and folders"<p>- create automations: whenever i visit someone's profile on linkedin, update my crm<p>- find automation opportunities: look at everything my team has done this week and turn it into a list of automation opportunities<p>Screenpipe data is stored locally, though we also offer an enterprise plan to discover automation opportunities and for that the company decides where the data lives. We built our own AI PII model to redact sensitive information, it runs locally on Apple MLX or Windows DirectML, we also support cloud confidential inference for low end devices, although our local models are meant to use <1% CPU and <400 mb RAM. Users can set apps, windows, and urls to filter, in addition to browser incognito mode.<p>We also support recording schedules and other privacy features.<p>Most of our codebase is written in Rust, MLX, Onnx, we like cidre or direct C call for Apple APIs and windows-rs for Windows API. We also experimentally support Linux.<p>We have a desktop app (<a href="https://screenpipe.com/how-to-install">https://screenpipe.com/how-to-install</a>) and a CLI:<p><pre><code> npx screenpipe record
</code></pre>
You can run that without creating an account. All the code is source-available at <a href="https://github.com/screenpipe/screenpipe" rel="nofollow">https://github.com/screenpipe/screenpipe</a>. We took the dreaded step of making our own Screenpipe Commercial License. I know HN strongly prefers OSI open source (MIT/Apache/etc.) but couldn’t find a sustainable way to keep developing Screenpipe while companies were using it commercially for free. So now personal non-commercial, nonprofit, educational, and research use is free, but commercial use requires a license.<p>Versions released before the license change remain available under MIT. We have a free tier, and other plans, including Enterprise which helps companies find automation opportunities.<p>Would love to hear any feedback, things you've done with screenpipe, or features you'd want