Product Hunt 每日热榜 2026-07-21

PH热榜 | 2026-07-21

#1
Lev8
Find, research, and reach the right people
493
一句话介绍:Lev8 是一款通过 AI 多智能体实时搜索网络,帮助用户快速定位并触达目标人群与公司的智能获客工具,解决传统 B2B 销售中“找人难、验证慢、触达乱”的痛点。
Sales Artificial Intelligence Marketing automation
AI销售助手 B2B智能获客 实时数据挖掘 多智能体搜索 人群触达自动化 用户意图监控 CSV数据 enrichment 多通道消息发送 销售智能平台 人脉发现工具
用户评论摘要:用户普遍认可其实时搜索能力优于静态数据库,但质疑数据源的透明度与准确性。核心问题包括:如何处理同名/多角色身份歧义、防止基于弱信号生成垃圾邮件、多通道触达的合规性及交付能力。建议增加结果置信度评分和证据层展示。
AI 锐评

Lev8 的核心价值不在于“找到更多人”,而在于“在正确时间找到正确的人并配以可验证的上下文”。其“对话式搜索+多智能体并行爬取+身份消歧+意图监控+多通道自动化触达”的链条,确实比传统静态数据库或单一搜索工具更接近销售场景的真实需求——用户需要的不只是名单,而是可行动作为(如“给 GitHub 星标用户发私信”)。

但产品面临的信任危机远比功能突破更严峻。评论中反复提及的“数据源是否合规”“AI 是否生成垃圾信息”“多通道触达是否触发反垃圾机制”并非杞人忧天:当工具同时掌控“爬取公开数据”“生成个性化话术”“自动化发送”三个环节,一旦任何一环失控(如爬取 LinkedIn 导致 IP 被封、AI 编造意图信号、批量发送触发邮箱封号),产品就会从效率工具变为风险炸弹。

值得注意的是,团队在回复中强调“使用授权 API”“跨模型验证”“分通道限速发送”,这恰恰说明他们清楚问题所在。但“合规依赖用户自行判断”的免责声明,暗示了产品在灰色地带上的某种滑行。对于真正追求可持续获客的团队而言,这个“快”字背后隐含的合规成本,可能需要额外几个“慢”来对冲。

真正让 Lev8 脱颖而出的,或许并非技术护城河,而是它精准拿捏了“销售团队对效率的焦虑”与“数据可用性与合规性之间的张力”。如果它能将“可验证的实时线索”这一核心壁垒,与“自动化触达中的合规设计”真正深度融合,而非仅仅作为吸引用户上车的皮,才算是从“快刀”蜕变为“利器”。

查看原始信息
Lev8
Chat with Lev8, the fastest way to find and reach your target people and companies, powered by live web search via parallel AI agents across every corner of the internet. Enrich CSVs with waterfall lookups, monitor intent signals, and send personalized multi-channel messages, all running automatically.

👋 Hey Product Hunters,


I'm Tony Zhang, co-founder of Lev8 , the fastest agent to find and reach your target people and companies across every corner of the internet and reach across multiple channels.

Here are a few searches our early users have thrown at Lev8:

  • “Find me VPs of Sales at fast-growing voice agent startups in the Bay Area that raised funding recently.”

  • “Find everyone who starred our GitHub repo, then reach out to them across multiple channels.”

  • “Find coffee shops in San Francisco with a 4.5+ rating and no website.”

AI has made it easier to build products. Getting the right people to notice them is still hard.

We felt this ourselves. Our team spent hours jumping between search tools, databases, spreadsheets, and enrichment services just to answer a simple question:

Who should we talk to, and why?

So we built Lev8.

What Lev8 does

Lev8 turns the live web into people and business intelligence.

Describe who you’re looking for in plain language, and Lev8:

  • Searches across public sources

  • Verifies identities

  • Adds relevant context

  • Surfaces the signals that explain who matters right now

What happens behind the scenes

Our crawler reaches sources that static databases miss. Parallel agents explore the web, while our identity system makes sure the facts belong to the right person or company.

The goal isn’t to create another oversized lead list.

It’s to help you discover the companies and people others miss—and provide reliable intelligence that both people and AI agents can use to make better decisions.

24
回复

🎁 For the Product Hunt community

Product Hunt members get 500 free credits today.

👉 Try Lev8.com with a search you couldn’t solve before, and tell me how it goes. I’d especially love to hear what you searched for and where the results could be better.

I’ll be reading every comment.

9
回复

@nora_yu Congrats mate
all the best

0
回复

@nora_yu Really impressive launch! One question: your live-web approach is a huge differentiator over static databases. How does Lev8 balance freshness with accuracy when multiple public sources conflict? Is there a confidence score or evidence layer that helps users understand why a particular recommendation or contact was surfaced?

0
回复

@Lev8 is one of the most technically grounded agent systems I’ve seen for people and company intelligence, and it works remarkably well.

The first time I tried it, I felt it had already gone beyond any general-purpose search tool I had used for finding and matching the right people and companies.

Search is an extremely long-tail problem. Lev8 handles it with a multi-agent system that can move quickly across a huge amount of scattered information and identify the exact people or companies you are looking for.

Quite often, it opens up a part of the market you did not even know existed.

It also does not stop at discovery. Lev8 connects the results to the social and outreach channels you already use, so the same workflow can continue all the way to the first conversation.

If connecting with people and companies is part of your work, give Lev8 a try. It may be the most accurate and efficient tool available for this job right now!

10
回复

@zaczuo Sounds very useful and the demo looks awesome. But what data sources does it actually use and how does it handle people/companies with similar (or the same) names? Can we see the underlying data to verify it?

5
回复

Congrats on the launch! I searched for a pretty unusual customer profile and Lev8 understood the request better than I expected. There were a couple of results I’d remove, but the overall direction was solid.

5
回复

@sandy_liusy Appreciate the congrats and the feedback! Really glad to hear Lev8 nailed the direction on a tough, non-standard search, that deep comprehension is exactly what we built our agent swarm for, going way beyond simple keyword matching to actually understand the context you're looking for. I appreciate you pointing out the couple of off-target results as well. We’re constantly fine-tuning our scoring and qualification layers to filter out that extra noise, so this helps us make the agent matching even sharper.

0
回复

the "who should we talk to, and why?" framing is exactly the difficult part. as a founder, finding names is usually not the bottleneck anymore, finding people who are relevant right now and having enough real context to write something that does not feel like mass outreach is

the live web search and intent signals sound especially useful for launch prep, partnerships, and early sales. Curious how Lev8 shows confidence and sources behind each result, and how you prevent the personalized outreach from becoming confidently written spam based on weak or outdated signals :)

2
回复

@andrasczeizel Finding names is easy, but having real context to reach out right now is the actual bottleneck. To prevent "confidently written spam," Lev8 ditches static databases to mine live web signals (like GitHub, forums, and tech stack shifts) in real time, with every insight linked directly back to its live source so you can verify it yourself. While AI hallucinations are a real challenge, we run cross-model validation across multiple LLMs—accepting a data point only when all models agree.

On top of that, we run Signal Scoring to filter out weak or stale inputs, ensuring outreach is only triggered when there is true urgency and ICP fit. Instead of rigid templates, the messaging is context-native and grounded in verifiable events, acting more like a research teammate than a cold-spam machine. Would love to get your thoughts if you're gearing up for a launch soon!

3
回复

Really like the 'pick a play, tell Lev8 to run it' framing instead of another dashboard to configure.
How much does it actually get right on the first try vs needing you to tweak the ICP/messaging after a few runs?

2
回复

@abod_rehman Lev8 is designed to get you started with very little setup. If you're new to outbound, a simple description of your goal is

enough for the agent to find relevant prospects, create personalized messages, and handle the outreach. If you're more

experienced, you can fine-tune the results by sharing your own review criteria and preferred communication style

0
回复

Tested the launch myself, and I really liked it.

Spent about 10 minutes chatting with it and got answers to a bunch of questions I had about competitor tools. I also liked the visual lists, they make the information much easier to explore.

Still getting familiar with everything it can do.

Good luck!

2
回复

@adana Awesome to hear! Super happy that the visual lists hit the mark and helped you get quick answers on competitor tools.

Have fun exploring the rest of the features, and thanks a lot for the good luck wishes :)

1
回复

How transparent are the results? Can users see the original sources behind each data point before using it for outreach?

2
回复

@jody_l_wyatt Yes. Lev8 includes the original source URL for the data it finds, whether from LinkedIn or another platform, so you can verify it during review. You can also review the hooks behind each personalized outreach message. If anything is unclear, you can directly ask Lev8 agent to investigate and cross-check a specific data point in more depth.

1
回复

Really liking the idea behind Lev8! It helped me discover relevant potential users for my project, which was exactly what I was looking for. One thing I noticed (and this could just have been my experience) was that none of the suggested users had email addresses available. I'm not sure if I was just unlucky or if it's because of the data source, but I thought I'd mention it. Overall, it's a really promising product, and I'm excited to see how it evolves. Great work! 🚀

1
回复

@archit_jha Sorry about that. Are you running into an issue? Feel free to email me at tony@lev8.com, and I’ll be happy to help.

0
回复
Congrats on the launch!
1
回复

@thisiskp_ Cheers!

0
回复

Can I input a list of Typeform inbound submissions and have Lev8 instantly enrich and score them before pushing to Slack?

1
回复

@jocky Yes. You can send your Typeform inbound submissions to Lev8, and it will enrich each lead, score them against your criteria, and push the results to Slack.

1
回复

Congrats on the launch @tony_zhang! The live web crawler approach is a huge step up from static databases that lag by 6+ months.

Because you're scraping real-time signals (like GitHub activity or forum posts), how does the identity system disambiguate someone who maintains multiple roles—e.g., a VP of Sales who is also an advisor or founder at a stealth startup? Does Lev8 tie the intent signal specifically to their primary domain/company context before generating the outreach hook?

1
回复

@franz_briones Yes. You can have Lev8 enrich the data you need, identify relevant email hooks, or search and scrape for specific signals you want to reference. Lev8 can then use those insights to generate highly personalized, engaging emails or outreach messages.

0
回复

the outreach-side deliverability question above is a good one, but I'm curious about the other direction: the crawler side. pulling live data from LinkedIn and similar platforms at agent speed is exactly the kind of activity those platforms actively try to detect and rate-limit or ban accounts for. is Lev8 hitting these sources through some kind of licensed/API access, or is it closer to scraping, and if the latter, does the risk of a flagged account sit with Lev8's infrastructure or with the user's own connected accounts?

1
回复

@galdayan  Lev8 uses authorized LinkedIn data APIs rather than scraping through users’ own accounts. That means users don’t need to connect or risk their personal LinkedIn accounts, and the data access is handled through Lev8’s compliant infrastructure.

2
回复

the three-layer qualification process is the interesting part here. most lead gen tools just give you a list and let you figure out if it's any good. the fact that you're verifying reachability before send is huge — biggest time sink in outbound is chasing contacts that bounced or changed roles 6 months ago. curious how the intent signals layer works in practice though, like what's the false positive rate on "hiring activity" triggers?

0
回复

@ozandag We don’t publish a false positive rate because intent signals vary by source and by the type of signal being tracked. Rather than relying on a single event like “this company posted a job,” Lev8 combines multiple publicly available signals to build confidence before surfacing an account. Hiring activity can be combined with signals like company growth, leadership changes, new product launches, funding, or other relevant events depending on the workflow.

We also don’t treat intent signals as an automatic “send” trigger. They’re used as context to help prioritize prospects and personalize outreach. Before sending, Lev8 researches each prospect, identifies the signal most relevant to your offer, and turns it into a personalized hook rather than relying on a generic template.

0
回复

Nice launch! Going beyond "find people" into personalized multi-channel outreach is a smart move.🌟

Curious how you handle deliverability and rate limits across channels as send volume scales?

0
回复

@aymi_malik Lev8 doesn’t blast contacts across every channel at once. Each channel has its own pacing, daily limits, and sending strategy, with outreach managed through controlled sequences. You can also define your own sequence and preferences directly in chat. This helps protect connected accounts and maintain deliverability as volume grows

0
回复

the multi-channel outreach part is where I'd want more detail before turning this loose - if it's finding someone's personal email/phone/social from public sources and then messaging them across several channels automatically, what's stopping that from tripping spam filters or straight up violating CAN-SPAM/GDPR rules on cold outreach consent? is there a built-in rate limit or opt-out handling per channel, or is that left to the user to figure out themselves?

0
回复

@omri_ben_shoham1 Good question. Lev8 doesn’t blast messages across every channel automatically.

For outreach, Lev8 uses a paced sending strategy rather than sending everything at once. Sending volume is limited based on the user’s account, and actions are spaced out over controlled time intervals to reduce account risk. Lev8 manages those intervals, and we’ll let users customize them in the future.

Compliance is still important. Lev8 is designed to help users find publicly available business contact information and automate parts of the outreach workflow, but users remain responsible for ensuring their campaigns comply with applicable regulations (such as CAN-SPAM, GDPR, and platform policies), including obtaining any required legal basis and honoring opt-out requests where applicable. We don’t position the product as a way to bypass spam protections or consent requirements.

0
回复

@omri_ben_shoham1 We don’t blast every contact across all channels at once. Lev8 applies a channel-specific sending strategy to each connected channel. Once you share your outreach preferences, Lev8 follows those rules to determine how and when each message is sent.

0
回复

If I base outreach on the wrong information once, I start questioning every result. How do you help users build confidence in the data they're seeing?

0
回复

@reda_roqai_chaoui All of our information comes from real searches, with source links provided so you can verify every result.

0
回复

Solid for getting past the manual scrapping grind, the parallel agents actually return fresher results than I expected. Waterfall enrichment on a messy CSV saved me a couple hours yesterday.

0
回复

@zafer175063 Thanks for the support!

0
回复

Finally a way to stop stitching together five different enrichment tools. One live webset beats a pile of stale exports.

0
回复

@1251912798 Exactly. That’s our ultimate goal: freeing people from complex, repetitive work so they can focus on what actually matters.

0
回复

The "monitor intent signals + send personalized messages automatically" combo is where these tools live or die. Intent data is only as good as it is fresh, and personalization at scale usually reads as templated the moment volume goes up. How do you keep the outreach from feeling automated once someone's sending hundreds a day? Nice launch.

0
回复

@david_marko That’s exactly the challenge we’re focused on. Lev8 pulls fresh signals at the time of research, lets users define the template and personalization angles, and grounds each message in real person- and company-level context. It also uses channel-specific pacing and sequencing rather than blasting everyone at once, so scale doesn’t have to mean generic outreach.

0
回复

This is good for research, but can you give some tips on what to do with that information?

0
回复

@carlos_mendez16 Lev8 is chat-driven. Just tell it in natural language what you’re looking for, whether that’s people, companies, or specific information, and it will search for you.

0
回复

This is exactly the kind of tool I didn't know I needed. I've spent way too many hours manually cross-referencing LinkedIn and company sites for outreach lists. Curious how it handles smaller/local businesses (like dealerships or regional service companies) where the web footprint is thinner than a typical SaaS company. Does the agent flag when confidence is lower on those searches?

0
回复

@andre_ajemian That’s a great question. Smaller and local businesses can definitely be more challenging because there’s often less public information available than you’d find for a typical SaaS company.

Lev8 cross-checks information across multiple public sources rather than relying on a single source. When the available data is limited or inconsistent, we don’t treat those results the same as well-supported ones.

Instead, the agent evaluates the quality of the available evidence and indicates which results are recommended and which are less reliable, so users know where they can move quickly and where a manual review may be worthwhile. Our goal is to be transparent about the strength of the data rather than give every result the same level of confidence.

0
回复

Interesting product! How does Lev8 verify conflicting or outdated information from multiple web sources before automatically sending outreach?

0
回复

@ram_sarjal 

From a technical perspective:
Lev8 uses a hierarchical validation framework that breaks each search request into more granular conditions. Each condition is verified independently, while a global validation layer evaluates the results against the user’s original query and intent. This multi-level scoring process helps deliver more accurate and reliable results.

To reduce AI hallucinations, Lev8 has also built an offline entity database. Each entity is assigned a structured ID card containing its core information, which significantly reduces entity-matching errors—especially confusion between people or companies with the same name.

From the user experience perspective:
Every result comes from real searches and includes its original sources, so users can review the evidence and verify the accuracy themselves.

Hope this answers your question. Feel free to try it out at Lev8.com.

0
回复
What are these ‘public sources’ you use, LinkedIn?
0
回复

@thomas_digaetano Live open-web search across LinkedIn, YouTube, Instagram, TikTok, GitHub, patent databases, niche forums, and more.

0
回复

The enrichment looks useful, but I’d want to know how the data is verified. Can users see the source for each field, especially job title and company? That would make it much easier to trust before any messages are sent.

0
回复

@liana_preston All of our information comes from real searches, with source links provided so you can verify every result.

0
回复

Congrats on the launch. Since outreach channels get connected by authorizing your own account, are those credentials scoped so only your own campaigns can use them, or is there a shared pool the agents draw from across users?

0
回复

@vollos Each user’s connected accounts and credentials are fully isolated and only used for their own campaigns. We never pool or share authorized accounts across users.

0
回复

Would love a way to save and reuse my best-performing search prompts as templates, maybe with variables for company size or role. Right now I rebuild the same complex queries every time I start a new campaign, and it feels like the kind of thing this could handle automatically if I could just hit run on a preset.

0
回复

@selimgurelu792 Thanks for the support!

0
回复

How does this compare to using Clay?

0
回复

@pc4media Clay helps GTM experts build powerful workflows. Lev8.com works like an AI GTM teammate,just tell it your goal, and it handles prospecting, enrichment, and outreach in one place. No complex workflows, no juggling multiple third-party tools - just a simpler way to build pipeline.

1
回复

Pretty cool. I tested it and the search criteria based on what I asked for is pretty on point.

0
回复

@inferhaven Thanks for your support!

0
回复

Congrats on the launch! This sounds like a super useful tool - one question though. How is it different from an apollo or hunter? As an avid user of both (although I've never been particularly interested in purchasing a paid subscription from either) I'm curious as to how you guys are differentiated.

0
回复

@ethan_cheng Thanks so much for the support! Appreciate you taking the time to ask—Apollo and Hunter are absolute staples for standard contact lookups, so it’s a very fair comparison.

The core difference comes down to Static Database vs. Live AI Teammate. Standard tools like Apollo and Hunter rely heavily on pre-indexed, static databases, which can quickly go stale or completely miss niche, hard-to-reach decision-makers. Lev8 operates as an integrated agent swarm that mines the live open web in real-time across 50+ providers, social platforms (LinkedIn, X, YouTube, Instagram), GitHub, and technical communities to deliver broader and far more up-to-date results.

Beyond just finding contacts, Lev8 handles deep contextual understanding and automated action. Instead of relying on basic keyword matches, it stitches together a person's fragmented online presence across platforms and time to reveal true intent, background, and preferences. It also acts as an always-on market monitor that tracks dynamic timing signals—such as tech stack changes, hiring shifts, and web discussions—so you know exactly why and when to engage. Finally, rather than just exporting a list of emails into a generic sequencer, Lev8 works as a chat-based teammate that qualifies leads, ranks urgency, and drafts context-native outreach across Email, LinkedIn, WhatsApp, X, and Instagram in a unified inbox.

If you just need a quick email address lookup from a standard database, traditional tools work great. But if you want a live AI teammate that finds hard-to-reach people, decodes deep intent signals, and turns dynamic data into multi-channel action, that's where Lev8 shines.

0
回复

Parallel AI agents for live search beats static databases — but how fresh is the data on average? And for niche B2C (music teachers, not tech execs), does waterfall enrichment actually find contacts or is it LinkedIn-only depth?

0
回复

@ringo_td5 Public data can be refreshed daily, while data from platforms such as LinkedIn may have a delay of up to around a month.

Lev8 already searches across multiple sources rather than relying mainly on LinkedIn. We’re also developing Open Search, which will enable more targeted searches for niche contacts whose information is scattered across different platforms.

0
回复
#2
ditto.site
Clone any website into clean code. Free & open source
379
一句话介绍:ditto.site 能以极高的保真度将任意公开网站克隆为干净、组件化的 Next.js 或 Vite 代码,帮助开发者快速获取网站起始模板,彻底告别传统克隆工具“Div 汤”式的混乱输出。
Design Tools Open Source Developer Tools GitHub
网站克隆 代码生成 前端工具 开源 Next.js Vite 组件化 设计系统 MCP REST API
用户评论摘要:用户普遍称赞输出代码“出奇地干净”,尤其好评其对交互状态(悬浮、手风琴)和设计 Token 的保留。主要建议包括:希望自动映射谷歌字体、提供与原始站点的差异对比视图,以及更智能地处理客户端渲染内容。关于版权问题,团队澄清主要用途是从锁定平台搭建迁移,并非鼓励抄袭。
AI 锐评

Ditto 在这个 AI 克隆网站遍地开花的节点上,硬生生劈开了一条“确定性”的血路。它的核心价值不在于“抄”得准,而在于“拆”得干净。当大多数同类工具还在给开发者一堆繁琐的“Div 汤”时,Ditto 凭借其编译原理级别的 DOM 解析能力,输出了直觉上就不需要人工重组的组件结构。这一点,直击了 AI 开发者最大的痛点:从零开始搭建比迁移旧系统更痛苦。

它聪明的做法在于,把自己定位成一个“迁移工兵”而非“内容小偷”。通过将 “克隆” 包装成一个帮客户从封闭平台(如 Wix、WordPress 等)逃生的合法工具,它巧妙地规避了版权争议,同时将技术门槛降到了“喂一个网址”这么低。这也意味着,Ditto 的真正战场并非静态克隆,而是为“网页转代码”这一工作流提供无需训练、成本极低的前端 IO 接口。

当然,它最大的天花板同样来自这份“确定性”。那些依赖复杂 JS 动态生成的内容、滚动触发动画或登录后才显示的组件,Ditto 是无能为力的。它试图通过捕获不同状态下的 DOM 快照来弥补,但终究是盲人摸象。此外,缺乏有效的代码差异对比机制,会在版本迭代维护时让开发者陷入迷茫。总的来说,Ditto 是一个完美的起点工具,但它需要搭配一个智能的“后处理代理”才能真正形成闭环。这不是缺陷,而是项目定位使然——它只承诺给你一张干净的纸,画工好不好,还得看你自己。

查看原始信息
ditto.site
Point Ditto at any public URL and get a faithful copy as clean, componentized Next.js or Vite code, in minutes. It's fully deterministic, so it's fast, cheap, and consistent. Components, tokens, interactions & hover states included. Free REST API + MCP server, open source.
Hey Product Hunt 👋 We're the team behind ion.design, and today we're open-sourcing ditto.site. Ditto takes any website and gives you back a clean copy in Next.js or Vite. We extract actual components, sections, design tokens, and an editable content model. We even do our best to get hover states, focus states, dropdowns, accordions, declaritive motion, web fonts, SEO metadata, lottie animations, and responsive layout. Ditto is fully deterministic, so it is fast, cheap, and consistent. An AI-based approach can't deliver this. Why open source? We've been running a version of ditto inside ion.design since the start of the year and it's a bit too useful to keep to ourselves. ditto is an extremely valuable starting point for AI builders, because you can get a ready-to-go clone of a customer's website as soon as they sign up rather than a blank canvas. What you get: 🧩 Clean, componentized Next.js or Vite + TypeScript. Optional Tailwind styling. 🎯 Deterministic. no AI guesswork, byte-stable output. 🔌 Free hosted REST API + MCP server, or self-host for best speed. 🎨 Tokens, interactions, hover states, fonts, SEO preserved. 🪪 MIT licensed, fully open source. Try it: https://ditto.site Star us: https://github.com/ion-design/di... We're around all day! Give it a spin, tear it apart, tell us what you want, and use the hell out of it. We're all in this together, let's create user value. 💚 - The ion.design team
9
回复

@samraaj_bath1 I'm curious about the deterministic approach. If a source website changes slightly over time, can Ditto intelligently regenerate only the affected components while preserving developer modifications, or is regeneration treated as a fresh export? That workflow could make it incredibly useful for ongoing migrations and redesigns.

0
回复

@samraaj_bath1 Curious, how do you ensure the extracted design system stays reusable instead of generating components that are too specific to the original site?

0
回复

@samraaj_bath1 Really interesting! How does Ditto handle complex interactions and animations while keeping the generated components clean and reusable?

0
回复

Tried Ditto on a few landing pages and the output was surprisingly clean. One thing that would make this indispensable for me is if it could also extract and map the original page’s Google Fonts or custom font stack into the generated project automatically. Having to manually re-add fonts after cloning breaks the flow a bit.

3
回复

@yldzkfqa will look into this!

0
回复

@yldzkfqa ah yea, it may have hit an edge case on font extraction. we'll look into it and make a PR

1
回复

How can you defend the fact that the page was copied? I often see people then start to argue about who is copying whom.

3
回复

@busmark_w_nika We're a tool, i don't condone people stealing other people's stuff.

the main use case for ditto (which we use it for) is migration from locked website platforms. So a customer gets a clone of their own stuff they can port anywhere.

4
回复

The interesting constraint with deterministic cloning is that it only sees the rendered DOM, so anything gated behind auth, scroll triggers or JS state never gets captured. Curious how you handle client-rendered sections that only appear after interaction.

Really cool though - congrats on the launch!

3
回复

@aidan_codefox Yea JS is way harder, we didn't write a decompiler or anything. We are able to infer some stuff by capturing DOM diffs at different states (timing for intro animations, after clicks for menus/accordians)

we're just doing the max we can deterministically. in prod we run an agent on teh output of this to get even closer.

2
回复

Can confirm this thing works amazing. Even nailed a complicated Lottie interaction and some complicated layouts.

3
回复

@adamperlis Heck yes thank you Adam!!!

1
回复

Back in the '90s, we used a program called Teleport to clone websites, and we were absolutely amazed by how powerful it was. Time really flies—and so do technologies.

But honestly, this is a brilliant idea. Now we can just send every client who says, "I want exactly the same thing, just with my own logo," straight to you. 😂

3
回复

@maxchen Yeaaaa i have heart of teleport! Great name :)

ahahaha yea interesting use case. We typically do this for migrations from other website platforms but getting an inspired starting point works!

1
回复

The componentization is what got me, not just the clone but the fact that tokens and hover states come through intact. That level of fidelity on the interaction layer usually takes weeks of cleanup work.

2
回复

@kemalergdeydc3 Yea exactly. If you pass an agent on top of the ditto output, it's even better.

1
回复

@kemalergdeydc3 You love to see it

0
回复

Pointed it at my portfolio site and got back surprisingly clean component code with the hover states actually working. The deterministic part is no joke, second run matched the first exactly. Solid open source move with the REST API included too.

2
回复

@velilryc Yep we see this as a great community resource :)

1
回复

@velilryc would love to see that site!

0
回复

pulled up a couple of sites and the generated next code was honestly way cleaner than i expected, like the component split actually made sense and not just a wall of divs

2
回复

@nazifesarawpwe yep we have the concept of "Recipes" on our cloner, so it tries to match sub patterns and pull them into a component.

We also do some like-component-tree compiler magic to pull other stuff out.

1
回复

@nazifesarawpwe yuh we put a lot of effort into making sure it was as clean as possible.

0
回复

Love how fast and clean the output is, especially that it keeps hover states intact. One thing that would make this even better is letting me diff the generated code against the live site as it changes, so I can spot regressions during maintenance without having to rebuild from scratch.

2
回复

@beril3ljq good idea, but the thing is this is more for a starting point. We imagine the diff will be really hard ot manage over a longer scale rather than during a snapshot

1
回复

the fact that hover states and interaction details actually make it into the generated components is such a thoughtful touch. so many site cloner tools strip those out and you're left with static-looking pages

2
回复

@nuraypehlit7yr yea we coded in a lot to make it extract all that

1
回复

A diff view between the generated code and a hand-written baseline would be super useful so you can quickly spot where Ditto makes questionable choices, especially around accessibility or semantic HTML.

2
回复

@abdullahboyu Yea this is good in theory but hard in practice. CSS is one of those things where there are many ways to do the same thing. so diffs naturally diverge when u go into semantic, compiled languages.

we can diff the rendered DOM, but there's often a touch of drift with same visual output

1
回复

honestly this looks really useful, but one thing i'd love to see is a diff view between the original site and the generated code so i can quickly understand what ditto decided to extract or skip

2
回复

@rmeysagenaylpi yea good idea. make a pr for it!

1
回复

Honestly this is kind of wild, I pointed it at a random landing page and got back actually clean componentized code instead of the usual div soup. Hover states and tokens coming through too was a nice surprise.

2
回复

@abdulsametjo5z awesome!

0
回复

Love how fast and consistent Ditto is. One thing that would level it up for me: a visual diff mode where I can paste my own hand-written component next to the generated one and see them side by side. Perfect for cleaning up legacy markup and spotting what the AI missed or got wrong.

2
回复

ran it on a cluttered landing page i had bookmarked and the component split came out way cleaner than i expected, hover states and all. the free rest api is a nice bonus for quick experiments.

2
回复

@ayewyec love to hear this!

0
回复

Would love a "rebrand swap" mode where I paste a palette and font choice and ditto restyles every component in the generated output, so I can skip the tedious find-and-replace pass after cloning a competitor's layout. Would be a killer add for solo builders migrating inspiration into their own brand.

2
回复

@hamzanazloiwet Smart idea. ya try making this as a pr!

2
回复

Nice one, congrats on the launch. This looks genuinely useful for getting past the blank-canvas stage, especially when rebuilding an existing customer site. How about the quality of the output once you start editing it. Does Ditto identify sensible reusable components, or is the main goal visual accuracy with some cleanup still needed afterward? Also, how does it handle sites with heavy client-side rendering, forms, or third-party scripts? Those are usually the parts where a clean clone gets difficult.

2
回复

@os_ishmael it does identify repeated, structurally consistent patterns (e.g. cards, nav items, logos, list rows, etc.) and extracts them into reusable React components with data arrays. For heavily client-rendered sites, Ditto captures the post-JavaScript DOM, computed styles, lazy-loaded content, assets, and observable interactions. Forms are also reproduced visually. Various scripts and stuff may not carry over, as we are then dealing with compiled JS.

0
回复

@os_ishmael Yea so like michael mentioned, we have "Recipes" for common patterns where we hav estructures for re creating logic. but that's as far as that goes. For client side rendering and forms you'll still have to bring that in yourself (or with a followup agent)

1
回复

the determinism angle is such an underrated move, makes the whole thing actually predictable instead of gambling on a vibe-based clone every run

2
回复

@erva914472 Yep, also fast and cheap

1
回复

@samraaj_bath1 Congrats on open sourcing this! For a marketing site with a few dozen pages, does Ditto crawl and clone the whole site graph in one pass, or is it strictly one URL at a time and I'd need to loop it myself? That can change the real time/API cost of an actual migration job!

2
回复

@clement_avq Yep we have a parameter for single page vs multi page clones.

in prod we start with single page so we can show the customer something fast, and do multipage in teh background.

for repeatsed patterns (cms blogs, products, etcc) we recognize that and just make one of those entries at random. Soon we have a CMS extractor coming too :)

2
回复

The deterministic angle is the smart differentiator here, most "clone a site to code" tools lean on an LLM and you get a slightly different result every run, which makes them impossible to trust in a real workflow. Fast, cheap, and consistent is a much better pitch. Capturing tokens and hover/interaction states rather than just static markup is also the part people usually skip and then regret. The free REST API + MCP server plus open source is a generous combo. Curious how it handles sites that lean heavily on runtime JS rendering. Congrats on shipping!

2
回复

@kelly_king3 We dont do more than very basic JS! We are meant to be a starting point :)

1
回复
Congrats team, generous move open-sourcing this! My use case is probably the boring one: our own site. We’re on a hosted builder and the lock-in itch is real. You mention an editable content model, so if I point Ditto at our own CMS-driven marketing site, does repeated content (blog cards, testimonials) come out as mapped data I can actually edit in one place, or as hardcoded sections I’d untangle by hand? If it’s the former, this is a genuine escape hatch from website builder lock-in and that’s worth a lot more than cloning competitors.
2
回复

@ridhwikvinod Try it out! It's free anyway :) and log any issues on gh

Our algorithm finds repeated content and maps it out into a file that is easy to edit. And ya we built this for customers of ion.design migrating from wordpress/webflow/framer/wix/shopify/etcc

In a production setting we start with this and then run an agent that goes and makes it even better. ditto gives us the speed to show the customer their clone, the agent makes is scalable.

1
回复

@ridhwikvinod This is seriously impressive.

What I love most is that you chose a deterministic approach instead of throwing AI at the problem. There's something refreshing about solving an engineering challenge with engineering rather than relying on probabilistic outputs.

I'm curious, after watching developers use Ditto, what's the most unexpected workflow that's emerged? Did people use it primarily for migrations and redesigns, or have AI builders become your biggest audience because of the ability to start from a real product instead of a blank canvas?

Huge congratulations to the team on open sourcing this. I have a feeling the community is going to build some incredibly creative things with it. Wishing you an amazing launch! 🚀

0
回复

Hello!

That's a very interesting idea. How do you handle the issue of accessibility?

2
回复

@francisco_rocha_cortes We get as much as we can from the rendered DOM

1
回复

deterministic output is the part that sells me. every other site-to-code tool spits out something different each run, so you can never trust it in a real workflow. clean componentized next/vite plus open source is a genuinely useful combo, not just a demo.

2
回复

@alex_watson2110 Ya we needed something fast and cheap so we built it.

1
回复

the deterministic angle is the part that stands out to me, most "clone a site" tools I've tried lean on an AI guess and you get slightly different output every run. one thing I didn't see covered in the thread yet: if the source site gets updated after I've cloned it, is there a way to re-sync or diff against the new version, or is this meant to be a one-time snapshot you then own and diverge from?

2
回复

@galdayan Yea this is meant to be a one time snapshot to migrate customers to a new platform!

1
回复

Seriously amazing that it can extract all sorts of states including hover + focus. We could only dream of such a thing a year ago! And it's open-source?! wild!!

2
回复

@teddyni hell ya thanks for the support Teddy!

1
回复

the deterministic angle is what caught me. most of these tools hand the page to a model and you get something different every run. curious how it holds up on messy sites where the css is basically soup. also glad you shipped an mcp server, that's honestly how i'd use this

1
回复

@terminal_candy Glad you like the MCP!

0
回复

The deterministic output is such a smart call, makes the whole thing feel reliable instead of just another flaky scraper. Love that the components and tokens come through intact.

1
回复

@hayrettinyhac that's how we do, fast cheap reliable

0
回复
0
回复

HELL YES!!!

1
回复

@yahia_bakour3 hell yeaaaaa

0
回复

Pulled a few landing pages through it and the component breakdown was surprisingly clean, hover states and all. Way less guesswork than I expected for a from-scratch rebuild.

1
回复
#3
CartAI
The AI agent that handles checkout.
371
一句话介绍:CartAI 是一款让 AI 代理能在任何电商网站自动完成结账的 API,彻底解决了代理演示中“到付款就停”的痛点,无需商家集成。
SaaS Developer Tools Artificial Intelligence
AI代理 自动化结账 支付API 电商自动化 无商家集成 代理支付 PCI合规 购物车自动化 开发者工具 智能代理
用户评论摘要:用户普遍认可其“无需商家集成”的核心价值,主要担忧集中在:1)商家结账流程变化时的适应性;2)失败结账时的重试与重复扣款风险;3)退款、纠纷等售后事件的生命周期追踪。建议发布成功率指标、增加失败状态的事件流和幂等性保障。
AI 锐评

CartAI 切入了一个几乎被所有 Agent 演示刻意忽略的“最后一公里”——结账付款。它宣称通过合作而非逃避机器人检测、使用 Visa/Mastercard 的专门代理支付通道,解决了技术合规(PCI)和风控(被屏蔽)两大难题。其“无需商家集成”的定位极其精准,因为目前要求商家适配 AI 代理的协议标准(如 UCP)几乎不现实。从社区反馈看,用户的注意力集中在失败语义(幂等性、重复扣款)和售后追踪(退款、订单变更)上。这恰恰暴露了 CartAI 的薄弱环节:它本质上是在商家系统的外部“模拟”人类操作,不拥有系统记录权。一旦订单确认但自身收不到回执,或发生退货,其“确认即终止”的设计缺乏事后对账的闭环。除非其代理支付通道能从银行侧提供额外的对账单,否则“重复扣款”的担忧并非多余。此外,仅依赖客户自己收到的商家邮件来防止重试,在完全自主运行的 Agent 场景中并不可靠。长远看,CartAI 的价值在于为缺乏商务支持的大模型(如 Gemini)提供了一个可立刻落地的“商务能力附件”。但它构建的更像是一个基于风险对冲的“自动化支付桥”,而非一个耐用的“商务基础设施”。其成功最终取决于有多少商家愿意在风险监控上给这个“持证代理”开绿灯,以及公众对 Agent 代下单的信任度。

查看原始信息
CartAI
Checkout is the hard part, and most solutions clear it only where the merchant has integrated. CartAI completes checkout on any live merchant surface with no merchant-side work, cooperating with bot detection instead of evading it. One developer-first API, four products: Catalog (search and live pricing across merchants), Checkouts (clear and track orders to confirmation), Payments (PCI off your stack via Visa and Mastercard agent rails), Monetization (commissions on every agent sale).

Hey PH 👋
I'm Manil, founder of CartAI. We have been building CartAI for more than a year now and today we're shipping the developer release.
The problem we kept running into: every AI agent demo ends right before the part that matters. Browser automation can navigate the web but can't pay. Payment APIs can move money but can't navigate. Your agent gets to the checkout button — and stops.

CartAI is the API that closes that gap. Integrate, point an agent at any web property, and it completes the transaction — checkouts, subscriptions, invoices, orders.

It's four products covering the full transaction lifecycle, with checkout as the wedge everything sits on:

Catalogfind the product: search across merchants, live variants + pricing, and checkout estimates before a cart even exists
Checkoutclear the order: complete on the live merchant surface, tracked to confirmation with normalized orders + webhooks
Paymentsmove the money: hosted, PCI-compliant sessions on Visa Intelligent Commerce and Mastercard Agent Pay
Monetizationshare the upside: affiliate commission captured automatically across 70,000+ brands, attribution preserved through to the sale

Three ways to use it:
Automate — your agent runs the full flow autonomously
Embed — drop our agent into your app workflow
Enable — add checkouts to surfaces that never had one. Make any surface commerce enabled.

What's actually hard here isn't the navigation. It's PCI-compliant payment handling, cooperative bot-mitigation so you clear checkout instead of getting blocked (we cooperate with Cloudflare/HUMAN/Akamai/Fingerprint via Web Bot Auth and signed agent identity using Skyfire — we don't evade them), and workflows that always terminate in a known transactional state. That's the part we built.

CartAI clears real orders on production merchants: start to finish, from one API call. There's also an open-source MCP server if you'd rather drop it straight into Claude, Cursor, or your own agent.

Would love your feedback, especially from anyone building agents that need to actually transact. What would you point CartAI at first?

Docs → docs.cartai.ai

16
回复

@manil_uppal Since CartAI operates across many different merchant experiences, how does the API improve over time when it encounters new checkout flows or unexpected edge cases? Is there a shared learning layer where successful transaction patterns make future checkouts more reliable across merchants, or is each flow handled independently? That network effect feels like it could become a significant long-term advantage.

0
回复

@manil_uppal Really interesting approach! How does CartAI handle situations where a merchant's checkout flow changes unexpectedly? Is the system able to adapt automatically, or does it require manual updates?

0
回复

@manil_uppal Really interesting! How does CartAI handle unexpected changes in a merchant's checkout flow? Can the agent adapt automatically, or does it require manual updates?

0
回复
Check out in any merchant with zero merchant-side integration is the whole ballgame. If that holds up, this is big.
4
回复

@anusuya_bhuyan 

Exactly—you've identified the crux of it.

Enabling checkout without merchant-side integrations is only half the challenge. The other half is ensuring every transaction is secure, authentic, and trusted.

CartAI handles that complexity by working alongside the merchant's existing trust and security ecosystem—including solutions from our partners like Akamai, Cloudflare, Skyfire and HUMAN— so AI agents can complete purchases while respecting the same protection mechanisms merchants already rely on.

We believe that's what will make Agentic Commerce truly scalable. 🚀

0
回复

I like that checkout works without merchant integration. How do you handle unexpected checkout changes on different stores? Publishing success rates by merchant could build more confidence.

3
回复

@alheri_murya 

Thanks! That's a great question.

Handling dynamic checkout experiences is one of the core challenges we're solving. CartAI continuously adapts to changes in merchant checkout flows through our automation and validation layer, allowing us to remain resilient even as storefronts evolve.

I completely agree on merchant-level success metrics. As the platform grows, we'd love to publish coverage and success rates to give developers better visibility and confidence.

Thanks for the thoughtful feedback!

2
回复

Me value the simple API approach for developers. What monitoring tools are available after checkout starts? Real time alerts could help teams respond faster.

2
回复

@hana_salazars 

Thanks!
Excellent developer experience has been one of our core priorities from day one.


Once a checkout is initiated, developers can track its lifecycle through our Webhooks, receiving real-time events for important state changes and integrating those into their own monitoring or automation workflows.


We also provide a portal where teams can monitor orders, inspect checkout status, and troubleshoot issues when needed.


We believe every team has different operational needs, so rather than locking developers into a predefined alerting system, CartAI gives them the building blocks to create monitoring and notification workflows that fit their own stack.

Native email and SMS alerting are already on our roadmap and will be rolling out soon, making it even easier for teams to stay informed without building custom notification workflows.

Really appreciate the suggestion—feedback like this helps us prioritize what developers need most. 🚀

1
回复

It would help a lot if the Catalog API exposed a confidence score or freshness timestamp with each price result, so we can decide when to refetch before sending the user to checkout. Right now we have to guess whether a cached price is still good enough to trust.

2
回复

@esmasagtekin 

Thanks for the thoughtful suggestion!

Our approach is slightly different from a traditional Catalog API.
In addition to serving product data, CartAI's Checkout Estimates API gives developers an understanding of the expected checkout outcome— such as pricing and other checkout signals — before initiating the actual purchase.


During checkout, we also validate the latest merchant data to ensure the transaction uses the most up-to-date information available.

That said, exposing explicit freshness metadata or confidence scores is an interesting idea, and we'll definitely consider it as we continue to evolve the platform.

Thanks for the valuable feedback! Love it!

0
回复

The 'checkout as the wedge, no merchant-side integration' framing is the interesting part for me, since that's the piece every agent demo skips. The first thing I'd test is failure semantics on Checkout: if the agent clears a cart but the confirmation webhook never lands — merchant surface changes mid-flow, or the session times out — does the API give me an idempotency key so a retry can't double-charge, and do I get a clean failed state with no partial charge? That reconciliation is where agent payments usually fall apart in support.

0
回复

The buyer-side and merchant-side approaches to this are converging fast, and the tradeoff you picked is the interesting one. Protocol-based integration only works where the merchant has already done the work, which is a thin slice of the web today, so completing on the live surface covers far more ground immediately.

The part I would watch is order state after confirmation. Refunds, partial cancellations and price changes get messy to normalize when you do not own the merchant's system of record.

Do the webhooks cover post-purchase events, or is confirmation the end of the tracked lifecycle?

0
回复

@paul_crinigan Today order confirmation is the terminal state. Once an order is placed the communication becomes direct between the merchant and the consumer. We do not insert ourselves in the middle of that conversation with email relay type techniques etc. We will see how this evolves over time

The protocol based approaches support post purchase messages a bit better(UCP for example) but there are no examples of UCP being pointed to 3rd party demand agents like @CartAI today so as you say, adoption is low at this point but we are hoping it picks up.

0
回复

completing checkout on any merchant surface with zero integration work is the hard part everyone else punts on, so props for actually shipping that instead of just a demo video. curious how disputes/refunds work in practice - if the agent picks the wrong variant or size because the merchant's page was ambiguous, does the normalized order data make it easy for the end user to catch that before it ships, or does it only surface after the fact like a normal order confirmation email would?

0
回复

@omri_ben_shoham1 So we do have a cart verification step where we have our QA agent check the work of our order placing agent to make sure the right variants, quantities etc were added to cart.

Also our order input api asks for the exact variant name and selection and we expect that to match what is on the website or the agent will not execute. The agent is trained not to "guess" or find "best match". It has to be a 100% match. That is also the reason we also have a catalog api service where we show the products and the exact variants available for sale with the merchants. We have made the path intentionally narrow to make sure there are no errors during execution.

0
回复

I'm curious how merchants feel about this. Do they see agent-driven checkouts as another sales channel, or do some still treat them as bots they'd rather block?

0
回复

@reda_roqai_chaoui Both, honestly, and it's a spectrum right now.

Some merchants still block all agents outright. Others run tiers, different levels of access depending on whether you're approved through their CDN or bot provider. So the same agent can be blocked on one site and waved through on another, based purely on who vouched for it at the edge.

Ultimately this only gets better. As the big surfaces like Google Gemini normalize buying on non merchant native surfaces, agent driven checkout stops looking like an anomaly to defend against and starts looking like a channel to serve.

1
回复

the cooperating with bot detection instead of evading it line is the most interesting part of this. i sell software online and my first thought was whether i'd even know an agent checked out on my site. what does the merchant actually see when cartai comes through?

0
回复

@terminal_candy At the security layer, you see us. If you run Cloudflare, HUMAN, Akamai, or Fingerprint, CartAI shows up as a declared, signed agent via Web Bot Auth. A known identity you can allow, challenge, or block. We arrive labeled, and we're listed in Cloudflare's bot directory, so it isn't guesswork.

In your admin, less. The order lands looking like a normal ecommerce order with the real consumer's email and phone, because order records don't yet have a field for "an agent placed this." So you'd know from your bot detection, but not from the order itself today.

0
回复

The agent-rails angle is genuinely interesting, especially how you're framing cooperation with bot detection as a feature rather than a workaround. One thing that would help me trust this more: add a public status page or webhook for failed checkouts, like when a merchant site changes its DOM and Catalog pricing or selection silently breaks. Right now I'd have no visibility into why an order stalled, which makes debugging painful. A simple event stream for "checkout abandoned at step X" would close that loop nicely.

0
回复

Hello @ece1080217 great points.

  1. We do have a webhook that shows Failed when a checkout could not be completed. https://docs.cartai.ai/reference/webhooks

  2. Also in our portal dashboard, we show the sequence of events in the checkout workflow and if there is a failure we show which step failed

  3. We also capture screenshots as the agent is checking out so you can see a full gif of the checkout for audit and troubleshooting purposes if needed. See some examples here https://www.cartai.ai/see-it-in-action

0
回复

Right, first-party contact means the shopper still gets the merchant's email even when your side logged a fail. What I keep circling is the fully-autonomous case where nobody reads that email: CartAI's own record stays 'failed' while the money actually moved, so the calling agent could retry and pay twice. Do you expose a reconciliation signal, a webhook or a status I can poll later, so the agent can catch a settled-after-failed order before it re-runs checkout?

0
回复

@dipankar_sarkar 

That's a really good question.

One important difference with Visa Intelligent Commerce and Mastercard Agent Pay is that the customer grants spending authorization before checkout, and that grant is scoped to a specific order.

For example, if the customer authorizes an agent to purchase a shampoo for up to $100, and the checkout succeeds for $99, that grant has been consumed. Even if the agent retries because CartAI didn't receive the merchant's confirmation, it can't simply charge the customer a second time using the same grant.

If a developer were to initiate a brand-new grant, the customer would need to approve it again. In practice, if they've already received the merchant's confirmation or know the purchase went through, they're unlikely to authorize a duplicate purchase.

You're absolutely right that the lack of visibility after a successful-but-unconfirmed checkout is an edge case we're actively working on with our payment partners. But the underlying agentic payment guardrails are designed to prevent an agent from repeatedly charging a customer for the same purchase, even if a retry occurs.

0
回复

Interesting because that is a pretty hard problem to solve without being considered spam.

I am curious how it'd work in Turkey as all of the payments are 2FA message guarded.

0
回复

@fiyu fair point, 2FA guarded markets are a real constraint. If a card payment forces a live OTP or bank verification challenge, the pure browser method can't clear it on its own. That's where the payment needs to flow through rails built for it, like direct merchant or PSP integration. On the standards side, Google's AP2 is the one aimed here. It handles payment authorization with signed mandates and a user approval step for the step up. UCP (Universal Commerce Protocol) standardizes the commerce side around it. Both are new and adoption is still limited.

0
回复

Congrats on the launch! Checkout is the part most agent demos quietly skip, so tackling it head-on is welcome. The interesting bit for me is the confirmation loop (keeping the human in control of what actually gets charged). Good luck today!

0
回复

@fujibee much appreciated 🙏

0
回复

Fail-closed is the right default, and I like that you'd rather under-claim than over-claim a success. The case that still bit us was the inverse: the charge actually settled on the merchant side while our side had already marked it failed, so the user retried and paid twice. Does the per-agent token give you a way to reconcile a 'we called it failed but it settled' mismatch after the fact, or does that fall to a manual dispute?

0
回复

@dipankar_sarkar 

That's a fair concern. In that scenario, the customer would already have received the merchant's order confirmation because we use the customer's own contact and shipping information as first-party data during checkout.

So while CartAI may have marked the checkout as failed due to the missing confirmation, the customer would still be notified directly by the merchant if the order was actually created. That significantly reduces the likelihood of an unnecessary retry leading to a duplicate purchase.

That said, improving reconciliation for these rare edge cases is something we're actively working on with our payment partners.

0
回复

"Cooperating with bot detection instead of evading it" is the smart call — the evasion arms race is unwinnable long-term. The part I'd worry about: an agent completing checkout on any live surface with no merchant integration is one wrong parse away from buying the wrong thing or the wrong quantity. What's the confirmation/guardrail before it actually pays? That's the line between useful and scary. Congrats on the launch.

0
回复

@david_marko 

We completely agree, that's exactly why we don't believe AI should make unchecked purchasing decisions. CartAI requires explicit consumer authorization for agentic payment, and merchants remain in control of how agent checkouts are handled.
Trust isn't just about getting through checkout; it's about making every purchase intentional, auditable, and aligned with merchant policies.

0
回复

The charged-but-no-confirmation case Waqas raised is the one I'd lose sleep over building this. When we ran agents against checkout flows we didn't control, the timeout was never clean: the backend committed but the confirmation render died, and our side couldn't tell 'it failed' from 'it succeeded and we didn't hear back'. Since the Visa and Mastercard tokens are scoped per agent, can you reconcile against the payment network's authorization record instead of the merchant's confirmation page, so a dropped render doesn't strand a real charge?

0
回复

@dipankar_sarkar 

That's a great question—and trust us, we've already spent plenty of sleepless nights thinking about this exact scenario. It's one of the reasons this has become a recurring discussion in almost every conversation we have with our payment providers.

Today, if CartAI doesn't receive a definitive confirmation from the merchant, we intentionally treat the checkout as failed rather than assuming success. From the developer's perspective, there are only two possibilities:

  1. The merchant successfully created the order, but the confirmation never made it back to CartAI.

  2. The merchant never completed the order.

In the first case, the customer is still the contact on the order, so they'll receive the merchant's normal order confirmation and post-purchase communications even though CartAI couldn't verify the final state.

Your suggestion about reconciling against the payment network's authorization record is exactly the kind of capability we're exploring with our payment partners.
We don't consider this problem fully solved yet, and we're actively working on making this edge case much more transparent for developers.

0
回复
Congrats on the launch. That’s really strong use case. How do you guys manage a failed checkouts or partial checkouts so that the agent can retry safely without getting duplicate charges or orders?
0
回复

@nischaydhiman 

Thanks for the question!
Before the checkout begins, two important things happen:

  • The customer provides explicit confirmation before the agent starts the checkout.

  • That confirmation includes a spending grant, which specifies the maximum amount the agent is authorized to spend for that particular order.

Payment authorization happens later in the checkout flow. This means that if the checkout fails before the order is successfully placed, the customer's card is not charged. Charges are only captured when the merchant successfully accepts the order.

On Partial Checkouts

  • CartAI supports placing orders containing multiple SKUs, whether they belong to the same merchant or different merchants.

  • Our Create Checkout API includes a flag that lets developers choose whether all SKUs must be successfully purchased or whether a partial order is acceptable.

If the checkout is configured to require all SKUs and one or more items are unavailable, the entire order is aborted, and no charges are made to the customer's card.

0
回复

The no-merchant-integration approach is impressive! How does CartAI handle last-minute price, shipping, or product changes before payment—and when does it ask for explicit user approval?

0
回复

@ram_sarjal 

Great question. We've intentionally split the flow into three distinct stages:

  1. Product Discovery

  2. Cost Estimation

  3. Checkout

During the Cost Estimation stage, we calculate the expected total, including the product price, shipping, taxes, and other applicable charges based on the user's address.
In most cases, these values remain stable through checkout.

Minor variations can still occur—particularly in shipping costs during peak periods. That's why we recommend developers authorize the agent with a 5–10% spending buffer to accommodate small changes.

If the final amount exceeds the approved grant, the checkout simply fails. The agent cannot spend beyond the user's authorized limit, ensuring spending always remains within the consumer's explicit approval.

0
回复

I like the idea, but checkout is one place where I’d still want a clear approval step. What happens if the price changes, an item is substituted, or the final total is higher than expected? Does CartAI pause and ask before placing the order?

0
回复

@liana_preston We do render an expected total based on item cost, shipping costs and tax using our checkout estimates api https://www.cartai.ai/product/catalog#estimates

When the agent goes to place the order and if the price encountered is higher, we have an upper bound setting that our platform customers can set around how much drift they can tolerate. If set to 0% and if the actual is even 1 cent higher, agent will not complete the order.

0
回复

@liana_preston 

That's a great point, Liana. We actually have a feature in our backlog called "Confirmation Required Purchase." If this is something more developers are looking for, we'd be happy to prioritize it.

For now, our model is what we call "Intervention-Free Purchase." It works by ensuring:

  • The customer gives explicit confirmation before the agent begins checkout.

  • That confirmation includes a spending grant with a maximum amount the agent is authorized to spend for that specific order.

  • If the final checkout amount exceeds that approved limit, the transaction simply fails—the agent cannot overspend.

This gives developers a predictable, auditable flow while ensuring every purchase stays within the consumer's approved spending limit.

0
回复

Interesting. What framework and apps do you support today?

0
回复

@chilarai Thanks!

Today, we have our capabilities exposed through APIs, and our MCP toolset.

We also made it developer-first. Sign up for free on our portal, get an API key, and test every product on the platform yourself. No sales call to start. Hit the portal or API and create your first checkout today.

0
回复

Nice Lanuch! How do you handle the nasty middle state where the card gets charged but the confirmation page times out or never comes back? Curious how you guarantee idempotency and reconciliation there, and how the final state a caller receives distinguishes "order actually placed" from "charged but the merchant has no record of it.

0
回复

@iamyuhanliu 

Great question.
The key thing is that CartAI never charges the customer's card itself.
The payment is still processed by the merchant through their existing payment provider, just as it is today.

CartAI uses Visa Intelligent Commerce (VIC) and Mastercard Agent Pay to obtain permission to spend using agentic payment tokens (agentic PANs).

The merchant processes these tokens like a normal payment credential, while the Card Network handles the mapping to the customer's actual card.


If a confirmation page times out after payment, that's not unique to Agentic Commerce, it's the same scenario merchants already handle today. Their payment gateway, order system, and the card networks remain the source of truth for whether a payment was authorized and whether an order was created.

CartAI simply works within that existing ecosystem rather than introducing a new payment flow or reconciliation model.

1
回复

Congrats on the launch, excited to try this out

0
回复

Thanks so much, @mark_okiki 

We really appreciate the support and hope you enjoy trying it out.

0
回复

Hi @CartAI Team,

First of all, Many many congratulations for such a great launch, 🎉

Really like the approach of working with bot detection instead of trying to sneak past it. Getting listed in Cloudflare’s verified bot directory is not something most teams can pull off.

One question though. Since you’re checking out on live merchant sites, what happens if a payment goes through but the site glitches before the confirmation shows up? Does the system know not to retry and charge the customer twice? Curious how you handle those messy edge cases, because that feels like the hardest part of making this reliable.

Rooting for you guys, the idea makes a lot of sense. Great job 👏

0
回复

@waqas_baloch4 

Great question !!

The agentic order is considered successful once the merchant processes and confirms it. Any failure during this period will fail the entire transaction.
We do not retry if the merchant site glitches before the confirmation and fails to display an order confirmation page.

1
回复

Checkout is exactly where agent demos need hard state boundaries. The valuable bit is not just completing a flow; it is knowing when to stop, what was authorized, what changed, and how a human can audit it later.

0
回复

@krekeltronics 

Completely agree.
Checkout is where AI has to stop being probabilistic and become deterministic.
Every state transition needs to be explicit, auditable, and backed by the merchant's existing commerce and payment systems.
That's exactly why CartAI is built as infrastructure.
We don't invent new transaction semantics; we work within the merchant's existing authorization, payment, and order lifecycle.
By working within the merchant's existing transaction boundaries, we have build a framework that provides strong auditability, clear state transitions, and complete operational visibility.

0
回复

Interesting, how are you handling payment failures ? Also, how do you handle websites blocking such AI bots (ex: cloudflare turnstile) ?

0
回复

@vinitvr 
Bot mitigation relies on establishing verifiable agent identity using industry-standard protocols and implementations, enabling merchant websites to identify CartAI as a trusted agent and whitelist our traffic.

In case of payment authorization failure, the entire transaction is aborted. Additionally, the product provides robust handling for various other payment-related failure scenarios.

0
回复

Congratulations on the launch, excited to try it out!

0
回复

Thank you! It means a lot.@ronakagarwal3434 

Looking forward to hearing what you think once you've tried it !!.

0
回复

Congrats on the launch @maniluppal! Solving the execution wedge with cooperative Web Bot Auth rather than scraper evasion is definitely the right long-term move.

On the payments side, when an agent executes a purchase via Visa Intelligent Commerce or Mastercard Agent Pay, how is chargeback liability and dispute resolution structured between the agent developer, CartAI, and the issuing bank? If an end-user claims their agent made an unauthorized purchase within its spending grant, does the tokenized agent identity serve as cryptographic proof of authorization during a dispute?

0
回复

@franz_briones Great question, and it's the part of agentic payments people still don't talk about enough.

The main thing agent identity changes is the evidence around a transaction. On both Visa Intelligent Commerce and Mastercard Agent Pay, the purchase carries a token scoped to the agent plus a signed record of the user's mandate: the spend grant, its limits, and proof the agent acted inside them. That works a lot like a 3DS authentication signal. A properly authenticated agent transaction that stayed inside its grant is designed to move fraud liability off the merchant, the same way authenticated card not present does today.

Where @CartAI sits: we're infrastructure, not the issuer or the merchant of record, and we don't adjudicate disputes. What we provide is the cryptographic authorization trail. Who consented, what the grant was, and proof that execution stayed within it. That gives the developer and merchant real evidence to represent a dispute.

So on your exact case, an agent buys something inside its grant and the user later calls it unauthorized: the signed agent identity plus the mandate is strong evidence the purchase was authorized, and materially better than anything in current card not present flows. It strengthens the issuer's picture rather than replacing their decision.

0
回复

Coming at this from the merchant side, we build support tooling there, so my head goes straight to what happens after a wrong order clears.

The thread's covered the price/stock case at checkout. The one I don't see: the order completes fine, and then it's wrong. Wrong variant, wrong address, size M when the person wanted L. On a human order the merchant's support desk just emails the buyer, confirms, sends a return label. The whole post-purchase flow assumes a human on the other end who can answer "which one did you mean".

When the buyer was an agent, who does the merchant's support team reach? The agent is an API, not an inbox. The end user is behind someone else's app. So a returns or "wrong item" ticket lands on a merchant with nobody reachable to resolve it.

Does CartAI keep a channel open back to the buyer after confirmation, or does the agent vanish at the confirmation webhook and leave the merchant holding a ticket they can't close…

0
回复

@jernej_jan_kocica Great question. CartAI places the order with no loss of fidelity between the consumer and the merchant. So the contact information(email, phone) etc that goes with the order is of the end consumer. Any post purchase or customer service activities including customer reach back etc will follow exactly your normal processes as the order we place will look no different than any other E-com orders that you get.

0
回复

the smart bit is cooperating with bot detection instead of evading it, everyone else plays cat and mouse and eventually loses. the thing i'd worry about is merchant checkout flows changing under you, but if you've genuinely solved that on any live surface, that's a real moat.

0
回复

@alex_watson2110 


Thanks Alex! That's exactly the philosophy we're betting on.

We don't believe Agentic Commerce becomes mainstream by trying to outsmart merchant security. It becomes mainstream by earning merchant trust and working within the existing security ecosystem.

On checkout changes, that's definitely one of the hardest engineering problems. CartAI continuously validates the checkout state, adapts to dynamic changes, and uses self-healing orchestration to recover from transient failures whenever possible, rather than restarting the entire journey.

Since checkout is all we do, we're constantly learning from real-world checkout flows and improving our engine to adapt as merchants evolve their experiences.

We believe that adaptability, combined with a cooperative trust model, is what creates a durable moat for Agentic Commerce.

0
回复
#4
Rerun
The easiest way to build AI agents for all your tasks
308
一句话介绍:
Rerun是一个让用户无需编码即可构建24/7全天候运行AI代理的平台,通过实时代理步骤可视化、敏感操作审批暂停和私有服务器部署,解决了现有AI代理“黑箱”决策、不可控且难以调试的核心痛点。
Tech Marketing automation
用户评论摘要:
用户高度关注失败恢复(是否可从故障步骤续跑)、审批自定义(如按金额阈值自动放行)、审计回放(时间轴浏览与日志导出)及长期学习能力(是否从历史决策中优化行为)。建议提升工具查找速度、支持分步重放与自定义敏感规则。
AI 锐评


Rerun精准切中了当前AI代理市场最核心的信任赤字——黑箱恐惧。它在“无代码构建”的便捷性与“可观测可干预”的控制力之间找到了一个稀有的平衡点,这比单纯堆砌自动化功能更有价值。

产品的生命力在于对“失败场景”的宣战。虽然“重放与回滚”机制仍是待解难题,但团队坦诚地承认限制,反而建立了可贵的信任资产。其“私有服务器”架构和“关键步骤审批”功能,是赢得企业级客户的最短路径——安全与可控永远是商业化的门票。

真正的挑战在于用户评论中反复提及的“学习循环”。一个仅靠预设规则运行的代理,长期看不过是昂贵的宏指令。如果Rerun能将用户每一次的“批准/拒绝”决策转化为模型微调的反馈信号,让代理在保持人控安全网的同时,自主优化执行策略,它将从“更好用的工具”进化为“自进化的数字员工”。此外,对审计和回放的刚性需求,暴露出当下AI代理尚未解决的责任链困境——Rerun必须将这些“事后追溯”功能做到极致,才能让企业在面对错误决策时有据可查,而非把责任推给“AI的灵光乍现”。值得警惕的是,用户对日志导出和回放的朴素要求,可能预示着一个比“自动化”更大的市场:AI代理的合规与审计基础服务。

查看原始信息
Rerun
Most AI agents are black boxes. Rerun isn't. Build no-code agents that run 24/7, chasing invoices, qualifying leads, clearing your inbox, and watch every step in real time. They pause for approval before anything sensitive. Each workspace gets its own private server. Start with the included model, or connect your own API key or subscription
Hey Product Hunt 👋 I'm Clément, founder of Rerun. We built Rerun because every "AI agent" tool felt like a black box. You fire it off and just hope it did the right thing. We wanted agents you could actually watch work, step by step, that pause before doing anything sensitive. Rerun lets you build no-code AI agents that run 24/7 (chasing invoices, qualifying leads, clearing your inbox), connected to the tools you already use. Live dashboards show every run, token, and decision. Each workspace gets its own private server, no shared tenancy. We'll be here all day. We'd genuinely love your feedback, questions, and ideas on what to automate next.
10
回复

@clement_janssens Curious about the failure mode more than the happy path. When we automated recurring ops work internally the problem was never building the thing, it was the agent quietly doing it wrong for three weeks before anyone noticed. What does Rerun surface when a run half-succeeds, and can I set a rule like "stop and ask me" on specific steps?

1
回复

@clement_janssens I'm curious about one capability:

As users run agents over weeks or months, does Rerun learn from approval decisions? For example, if I consistently approve certain actions and reject others, can the platform adapt its behavior or confidence thresholds over time while still keeping the human in control? That learning loop feels like it could become a major long-term differentiator.

0
回复

@clement_janssens Love the transparency-first approach! How does Rerun handle long-running agents that improve over time? Do they learn from previous executions, or is each run treated independently?

0
回复

Wow, this looks extremely powerful! Definitely have to check this out!

2
回复

@bennyqp Thanks Benny, we're super proud of the product, can't wait to get your feedback once you've tried it!

0
回复

"Watch every step in real time" is the part I care about most, because the black-box complaint is real. When a step fails halfway through a run, can you fix it and resume from that step, or does the whole agent re-run from the top? Resuming from the failed step is what makes long runs actually usable, and it is the hardest part to get right.

1
回复

@cmumulle This is honestly one of the things we're always working on

The agent is super independent, if a step fails, it'll find a way to make it work one way or another (like switching APIs or something)

But if it gets stuck, you’ll get notified and can fix it with the agent

0
回复

IMO the straight 'there's no rollback' answer is worth more than a polished one, most tools in this lane would have claimed replay works and hoped nobody tested it. QQ: when the agent cancels changes on request, is it walking back through the tool calls it made (voiding the invoice it created, deleting the record it wrote)? Congrats on the launch!

1
回复

@artstavenka1 Exactly, the agent knows what they’ve done which lets them undo changes (if possible)

That's why it's better for the agent to ask for approval BEFORE messing things up

1
回复
The part I'd actually care about is it stopping to ask before anything sensitive. Letting an agent send emails or move money on its own is the scary part, so having it wait for a yes first is what makes it usable. Congrats on the launch!
1
回复

@etiennegarciaThank you, man! It’s definitely possible. You can also create your own personalized widget that aggregates all the tasks requiring your approval in one place.

1
回复

@etiennegarcia Exactly, especially during the first automation runs. Even if the skill is accurate, it's not perfect on the first try. Making sure you’ve got control before critical phases is key for safety and for iterating to get an agent that truly fits the use case

1
回复

The hard part with always-on agents isn't watching them work, it's what happens after an approval. If a human approves step 4 and step 7 goes wrong, can you replay from the pause point with corrected context, or does the whole run start over?

1
回复

@aidan_codefox Totally agree.

Since the flows are unique, there's no way to "rollback."

However, the agent is smart enough to cancel changes if the user asks.

For example, with Rerun, any agent handling important processes has two key steps:

First, every critical or risky step requires user approval explaining exactly what it’s going to do. This ensures it hasn’t gone off track before taking action

Second, the agent notifies users once the step is done to confirm everything went smoothly.

If not, you can make it redo or cancel the changes

1
回复

one feedback, finding tools is slow, which i feel is not supposed to be, i feel you can fix that

1
回复

@daniel_oyetunde Hey Daniel, we’ll check it out right away, thanks for reporting it!

0
回复

Real-time step logs, approval gates, and a private server per workspace make always-on agents much easier to trust. Could teams set per-agent token budgets or automatic stop thresholds so a looping workflow cannot burn through connected model credits?

0
回复

the live dashboard showing every run/token/decision is the part I'd actually use daily, most agent tools just give you a final log line and nothing in between. how configurable is the 'pause for approval on anything sensitive' rule - is that a fixed list of action types you ship with, or can you define per-workflow what counts as sensitive for your own use case?

0
回复

I feel like debugging AI agents is becoming as important as building them. Have you found that users spend more time creating agents or understanding why they failed?

0
回复

the pause for approval before anything sensitive part is the right call. i live in coding agents all day and the failure mode is never the work, it's the thing it does confidently that you didn't want. how do you decide what counts as sensitive, is that configurable per agent?

0
回复

Real-time visibility is a huge plus here, love that. One thing I'd love is the ability to set custom approval thresholds per agent, so something like "auto-approve invoices under $500 but ping me for anything above that." Would save a lot of clicks on the routine stuff while still keeping the safety net for the bigger decisions.

0
回复

One thing that would really help me trust the agents even more is a simple replay button for past runs, so I can scrub through and see exactly what decisions were made when something went sideways. Live monitoring is great, but most of my debugging happens after the fact and right now I'd have to dig through logs to piece it together. A timeline-style playback view per run would make audits and post-mortems so much easier.

0
回复

honestly the real-time step view is genuinely cool, but i think a simple way to export the full run log as csv or json would be super useful for auditing later on

0
回复

honestly the live step viewer sounds super useful for debugging agents. one thing i'd love is the ability to rewind and replay a specific portion of an agent run instead of just watching it forward, especially when something weird happens and you wanna trace back what triggered it

0
回复

@azizileskaxtxg Rerun was really born from this simple realization. We’ve got agents, but we don’t fully get their journey. So, we built tools to monitor that. And honestly, it’s a joy to use every day

0
回复

The live step viewing is genuinely useful, way more transparent than other agent builders I've tried. One thing that would seal the deal for me is a simple test mode where I can replay a past run with tweaked inputs, so I can fine tune prompts without spinning up real tasks each time.

0
回复

@kerimkkavuggeq Oh good catch, hadn’t thought of that yet! Adding it to the roadmap

0
回复

Love that I can actually see what the agent is doing in real time, that's the missing piece in most no-code tools. One idea: add a way to export the full execution log as a shareable link so I can send a client a recap of what the agent did without screen recording. Would save a ton of back and forth.

0
回复

@tubauwvo Thanks for your feedback, it’s already possible with the “Dashboard” option

You can drag’n drop pre-made widgets for monitoring, analysis, approval, etc.

Feel free to reach out to support for guidance

0
回复

Love how every workspace gets its own private server, that kind of isolation is honestly rare for no-code tools and shows you clearly thought through real business use cases, not just the demo.

0
回复

@ozanzbaykulvdq Yep, every workspace is on a private server for full control and total security

Excited to see what you're gonna build in the next few weeks

0
回复

I follow the founders on X and they’re both very inspiring and creative

I already created my account on their platform and it’s a gem

Go go guys

0
回复

@fberrez1 Thanks Florent for your support! Can’t wait to see what you’ll build on Rerun in the coming weeks! Let’s keep in touch

0
回复

Whats the top 3 most popular use cases for this system?

0
回复

@conduit_design Marketing automation (posting & scheduling), autopilot SEO, lead magnets creation, lead scraping, managing meta ads campaigns... way too many use cases to just have 3

0
回复

honestly the step-by-step visibility is what sold me, like finally i can see what the agent is actually doing instead of just hoping it works. the approval pause before sensitive stuff feels solid too

0
回复

@kaanzclu You totally got how Rerun works, and it's such a joy to use! Can't wait to see what you've built with it

0
回复

the live step-by-step view is genuinely useful, you can actually see what the agent is doing instead of guessing. pausing for approval before sensitive actions feels like the right default.

0
回复

One thing I'd love is the ability to set custom rules for when an agent should pause for approval, beyond just "sensitive" actions. For example, letting me define thresholds like "pause if this invoice is over $500" or "ask before sending to a contact I haven't emailed before" would make the approval flow way more useful for my actual workflows.

0
回复

@celalasravxqhf You totally got how Rerun works, and honestly, it’s one of the most used features, even by us at Rerun

Can’t wait to see what you’re gonna build!

0
回复

The live step-by-step view is genuinely useful, but it would be great to have a simple way to roll back or replay a specific run if something goes wrong mid-task. Right now if an agent gets stuck or makes a bad call early on, I have to start the whole thing over. A "restart from step 3" button or a rewind slider would save a lot of time when debugging automations.

0
回复

@tahsin7mio Thanks for your feedback, you're totally right, we'll put that at the top of the roadmap

0
回复

honestly the live step-by-step view is kind of addictive, like watching a little worker do my tasks in real time. the approval pause thing also feels genuinely thoughtful rather than just bolted on.

0
回复

@cemraca51348 With Rerun, we really want work in the future to feel like playing a video game

You set up agents, watch them work, approve, tweak, and validate

Future of work

0
回复

The pause-for-approval step before sensitive actions is such a smart call, honestly. A lot of no-code agent builders skip that and just let things rip, so it feels like the team actually thought through what people would be nervous about handing over to automation.

0
回复

@selim2fu3 Exactly Selim, thanks for your feedback

0
回复

"Watch every step in real time" and "pause for approval" are the two things keeping me from trusting AI agents with actual business tasks. Most tools just give you a success/fail flag and hope you don't notice what happened in between.

0
回复

@ringo_td5 Totally, that’s Rerun’s strength, being able to really control your agents, monitor them, and step in if they go off track

Thanks for the feedback

0
回复

Love the "pause before sensitive actions" approach. Curious how it plays with the 24/7 long-running tasks though, when an agent pauses for human confirmation, does it just hang and hold resources, or is there a timeout / fallback if no one responds in time?

0
回复

@iamyuhanliu For now, the agent's stuck waiting for the user's reply

0
回复

Nice approach. Does the approval gate still work if an agent runs while you're asleep?

0
回复

@dhiraj_patel5 Exactly, the agent notifies you and in the morning, when you wake up, you just reply!

0
回复

Rerun sounds aimed at a pretty wide set of use cases, from Marketing automation to Engineering & Development and Productivity. How do you think about helping new users choose the right starting point for their first AI agent? Is the product more template-driven by category, or does it guide people from a task description into an agent setup?

0
回复

@ivory_xuxuxu One of the challenges with Rerun is that it's super flexible, so it can work for tons of use cases

We've put a lot of effort into a multi-step onboarding so you can run a 100% personalized agent without needing any tech skills

Feel free to try it out, would love to hear your thoughts

0
回复
#5
Routine AI
Control work with your voice. The Siri for work.
225
一句话介绍:Routine AI是一款利用语音指令控制工作任务、日历、笔记和项目的AI生产力工具,解决了用户在多个应用间切换、手动操作效率低下的痛点,让办公像与真人助理对话一样自然高效。
Productivity Artificial Intelligence Audio
AI语音助手 智能办公 任务管理 日历管理 笔记与项目管理 工作自动化 跨应用同步 个人助理 效率工具 Product Hunt
用户评论摘要:用户普遍赞赏其整合日历、任务和笔记的功能设计,并认可语音操作的实际价值。主要建议包括:增加仪表板和自定义工作流灵活性;希望AI能根据截止日期主动建议优先级;完善日历的甘特图或时间线视图;关注语音误操作时的确认机制与纠错方式。
AI 锐评

Routine AI在语音助手泛滥的当下,难得地找到了一个“有上下文”的切入点。它没有试图做一个无所不知的通用AI,而是聪明地扎根于用户已经存储的工作数据——日程、任务、笔记——让语音命令不再是盲人摸象。这才是真正的价值:不是让你更快速地打字,而是让你不用再打字。创始人Julien在评论互动中强调“语音命令只是开始”,后续还有AI代理和自动化,方向正确,但挑战不容小觑。评论中反复出现的“误操作确认”“纠错方式”“任务链失败处理”问题,恰恰戳中了当前大多数语音助手的致命伤——意图解析的鲁棒性。用户能够忍受打字缓慢,但绝无法容忍安排好的会议被一句话毁掉。Routine目前用“确认清单”的方式规避了这一风险,但体验是否足够流畅仍是未知数。此外,正如有用户所言,改变成年人的工作习惯极其困难。即便你的产品比现有的更好,心理迁移成本依然是最大的隐形成本。Routine若想成为“Siri for work”,不仅要解决功能集成与自然语言理解的问题,更要通过强大的跨平台同步和低迁移门槛,让用户敢于换个活法。否则,它只是一款更好的“整合版日历+待办清单”,而配不上“AI”这个定语。

查看原始信息
Routine AI
Routine AI lets you control your tasks, calendar, notes and projects using your voice. Just talk naturally to schedule meetings, add reminders, capture ideas, write notes, search your knowledge, update projects and automate repetitive work.
👋 Hey Product Hunt, I'm Julien, founder of Routine. We started Routine a few years ago with a simple idea: your calendar, tasks, notes, meetings, projects, and contacts should live in the same app. Earlier this year, we began building AI agents and automations that could understand all that context and take action on your behalf. Then we realized something surprisingly simple. Because Routine already knows about your work and your day, why shouldn't you be able to talk to it like a real assistant? With Routine AI, you can capture ideas, create reminders, schedule meetings, organize your day, and much more using nothing but your voice. For example: "Move my afternoon meetings to Friday, block two hours tomorrow morning to finish the pitch deck, and remind me on Monday to submit my YC application." A few seconds later, it's all done. Voice commands are just the beginning. Routine AI also includes AI agents and automations that help eliminate repetitive work across your tasks, projects, calendar, notes, and more. It's free to try, and we'd genuinely love to hear what you think.
11
回复
@jmq I really like the overall direction. Combining tasks, calendar, notes, AI meeting notes, and time blocking into a single workspace makes it feel like a true productivity hub instead of just another task manager. The clean UI and cross-platform support are also big positives. One suggestion: I’d love to see more flexibility in dashboards and custom workflows. It would also be great if AI could proactively suggest priorities based on deadlines, meetings, and workload instead of only responding to prompts. That would make the experience even more valuable for teams and power users. Overall, impressive work!
0
回复

@jmq Huge congrats on the launch!

3
回复

@jmq Finally got an assistant on steroids, awesome! 👏

1
回复

Voice control for work is powerful when it treats speech as intent, not just text input. The useful boundary is letting the agent capture messy instructions, then make the next action reviewable before it mutates calendars, tasks, or customer-facing work.

2
回复

Hey @krekeltronics . Definitely. Mutating without guards is risky. That's why you talk to Routine, it lists what actions it will perform. You can change, remove etc. And when ready, tap a button and everything is taken care of.

0
回复

The combination of tasks, docs, and calendar in one place actually feels useful rather than bloated, and I liked how quickly I could switch from a project view to my schedule without losing context.

1
回复

Hey @oktaytc34 . It's always been our vision to bring together information that can be leveraged in another context. And because meetings, contacts, tasks and notes all work together, it makes more sense that have all of those in one app.

0
回复

The new vocal assistant has seriously upgraded my work as a freelance since @jmq gave me early access. I've always dreamed of handling more work simply with my phone (I use Routine to coordinate my work with all my clients).

I would love to customize some parts of the interface to make it fit even better with my work!

Keep up the good work and excited to see what's next!

1
回复

Thanks @brieuc1 

0
回复

The calendar view feels a bit basic compared to the rest of the platform. Adding a timeline or Gantt-style view would make it way easier to see how tasks and projects overlap, especially when juggling multiple deadlines. Would love that as an option alongside the existing layout.

1
回复

Hey @sla919522062231 . We will definitely be adding more visualization later.

0
回复

Cool app, congratulations on the launch, it this a completely new app or does this integrate with existing google/apple cal, tasks etc ?

1
回复

Hey @vinitvr . It integrates with other apps. For instance, your calendar in Routine bidirectionally synchronizes with Google Calendar. Then you can also 2-way sync tasks from Todoist, Gmail, Notion etc. Same with your Google Contacts and contacts in Routine.

1
回复
Most voice assistants are useless because they have no idea what you're actually working on. At least this one already lives where all your stuff is. Congrats on the launch!
1
回复

You're exactly right @etiennegarcia . Thanks.

1
回复

the sidebar layout is genuinely satisfying, everything lives exactly where you'd expect it to. rare to see this many features packed in without the UI feeling cluttered.

1
回复

@kranurar4m Thanks. It took us some time, I won't lie :)

0
回复

the example command in the launch bundles three separate actions in one utterance (move meetings, block time, set a reminder). if one of those three fails to resolve - say the Friday slot is already double-booked - does it still summarize and confirm the two that worked, or does the whole chain get held up waiting on the one that's ambiguous?

0
回复

Love that a voice command can actually update your tasks and calendar

0
回复

voice control for calendar and tasks sounds great until you say something ambiguous and it moves the wrong meeting or deletes the wrong reminder. with something like 'move my afternoon meetings to Friday' - does it show a confirmation of what it's about to change before committing, or does it just execute and you find out after the fact if it misheard something?

0
回复

I wonder if the biggest challenge is changing habits. Most people already have a way of managing tasks, even if it's inefficient. What has convinced users to switch?

0
回复

honestly looks solid, the calendar + tasks combo is what i usually end up juggling across like three apps. one thing that would be really useful is a quick "focus mode" that hides everything except your current task and its linked docs, basically a stripped down view when you actually need to get stuff done without distractions.

0
回复

The way you bundle tasks, docs, and calendars into a single clean interface without it feeling cluttered is really impressive. Took me about ten seconds to figure out where everything lives.

0
回复

voice for capturing tasks makes total sense, i lose half my ideas walking between meetings. the part i always doubt is corrections. when it mishears a name or a date, can i fix it by voice or am i back to typing? that flow is what would make or break it for me

0
回复

honestly looks solid, the all in one angle is appealing. one thing i'd want is a built in pomodoro or focus timer that syncs with the task list, kind of a small quality of life thing but it would make routine feel more like the hub for actual work rather than just organization.

0
回复

@keremdiben, Routine does have a Pomodoro timer, and task-specific times as well. All synced with tasks and events.

0
回复

One thing that would make Routine a lot more useful for me is a built-in time-tracking view that lives inside tasks, so I can see how long things actually take without bouncing to a separate timer app. A simple weekly breakdown by project would be enough.

0
回复

When you open a task, you have the list of allocations i.e the time you blocked to work on that task. Plus you can run a timer for a specific task. When it stops, it records the time in your calendar as an additional allocation. So basically exactly what you describe.

0
回复

honestly looks solid, the all-in-one angle is super appealing. one thing that would push this over the edge for me is a proper meeting notes template that auto-links to the relevant task or project page, basically turning a random note into actual traceable work without extra clicking.

0
回复

Hey @demet75124 . Well even better, meetings are automatically recorded and transcribed. Then, you can define an AI automation to run to do whatever you want with the transcript.

The default automation summarizes the transcript and updates the event with the summary above the transcript.

But you can do whatever you want. You could create tasks in your team workspace and assign the tasks automatically.

0
回复

Honestly kind of surprised how well the calendar and tasks fit together, like everything just lives in one place without feeling cluttered. The wiki is pretty barebones though.

0
回复

Thanks @melisnceda1tw6 . What's missing in the note editor in your opinion?

0
回复

Does Routine AI ask for confirmation before making changes like moving meetings or updating tasks, or can it do everything automatically?

0
回复

You can decide. In the assistant/agent configuration, you can indicate if you always want to validate before something is done, or just related to events for instance. Of if you want it to execute by itself. Everything is configurable through the user preferences.

0
回复

Congrats on the launch, Routine looks like a really clean way to keep everything in one place. One thing I'd love to see is native time blocking inside the calendar view, like dragging a task from a project onto a time slot and having it actually reserve that time. Would make planning my week way less of a juggling act between tools.

0
回复

Hey @nermina7zv . This is already supported. See https://www.youtube.com/watch?v=Ym5yBC-QtZU

0
回复

The all-in-one approach is appealing, but a native time-tracking feature built into tasks would really set it apart. Being able to start a timer on a task and see aggregated reports by project or client without needing a separate tool would make Routine a true one-stop shop for freelancers and small teams like mine.

0
回复

Hi @fikretkoar6dg . There is already task-specific time tracking. See here: https://www.youtube.com/watch?v=cJkbaBfHGqE

Then you can ask the AI assistant or an agent to compute the sum of the time spent across all the allocations (as it is called in Routine) i.e the times you spent on a task.

0
回复

Love how Routine pulls everything into one place, the calendar + tasks combo has already replaced three tools for me. One thing I'd love to see is a lightweight inbox-style view that surfaces comments, mentions and assigned items across projects in a single feed, so nothing slips through when work moves fast.

0
回复

Hey @berkehdyw . The inbox already surfaces mentions and collaborative work assignments. However you are right, we do not support comments yet. But those will come soon. And when they do, you will receive a notification in your inbox.

0
回复

Love how you can keep tasks, docs, and calendars in one place. One thing I'd love to see is a quick "weekly review" view that pulls together completed tasks, updated docs, and upcoming meetings so I can actually reflect on the week without jumping around.

0
回复

Hey @asminy6i5 . Well that's simple. Create a weekly AI automation that says something like "Analyze all the tasks I've worked on during the week and meetings I attended. Prepare a weekly report for me to reflect upon". Done. The AI will notify you once done with the report. Or you could even ask the AI automation to save the report in a database or document.

0
回复

Finally checked Routine out and the calendar hooked up nicely with the task list, which is more than I can say for most tools. Wish the wiki formatting was a touch richer.

0
回复

Hey @dilaraekme8ruh . Thanks. What would you expect to see in the document formatting?

0
回复

Love that Routine combines tasks, docs, and calendars in one place, it would be great if you could add a "focus mode" that hides everything except your current task and a simple timer to track deep work sessions without leaving the app.

0
回复

Indeed, we've thought about it for several years@metehangtcx . But not many people asked for it actually: https://feedback.routine.co/p/focus-mode-4

0
回复

This video looks so good, but does it support any other language, or how does it remind us by notification or something?

0
回复

Hey @flytosky . Yes you can speak in any language. It will determine the language after a few words. As for reminders, it basically creates data in Routine. Because Routine manages your tasks, notes, calendars, contacts, projects etc. you can create and manage everything in your life through voice.

0
回复

One thing that would really help me as a user is a built-in time tracking feature inside each task, so I can see how much time I actually spend on a project versus what I estimated. Right now I have to jump to another app for that.

0
回复

Hey @boran993212 . You can already do that with allocations. Drag & drop tasks in your calendar. Or use task-specific timers with the menu bar widget to record in your calendar time you've spent on a task. Then ask your assistant to sum it up. That's it.

1
回复

voice for capturing tasks is the rare spot it genuinely beats typing, cause the friction of opening an app is exactly when you forget the thing. 5th launch, clearly iterating hard. nice one.

0
回复
0
回复
Great product! Can it help search my emails too?
0
回复

@raphael_goldsztejn Sure. Just connect Gmail to Routine. Then tell your AI assistant or agents to search with Gmail.

1
回复
#6
Phantomstory
Launch a third-party blog to win AEO with just two clicks
211
一句话介绍:Phantomstory通过两键快速部署第三方独立博客,以增量式AI就绪内容帮助B2B企业在AI搜索(AEO)中赢得精准推荐,解决传统AEO策略难以落地和量化的问题。
Marketing Growth Hacking Developer Tools
AEO优化 第三方博客 AI搜索可见性 内容营销自动化 AI内容生成 品牌诚实植入 增量发布 域信任 EEAT B2B营销
用户评论摘要:用户普遍关心:第三方博客的独立性与品牌关联是否披露,如何防止成为穿马甲的PBN;域信任如何随时间积累而非速成;AI模型是否会看穿“赞助关系”(如医疗等领域);能否量化AI引用归因;以及EEAT信号在YMYL领域(如健康)的落地。也有用户肯定增量发布而非内容泛滥的思路。
AI 锐评

Phantomstory切入的是一个真实且日益迫切的痛点:当ChatGPT和Claude成为B2B买家的第一搜索引擎,传统的SEO和付费投放正被AI的选择性“摘引”瓦解。创始人点中了一个关键洞察——大模型天然信任第三方来源而非品牌自说自话。在此基础上,Phantomstory不是要做又一个AI批量生成垃圾站的工具,而是试图用“诚实植入+增量发布+事实核查”来构建内容可信度,这比单纯的“AEO监控仪表盘”显然要聪明得多。

但真正值得深究的是产品长期的可持续性问题。评论中已有用户直言不讳地指出:Google花了15年学会打击PBN(私人博客网络),今天的AI引擎又比搜索引擎更难以用“引用深度”来欺骗。Phantomstory声称“只建1-3个强站点”,这意味着用户本质上在经营一个有归属关系的“媒体”,而非匿名网络。这考验的是:品牌是否敢于接受一个“不完全偏向自己的第三方内容”?如果系统本质上是品牌付费的,编辑独立性将永远是悬在头顶的达摩克利斯之剑——AI模型的信任算法迟早会进化出更精密的“本体关联检测”。

此外,产品当前最强的叙事是“启动速度”和“自动化”,但这恰恰掩藏了真正的门槛:持续的运营质量与市场适配。对于医疗、金融等YMYL领域,EEAT(经验、专业、权威、信任)的核心在于可信的作者背景和可追溯的署名,而非“一个小时内由AI代理完成研究-写作-验证”的流程。创始人坦诚“配置化披露”和“行业差异”,说明产品仍处在“工具”阶段,尚未形成模式验证——用户必须自己拿捏从“赞助网站”到“独立媒体”的灰度。

Phantomstory真正的价值,不在于它是否比传统AEO工具“更智能”,而在于它为品牌提供了一套低成本、低风险的AEO实验平台。它能帮你快速跑通“内容-引用-转化”的最小闭环,前提是你愿意接受它的本质:在AI的认知盲区里,用可控的“白手起家媒体”完成一次面向机器评委的讲稿。这需要产品在数据归因和信任算法上沉淀出清晰的可度量指标(如“LLM引用率 vs 同类竞品”),否则将沦为又一个“看起来很美”的自动化内容工厂。

查看原始信息
Phantomstory
Phantomstory helps teams win AEO by launching third-party blogs on fresh domains in under 5 minutes. With just two clicks, Phantomstory provisions credible content hubs that incrementally publish AI-search-ready articles. The secret? A portion of these expertly-written articles are designed to help LLMs discover, understand, and recommend your company. Stop thinking about AEO as just a monitoring problem; with Phantomstory, transform AEO into an actionable plan that actually drives results.

Hey Product Hunt! 👋

I'm Mathew, founder of The Letter Company. Today we're launching Phantomstory: our coolest product yet.

I've been a content marketer for the last four years, and I kept seeing the same pattern: one of the most effective strategies for winning AEO is third-party content. ChatGPT and Claude are smart cookies. They know a brand will always rank itself #1, so they trust third-party sources instead.

But here's the uncomfortable part: most "third-party" coverage today is a paid product placement, sold to the highest bidder.

What if third-party content could actually be honest?

That's why we built Phantomstory. With just two clicks, you can launch a complete third-party publication: a real media site that produces genuinely helpful content for readers, with honest placements of your brand on the articles where it actually belongs. It uses incredible technology: adversarial writing kernels, a smart KD-sensitive priority queue, and an amazing canvas tools to create stunning visuals.

Here's what you can do with Phantomstory:

🚀 Launch a publication in two clicks
Spin up a full third-party media site, complete with its own editorial identity. No agencies, no writers' rooms, no six-month content calendars.

✍️ Write with editorial voice, not marketing voice
Our writing kernels produce copy that reads like a practitioner wrote it for other practitioners. No fluff and no sales language.

🔍 Verify everything before it ships
Research agents fact-check claims and statistics so your publication earns trust instead of burning it.

🎨 Enhance every article automatically
Enhancement agents add visuals, citations, comparison tables, and the metadata and structured data that LLMs actually parse.

🤝 Place your brand honestly
Your product gets mentioned on the articles where it genuinely deserves a mention, next to fair coverage of alternatives. Readers get real value, and AI assistants get a source worth citing.

Phantomstory is built for devtools comparing themselves against alternatives, platform companies explaining a category they're creating, and any B2B brand whose buyers now ask ChatGPT instead of just Google.

Dozens and dozens of companies are already using it since launch last week, and we'd love for you to be next.

👉 Try Phantomstory: https://phantomstory.com

Drop your questions in the comments. I'll be here all day and would love to hear how you're thinking about AEO. 🙌

Mathew

10
回复

@pregasen From the buyer side this is the part nobody can price properly. We've paid agencies for "AEO" work and the deliverable was a spreadsheet of placements with no way to tell which ones models actually read. How do you decide which third-party sites are worth publishing on, and can I see attribution per placement or is it a black box?

8
回复

@pregasen Since Phantomstory can manage multiple autonomous blogs, how do you prevent those publications from becoming too closely associated with a single brand? Is there an editorial diversity layer that varies writing style, topic selection, and linking behavior so each publication develops its own credibility over time? That seems like a key factor for long-term AEO success.

0
回复

@pregasen Curious, what metrics have you found to be the strongest indicators of success for AI search optimization beyond traditional SEO rankings?

0
回复

I like the idea of building third party blogs so quickly. How do you protect domain trust over time instead of creating content too fast? Publishing steadily with clear expertise could make recommendations more reliable.

6
回复

@advin_jadis On average, people use our platform to publish 1-2 articles per day but never more! They will target very specific questions that there isn't content for, build domain authority that way, and then spiral upwards from there. Open to exploring the platform with me?

1
回复

I've always believed consistency beats volume, so the incremental publishing angle caught my attention. It sounds like a smart way to build credibility over time instead of chasing quick wins.

5
回复

@sheikh_umair1 Incremental publishing is extremely important. It's how you build actually helpful content!

1
回复

Mathew, I run marketing for a healthcare AI company, so I read this with equal parts curiosity and caution.

In our category the sources LLMs cite are held to a higher bar: health content gets extra scrutiny from both models and regulators, and a fresh domain with no named authors rarely makes the cut.

The launch copy leads with honesty, which I respect, so here is the question that decides it for me: is the relationship between the publication and the brand disclosed to readers, and how do you think about that for regulated industries like healthcare?

4
回复

@clemente_lopez1 Fantastic question. That is entirely configurable. For some of our customers, there is no disclosed relationship. It's effectively a sponsored website in the same way that Hearst might own a magazine that you read without their logo attached. For others, the sponsorship is direct or prominently disclosed. It depends what you're trying to do, and who / what models you're trying to win over.

Let's chat? I'd really like to learn more about your space.

1
回复

I'm seeing more marketers realize that AI visibility depends on discoverable, trustworthy content. What stood out for me is the focus on incremental publishing rather than dropping everything at once. That feels more natural.

4
回复

@thomas_jack3 Absolutely. Our business isn't in the game of creating a bunch of spammy, un-helpful content. Even when using our full AI-writing pipeline, the agents are taking over an entire hour to produce an article. There's a lot of research, fact-checking, and enhancement going on.

Open to seeing the platform with me? I'd really appreciate that.

1
回复

I really like the idea that a brand is only included where it genuinely belongs, rather than being forced into every article. Especially in a space where credibility is so important, that feels like a thoughtful approach and much closer to a real publication than a typical branded content operation.

What signals does Phantomstory use to decide whether a customer’s product should be mentioned in a specific article, and when does it deliberately leave the brand out?

4
回复

@nico_mandera Really amazing question. I rarely get asked this question, but I'm glad you did because it's something we really thought about. There are two reasons that a brand gets mentioned: the first is that an article is tackling a more pointed topic around what the brand is actually building. For example, if you're an LLM router, then you might get mentioned in an article that's comparing LLM Routers versus MCP Gateways, but not in an article that's discussing what LLM Hallucinations are. The second reason is that an article on the phantom is doing really, really well, and that light placement doesn't hurt.

Open to exploring the platform with me over a call?

1
回复

Definitely adding this to my list of marketing tools to explore. Curious to see how early users leverage it for AI search visibility.

3
回复

@1mirul thank you. let me know if you ever want to chat.

0
回复
Congrats on the Product Hunt launch! Phantomstory looks like exactly what the industry needs right now.
3
回复

@odeth_negapatan1 Thank you so much. Open to exploring it with me? Would love to learn more about Lancepilot and if we could help.

0
回复

Congrats on the launch, this is a clever take on AEO. One thing I'd love to see is a way to pick the niche or topic cluster before the domain is provisioned, so the auto-generated content actually matches the voice and positioning of my product instead of feeling generic. Right now I worry about getting a credible-looking blog that still drifts off-brand.

2
回复

@derinalmargwoo You can do that! Can I show you?

1
回复

Neat product! The 1 to 2 articles a day pacing is the detail I'd personally have gotten wrong if I built this:), every instinct says flood the fresh domain on day one! How's the configurable disclosure? Have you guys measured whether the disclosed publications get cited less by ChatGPT and Claude than the undisclosed ones?

2
回复

@artstavenka1 Great question. Disclosed publications are cited the same but then referenced less in ChatGPT and Claude's response. But they still get cited more than first-party content since models steer away from marketing sites! Open to exploring the platform with me?

2
回复

clever tactic, but the thing i'd watch is whether AI answer engines eventually treat coordinated blog networks the way google learned to treat PBNs. search spent 15 years getting allergic to exactly this shape. if the answer engines stay citation-hungry and don't care, real edge for a while. if they catch the same allergy, fresh-domain hubs age fast. genuinely curious which way it breaks.

2
回复

@alex_watson2110 It's a really great question. We aren't trying to help companies create thousands of these websites—just one to three strong phantoms. They'll have good quality content and organically grow. The better question is how good will LLMs get at suss'ing out what is and isn't sponsored under-the-wraps, and that's something we'll always be on our toes to watch out for and make changes as needed.

1
回复

i think the simple workflow makes adoption eaiser. how do you monitor domain performance after launch? a dashboard with growth suggestions could help teams improve results over time.

2
回复

@donna_gerrard Great question. We track four metrics: impressions, clicks, average position, and share-of-voice of the underlying brand. Open to seeing it with me?

1
回复

There is definitely something interesting in the idea of building useful third-party content rather than relying only on a company blog. The part I’m curious about is independence. If the brand creates and funds the publication, how do you stop it from becoming branded content wearing a media hat? Do publications clearly disclose who owns them, and can the system ever produce a conclusion that does not favour the company behind it? That feels important if the goal is genuine trust rather than just another AEO tactic. Congrats on the launch!

1
回复

@os_ishmael These are entirely configurable! Depends on the company's approach :) some would rather faint than ever be negatively portrayed, others embrace being balanced. Our platform itself is agnostic.

0
回复

Interesting angle. We run a content-heavy healthcare directory and invest a lot in original medical content written by physicians — curious how you handle E-E-A-T signals on a third-party blog? For YMYL niches like healthcare, author credibility seems like the hard part, not publishing speed. Congrats on the launch!

1
回复

@thetozkar Amazing question. EEAT for the most part works the same on third-party blogs as it does on first-party blogs. Credibility is everything which is why we integrate with the world's best research agents, including those that index papers. Let's chat?

0
回复

Phantomstory sounds aimed at teams trying to win AEO without building a whole publishing setup. When you say “launch a third-party blog,” how much control does the user get over branding, topics, and the structure of posts? Curious where the two-click setup ends and where ongoing editorial choices begin.

1
回复

@ivory_xuxuxu Great question. You have full control. The two clicks is to stage the blog, and if you want to be completely hands-off after, you could. But most of our users get their hands dirty by configuring their blog clusters, the voice, the branding, the structure, the graphics, etc. Our platform is full of tools. Open to seeing it with me? It's probably one of the most stacked products on the market when it comes to customization :)

1
回复
I get the concept but “honesty” and owning the “third-party” blog seem to directly conflict. With Search engines still driving traffic, what happens when you get penalized for having a PBN?
1
回复

@mattdowis Valid question, but this isn't a PBN. We aren't creating thousands of websites that link back to a single blog. We're creating 1-3 new blogs—in the same way you'd create your own company blog—that independently build domain rank and do marketing heavy-lifting. Is it still biased? 100%. But 80% of the content on a phantom site is entirely focused on answering a question without brand placement, and a portion of the articles are then used to build brand awareness.

Open to seeing on a call with me?

1
回复

This is an interesting concept in the world of AEO, especially since you can do it fairly quickly. I don't see it on your website, but how does pricing work for this? Also, how do you handle protecting domain health and authority? Or is this is a "quantity of AI blogs" as a strategy type of approach?

1
回复

@denitsapenchevavaltchanova Our pricing is quite approachable. It's just $500/mo and we'll manage up to 1-3 phantom sites with you with reported domain authority (impressions, clicks, average position, and share-of voice).

This is not a quantity game. It's all about quality. We're not trying to flood the Internet with a bunch of rubbish articles, but rather create really pointed and well-organized content in your space. Open to seeing the platform with me?

2
回复

Wow but i would love to see website in light theme also

1
回复

@abhijeetkumar_founder You might like our main storefront then :) https://letterstory.com

0
回复

I like the positioning around honest brand placement instead of forcing a product mention into every article. If AI search continues rewarding credibility over promotion, that philosophy could become a real advantage.

1
回复
1
回复

The combination of automated publishing, editorial quality, and AI-search optimization makes Phantomstory stand out in a crowded marketing space. Wishing the team a successful launch! 🚀

0
回复

the Hearst comparison undersells the difference though - Hearst's ownership of a magazine is public record even if the logo isn't on every page. with a fresh domain and no named authors, is there actually anything a skeptical reader or a journalist could find that traces the site back to the brand, or is the undisclosed version designed so that trail genuinely doesn't exist anywhere?

0
回复

Honestly the speed surprised me, I had a full blog up on a fresh domain in just a few minutes and the AI-search angle feels genuinely useful rather than gimmicky.

0
回复

honestly this looks really useful for anyone tired of just tracking their brand mentions in AI search. one thing that would make it way more valuable for me is a built-in dashboard showing which articles are actually being cited or pulled into AI responses over time, basically closing the loop between publishing and real LLM pickup.

0
回复

Setting up a fresh blog on a new domain took literally a few minutes, which still surprises me. The articles read naturally and weren't obviously stuffed with brand mentions.

0
回复

Spun up two blogs on new domains in under five minutes and the articles actually read like they were written by someone who knows the space, not generic AI filler. Curious to see how it holds up over a few weeks of incremental publishing.

0
回复

The disclosed-vs-undisclosed citation gap you shared is the most interesting number in this thread. One thing I'd watch: a fresh domain isn't in any model's training data, so every citation it earns comes purely through the live retrieval and grounding layer, not baked-in knowledge. That makes it far more exposed to a model swapping its search backend than an established domain is. Have you seen citation rates jump around when ChatGPT or Perplexity change how they ground their answers?

0
回复

honestly the speed kind of surprised me, had a blog up on a fresh domain in like three minutes flat. still wrapping my head around the aeo angle but it feels more actionable than just tracking metrics.

0
回复

the "third-party blog on a fresh domain" part is what gives me pause, if the whole point is that it reads as an independent voice recommending your product, isn't that an undisclosed conflict of interest by design? curious how you're thinking about disclosure here, or whether disclosing kind of defeats the purpose since the credibility comes from the blog not looking affiliated in the first place.

0
回复

any real examples of this you can share? how is phantomstory using phantomstory to win AEO?

0
回复
#7
Jockey by TwelveLabs
The video AI agent that understands your whole library
193
一句话介绍:Jockey是一款由TwelveLabs开发的视频AI代理,能像人类一样理解整个媒体库,通过人物、场景或语境语义搜索海量视频和图片,彻底解决传统元数据搜索无法精准定位具体时刻的痛点。
Productivity Artificial Intelligence Video
视频AI代理 语义搜索 媒体库管理 视频理解 多模态检索 MCP集成 时间轴定位 内容创作 AI剪辑 TwelveLabs
用户评论摘要:用户关注点集中在:1)隐私合规和确定性问题(如Pegasus分割一致性);2)复杂时序查询(如“门打开后”的因果推理)的分解能力;3)产品功能改进(可视化时间轴、帧精确跳转、浏览器插件);4)长期库更新时可否增量处理。多数用户对语义搜索的准确性表示惊喜。
AI 锐评

Jockey的真正价值不在于“更智能的搜索”,而在于它从“单文件理解”跃迁到“语料库级别推理”的架构性突破。传统视频AI本质上是“增强的标签系统”——用模型替代人眼做多标签分类,但依然是片段级认知。Jockey通过嵌入模型Marengo(检索)与视频语言模型Pegasus(分段)的协同,加上一个独立的记忆层,构建了一个可跨文件分解、推理、合成的智能体系统。这意味着“给我剪个高光集锦”这类任务从人工体力活变成了纳秒级的AI决策。

但产品当前状态更像一个“惊艳的技术演示”而非生产力工具。评论中高频出现的问题是:时序因果推理的稳定性、返回结果的确定性(对合规场景致命)、以及用户渴望的帧精确交互界面。Jockey强在“理解”,弱在“可操控”——用户获得的是文本时间戳,而非可直接拖拽编辑的视觉蒙太奇。MCP接入Claude虽能渲染播放界面,但这是一种“用AI的AI”的折叠路径,对于非技术用户仍不够直白。

更值得警惕的是,TwelveLabs用“改进自动生效、无需重新集成”来淡化模型迭代的潜在副作用——语义理解的标准在不同版本间是否漂移?对合规审计来说,昨天的模型说“握手”是第15秒,明天的模型说在第16秒,这种非确定性是致命的。此外,隐私问题被轻轻带过:图像/视频的本质是生物特征与场景的高保真数据,上传到云端训练推理,用户控制权几何?

Jockey已经摸到了“视频即数据库”的下一级入口,但要让每个普通创作者都用上,还差一个直觉界面、一份隐私白皮书,以及一套可审计的确定性承诺。它的未来不在科幻,而在工程细节的暴力打磨。

查看原始信息
Jockey by TwelveLabs
Jockey is the first AI that understands your entire media library just like you do, searching by person, moment, or context across every photo and video you've captured. Powered by TwelveLabs' advanced model stack, Jockey improves automatically with every update. Whether you need to connect via MCP for Claude/ChatGPT or build custom applications using our API, Jockey makes your media library instantly searchable and accessible.

I'm Aiden, co-founder and CTO at TwelveLabs. Super excited to share Jockey with you today.


Most media search is still metadata search: filenames, timestamps, maybe some object tags from an older computer vision model. None of that captures what's actually happening in your footage: who's in a scene, what they're doing, the context, dialogue, on-screen text. For that you need models that natively understand time and space in video, not a bag of sampled frames.


That's our stack. Marengo, our embedding model, resolves a query like "the moment we almost missed the flight" to real retrieval across video and images, not keyword matching. Pegasus, our video-language model, segments an entire video on a schema you define and returns structured, timestamped moments. Jockey is a unified agentic system that reasons across your videos and images: a reasoning model plus a memory layer that builds a knowledge store from your corpus, so it can decompose a query, retrieve, segment, and reason across the whole thing.


The point is a corpus-level understanding you can act on. Point Jockey at thousands of videos and images, say "cut me a highlight reel" or "pull the best viral moments," and it comes back with timestamped cuts you can use. The model-only approach can't do that as dumping one video into a context window is bound to a single file, and a single forward pass runs out of room fast. It can tell you about one video; it can't reason across your catalog or build a reel from thousands.


Because the models are what we ship and improve continuously, Jockey's reasoning and retrieval quality improve as we push new versions, meaning no re-integration on your end.


Two ways in:

  • MCP server: connect Jockey as a tool in Claude and query your library directly. ChatGPT coming soon.

  • API: full programmatic access to build custom retrieval or agent workflows on your own library.

This is a research preview, so if you hit edge cases (ambiguous queries, retrieval misses, latency) I want to hear about them. Let us know anytime! 


Best,

Aiden 

13
回复

@aiden_lee7 Congrats on the launch, Aiden. For teams using Jockey at scale, what’s been the biggest real-world surprise when the system reasons across thousands of videos? Any specific edge cases you didn’t expect that others should watch out for?

0
回复

@aiden_lee7 Congrats on the lauch!I just visited TwelveLabs’ webpage—the design is incredibly cool! Since TwelveLabs can perform more specific data-scene retrieval on images and videos within my media, could this potentially infringe upon users’ privacy?

0
回复

Great name Aiden! The "no tags needed" pitch is the part I'd want to/ am going to stress test. Natural language search over video is great for discovery, but for compliance/legal use cases you usually want deterministic, auditable categories, not a model's best guess at a scene. Curious how consistent Pegasus's segmentation is run to run on the same clip.

2
回复

The compositional queries are where library-scale video search tends to break. Single-concept stuff like 'red car' works fine off embeddings, but 'the moment right after the door opens' needs temporal grounding that flat similarity search can't reach. When we built multimodal search over video the causal and ordering queries were exactly where recall fell off a cliff. Does Jockey's agent decompose those multi-step queries and reason over ordering, or is retrieval a single embedding lookup under the hood?

1
回复

hi@dipankar_sarkar, great question. Through responses API (or query MCP tool), the agent decomposes the query into planned multi-steps to return the result. We separately offer a primitive knowledge-store search, which is the latter you asked "a single embedding lookup under the hood".

1
回复

honestly the search is way sharper than i expected, like i threw in a random cooking video and it pulled out the exact moment they mentioned "fold in the eggs" without any tagging. pretty cool to actually see video understanding feel useful

1
回复

would be cool if you added a built-in timeline view so you can jump straight to the exact moment a search result happens in the video. right now i think you only get timestamps in text, which is helpful but still requires manual scrubbing. basically a visual scrubber with the matched segments highlighted would save a lot of clicks.

1
回复

@buketmdrp to your point, the existing LLM convention of text responses is often insufficient, especially when working with video. this is exactly why we built the MCP! as part of the tool set, we've added the ability for Jockey to render a response in a format much more suited for video. so instead of getting back text and timestamps, you can actually visualize with playback the exact moment that Jockey is referring to. add it to your claude as a connector (https://mcp.twelvelabs.io/jockey/mcp) and let us know if you have any feedback!

0
回复

the corpus-level reasoning is the interesting jump here, most video search tools stop at "find the clip" and Jockey is going further to "reason across everything and build the reel." one thing I'm curious about that's different from the privacy/consistency questions above: as you keep adding new footage to a library over time, does Jockey only need to embed the new additions, or does growing the corpus mean periodically reprocessing the whole thing to keep the reasoning layer coherent?

1
回复

Hey @galdayan, great question! Jockey only needs to embed the new additions, but it does it an intelligence manner to keep the entire reasoning layer coherent.

0
回复

The "understands your whole media library" angle is really compelling. I already have product screenshots, launch videos, screen recordings, demos, and random clips scattered everywhere, and finding one specific moment usually depends on remembering the filename or roughly when it was created.

Searching by person, moment, or context feels much closer to how people actually remember media. connecting that through MCP is interesting too, because an agent could finally find the exact clip or screenshot needed for a task instead of asking me to dig through folders.. :) Curious how Jockey handles privacy and indexing for very large personal or company libraries.

1
回复

Was lucky enough to get access to Jockey a week or so ago (thank you team!), and have been blown away at the use cases we've already uncovered. It's changing the way we look at creative strategy across our roster of clients, and we've found some novel ways to extrapolate learnings that are informing some of our performance marketing campaigns that are already showing positive uplift.

Go TwelveLabs!

0
回复

Finally gave this a spin on a few clips and the semantic search actually nails what's happening in the scene, not just matching keywords. Impressed it picked up on subtle actions without me tagging anything.

0
回复

A live collaboration mode would be huge, letting a team tag and comment on different timestamps together while the AI pulls those notes into a shared summary report.

0
回复

Honestly impressed by how well it picks up on subtle visual cues in long videos, not just obvious keywords. Searched a 40 minute documentary for a specific gesture and it nailed it in seconds.

0
回复

One thing that would make this way more useful for me: a timeline-based search view where I can scrub to the exact moment a concept appears, instead of just getting text hits. Right now it sounds like the results are summaries, but for video editing workflows I really need frame-accurate jumps. Would love to see that built in.

0
回复

honestly the search across hours of footage feels almost scary good, like you throw in a rough idea and it actually pulls the right moments out without much fuss.

0
回复

honestly the search is way better than i expected, threw in some random clips and it actually pulled out the exact scenes i was thinking of. pretty wild that you can basically ask it questions about what's happening in the video.

0
回复

A browser extension that lets you right-click any video and instantly get a chapter breakdown or summary using your Pegasus model would be huge, especially for longer YouTube content or lectures where I don't always want to scrub manually.

0
回复

finally a video ai that actually finds the moment i describe instead of just dumping timestamps. tried it on some old footage and the semantic search picked out exactly the scene i was thinking of.

0
回复

Searched a bunch of old vacation clips by typing "sunset over water" and it actually pulled the right moments, which kind of startled me. The natural language search feels way more useful than tagging everything manually.

0
回复

The way the search results show exact timestamps with preview thumbnails makes me feel like I'm scanning a real video library, not just a text index. That attention to temporal precision shows serious craft.

0
回复

honestly the search looks solid, but it would be super helpful if you could save and share specific video clips or search results with timestamps baked in. basically a way to send someone straight to the exact moment in the video rather than just linking the whole thing and saying "go to 4:32". that would make it way more useful for team collaboration

0
回复

A live collaboration mode where teams can annotate and tag specific video segments together in real time would be huge. Right now analysis feels like a solo task, but most of our video review happens in group settings where marketers, editors, and strategists need to align on what they see.

0
回复

Would love a timeline-based annotation view where I can click any moment in a search result and instantly see the surrounding visual and audio context. Right now results feel like a black box of timestamps, but seeing a quick visual snapshot before clicking through would make reviewing long footage way faster.

0
回复

@pek4yco thanks for the feedback! would love for you to use the MCP (https://mcp.twelvelabs.io/jockey/mcp) which has a "render" tool that visualizes responses in a much more video-friendly way. instead of a block of text with timestamps, you can see results with playback.

0
回复
#8
CreateOS Sandbox
Instant, hardware Isolated Sandboxes for AI agents
175
一句话介绍:CreateOS Sandbox为AI代理提供毫秒级启动、硬件隔离的运行环境,通过内核级eBPF网络策略解决不可信代码执行时的安全隔离与合规痛点。
SaaS Developer Tools Artificial Intelligence GitHub
AI代理沙箱 硬件隔离 eBPF安全策略 Firecracker微VM CI/CD管道 Claude插件 私网网格 BYO基础设施 SDK/CLI 合规认证
用户评论摘要:用户验证了30ms冷启动和eBPF强制外联控制的实际效果。高频提问集中在:沙箱fork后的调试工作流、生产环境迁移、DNS解析绕过风险、销毁审计日志细节。创始人明确回应了外联规则基于iptables而非eBPF,并对DNS安全边界、fork策略继承时序做了技术澄清。
AI 锐评

CreateOS Sandbox真正稀缺的价值不是“快”,而是它在“隔离的安全层级”上做了正确的取舍。市面上多数AI代理沙箱要么依赖进程级命名空间(容器),要么就是纯粹的虚拟机(慢且笨重)。它基于Firecracker微VM,让每个代理拥有独立内核,将安全边界上移到Hypervisor层,同时将启动时间压缩到与容器相近的30ms——这在工程上是可行的,但更关键的是其网络策略的实现哲学:不将信任寄托于沙箱内部的代码(即便它被攻破),而是在主机侧通过iptables+透明代理从外部强制实施。这解决了AI代理场景下最尴尬的悖论——你让一个可能受污染的代理去执行命令,却又指望它自己遵守规则。

值得警惕的是,其技术栈的稳定性。创始团队在评论区纠正了eBPF在策略中的实际角色(仅用于跨租户VM隔离),这在初期容易造成误解。对于重度依赖HTTPS以外流量的用户,明文HTTP的DNS绕过风险已被文档标记为已知缺口,这需要在安全审计中被重点考量。另外,尽管BYO硬件和SOC 2认证对合规型企业是强卖点,但产品目前仍处于Alpha阶段,其调度器的多租户抗噪声能力、大规模fork时的资源竞争、以及销毁后的数据残留审计仍是考验点。

总体而言,这不是一个“玩具沙箱”,而是为“自主执行且不完全可信”的AI原生负载设计的底层基础设施。不过,其真实成败将取决于:当AI代理以毫秒级频率创建和销毁成千上万个沙箱时,它所承诺的每一条安全边界是否依然能毫无妥协地守住。现在积累的50+示例库和500免费额度,不过是它吸引用户来“亲手打破它”的诱饵。

查看原始信息
CreateOS Sandbox
CreateOS Sandbox gives AI agent builders their own fast, secure, hardware isolated, sandbox in ~30ms (p90). We have suit of CLI, SDK, 50+ SDK real world examples, claude plugins, computeSDK integration, and more

Hey Product Hunt 👋

I'm Pratik, CTO and co-founder at CreateOS.

CreateOS Sandbox gives every AI agent its own isolated environment with its own guest kernel. Egress is enforced in-kernel via eBPF, from outside the sandbox. Code running inside cannot route around the policy, even when fully compromised.

What's in it:

eBPF egress control — allowlist by host, IP, or CIDR, enforced in kernel

Fork — snapshot a running sandbox, clone it in milliseconds

Encrypted p2p mesh networking, plus VPN back to your local machine

BYO-S3 — mount any S3 service as a shared disk across sandboxes

CLI integration — `stripe projects add createos/project` provisions a project with credentials already in your .env

BYO-Infra — run sandboxes on your own hardware

We shipped 51 real-world examples: Claude-managed agents, multi-node clusters, batch inference, ffmpeg transcoding, and more.

We also shipped a Sandbox plugin for Claude this week. Claude can run untrusted code in a box that self-destructs when idle.

Get started:

CLI: `createos sandbox create`

Dashboard: https://createos.sh/app/sandbox

Quickstart: https://nodeops.network/createos/docs/Sandbox/Quickstart

500 free alpha credits, no card.

One thing I'd like feedback on: which of your production egress rules would this sandbox break? That tells us more than a star rating.

Links:

Docs: https://nodeops.network/createos/docs

Dashboard: https://createos.sh/app/sandbox

CLI: https://github.com/nodeOps-app/c...

SDK: https://github.com/NodeOps-app/c...

SDK examples: https://github.com/NodeOps-app/c...

Claude plugin: https://github.com/NodeOps-app/c...

Sandbox as self-hosted GitHub Actions: https://github.com/NodeOps-app/c...

20
回复

@pratikbin Since CreateOS Sandbox can instantly fork running environments, do you see this evolving into a native debugging workflow for AI agents? For example, could an agent automatically branch a failing execution, test multiple recovery strategies in parallel, and merge the successful path back into production without affecting the original session? That feels like a powerful use case beyond security alone.

0
回复

@pratikbin Really interesting! How does CreateOS Sandbox handle the transition from experimentation to production? Can users seamlessly deploy workflows they've built in the sandbox?

0
回复

Super proud of the entire team for shipping this 🤩🎉

As the marketing lead here at CreateOS, seeing all the hard work go into building CreateOS Sandbox for secure, high-performance AI agent isolation has been incredible.

Can't wait to see what amazing agents developers build with sub-30ms provisioning and bulletproof eBPF egress control. Let us know what you think.

12
回复

Team effort top to bottom -> kernel work, gateway, scheduler, docs, everything. Sub-30ms provisioning was the constraint we refused to break, and it shaped nearly every decision below it. Appreciate you rallying launch day. Now let's see what people build with it.

1
回复

Been on the infra side of this build, and it's honestly wild how much our CI pipelines sped up once we moved ephemeral test runners over to these sandboxes. The self hosted GitHub Actions integration saved us a decent chunk on runner costs, and not having to hand write egress rules anymore because eBPF handles it at the kernel level has been a relief for our on call rotation. Curious to see how the BYO infra option holds up once more teams start running their own hardware behind it, that was one of the more requested features from folks on our side

10
回复

@ashw6q It’s a different approach that we went with utilizing sandboxes with github actions, and it turned out to be a game changer, especially for ops. glad that you like it

2
回复

Been using CreateOS Sandbox, some of the things that stood out to me
1. Seamless integration with my existing managed agents to have an isolated execution environment
2. I can run these sandboxes on my own Cloud, existing compute layer
3. The examples repository is really good with variety of use-cases -> https://github.com/NodeOps-app/createos-sandbox-sdk/tree/main/examples

Great work by Pratik & team.

8
回复

@naman_nodeops Thanks for the detailed writeup. Glad the BYOC story and the managed-agent integration path are landing the way we hoped, that's the gap we built this to close. Appreciate the shoutout on the examples repo too. The team put real effort into covering varied use-cases there, beyond a single hello-world. Thanks for using it and writing this up.

0
回复

If your AI agents run code you don't fully trust, this is worth a look. Each agent gets its own isolated sandbox with a real guest kernel, ready in about 30ms, so isolation stops being a tradeoff against speed.

The smarter part is the egress control. Rules are enforced in kernel via eBPF, from outside the sandbox, so even a fully compromised agent can't get around your allowlist. Add in millisecond sandbox forking, your own S3 as shared storage, and the option to run it on your own hardware, and it covers most of what actually breaks in production agent setups.

There's already a Claude plugin so Claude can run untrusted code in a sandbox that wipes itself when idle, plus 50+ real examples to build from. 500 free alpha credits, no card needed, so it costs nothing to test against your own workload.

8
回复

@navedux There is a sandbox feature in Claude-Code and Codex, but apparently it's a process level and has a lot of gotchas, and you won't feel comfortable. That's where you can utilize sandboxes with your existing cloud or Codex harness to run any kind of untrusted nodes and do all kinds of testing, benchmarking, and development.

So if you want to take it to one more level up, you can utilize SDK and create your own agent. Also you can use Anthropic-managed agents with self-hosted sandboxes.

1
回复

The in-kernel eBPF egress enforced from outside the sandbox is what makes this actually trustworthy to me — most sandboxes run policy in a userspace proxy that compromised code can just route around. Since eBPF sees packets at L3/L4, how does the host allowlist resolve names: is DNS resolved outside the box and the resulting IPs pinned to the rule, or can a compromised process inside do its own resolution and point an allowlisted hostname at an attacker-controlled IP? And on Fork, does the clone inherit the parent egress policy the instant it snapshots, or is there a window before the eBPF rules attach to the new sandbox?

7
回复

@hi_i_am_mimo Valeria's DNS question is the one I'd want answered too, but the fork/pause-resume angle is what I'm actually curious about. When a paused sandbox resumes into a new micro-VM, does the eBPF egress policy get reattached before the guest kernel starts executing, or is there a window where the resumed process is live but not yet fenced? That gap, if it exists, is usually where a fork-heavy multi-agent setup leaks something it shouldn't.

0
回复

@hi_i_am_mimo Good catch, worth being precise here. The egress allowlist itself isn't eBPF, it's a hybrid of kernel iptables (IP/CIDR/port rules, per-VM chain, DROP-by-default) and a transparent proxy that reads SNI or Host header for domain rules on ports 80/443. eBPF in our stack does a different job: cross-tenant VM-to-VM isolation on the WireGuard mesh, and optionally bandwidth metering. Both run outside the guest, so the "compromised code can't route around it" property still holds. The enforcement layer just isn't eBPF for this specific path. Appreciate you pushing on the exact mechanism instead of letting the buzzword slide.


On DNS: resolution happens inside the guest, against resolvers we auto-allow (1.1.1.1, 8.8.8.8) the moment a domain rule exists. The proxy doesn't do its own lookup. It intercepts the connection via `SO_ORIGINAL_DST`, so it sees whatever IP the guest's own resolution already picked, then checks the SNI or Host string against your allowlist. That asymmetry matters:


- HTTPS: safe. A compromised process can point `pypi.org` at an attacker IP, the proxy allows the connection, but the TLS handshake fails, the attacker can't present a cert for `pypi.org`. The wire-level identity check saves you, not the proxy.

- Plain HTTP: not safe. without https, sails through, no protocol-level identity check exists on HTTP. Documented gap, not a hidden one. If HTTP matters for your threat model, use IP/CIDR rules instead of domain rules, or stay HTTPS-only.


On fork: the new row copies the source sandbox's full egress list at fork time, so policy inherits by default unless you pass an explicit egress override in the fork request. At resume, the host installs the iptables chain and proxy rule before the resume call waits for the guest agent to answer, and before the sandbox gets marked reachable. No window where a fork you can already reach is missing its policy. Same timing model as a cold create, not a weaker one.

0
回复

Hey everyone, Mrudul here from CreateOS. I spend most of my time talking to companies that want to run AI‑generated, untrusted code in production, and the same obstacle keeps appearing. It’s rarely about whether the code runs; it’s about what the code can reach, where the data resides, and who can see it.

That’s the problem I care about with CreateOS Sandbox. Each sandbox is an isolated environment, and the most useful feature is the ability to create a private network across many of them. You can stand up a cluster of agents that communicate over private DNS while remaining sealed from everything else, something most teams assume requires a full VPC and weeks of setup.

The other key advantage is control. You can run the entire stack on your own hardware, including the control plane and storage, and we are SOC 2 Type II and ISO 27001 certified. For teams undergoing a serious security review, that distinction can be the difference between a trial and a polite “no.”

7
回复

@mrudul_gole This is the conversation on repeat, exactly. Most sandbox tools solve "does the code run" and stop, the private-network-per-cluster piece was the part we debated longest internally, because it's easy to fake with DNS tricks and hard to do right with actual per-node isolation. On the compliance side, self-hosted control plane + SOC 2/ISO wasn't a checkbox exercise -> it's literally what gets us past the security review instead of a polite no. Happy to go deeper on the private-networking design if anyone wants internals.

0
回复

As AI agents become more autonomous, we kept running into the same problem: how do you let them execute arbitrary code without putting the rest of your infrastructure at risk? We built CreateOS Sandbox to answer that. Every agent gets its own isolated environment with a dedicated guest kernel, and network access is enforced externally through eBPF, so the policy still holds even if the code inside the sandbox is fully compromised.

Beyond isolation, we focused on making it practical for real workloads. You can fork running sandboxes in milliseconds, mount your own S3-compatible storage across environments, connect securely through encrypted networking, or even run everything on your own infrastructure if that's what your deployment requires.

6
回复

@rahilmavani What's actually true: the egress allowlist runs on iptables (kernel, per-VM chain, IP/CIDR/port) plus a transparent proxy reading SNI/Host for domain rules on 80/443 — docs/egress.md. eBPF exists in the stack, but it enforces cross-tenant VM-to-VM isolation on the WireGuard mesh and optional bandwidth metering, a separate subsystem. The "enforced outside the sandbox, holds even if the guest is fully compromised" property is still true. The mechanism named in comment 1 isn't. Your call whether that's worth a follow-up correction on the thread; flagging it since it's a public technical claim under your name.

0
回复

On the product side, the whole bet here was refusing the usual tradeoff.

Every sandbox tool makes you pick: fast provisioning or real isolation. We wanted both ~30ms to spin up, and a real guest kernel per agent with egress locked down in-kernel via eBPF, enforced from outside so a compromised agent can't route around it.

If your infra can't move at the speed your agents think, the agents aren't fast, they're just waiting. That's the problem we set out to kill.

500 free sign-up credits, no card. Genuinely want to hear what breaks against your workloads.

6
回复

@sid_625 Real bar here: infra that makes agents wait defeats point of having agents. 30ms + eBPF egress was the fun problem, prewarmed snapshots, cgroup tuning, per-VM egress enforced from outside guest kernel so compromised agent can't route around it, only through it. Credits live, no card. Send us workload that breaks it, fastest way we get better.

2
回复

been using this from last couple of weeks and the best thing about this is that its insanely fast, agents can spin up the sandbox in sub-milliseconds
great work by team

5
回复

@saurra3h Glad it's working well for you. Sub-millisecond's a stretch, but most agent loops never notice the gap. Thanks for writing this up.

0
回复

Congrats on your product and the launch. This sounds interesting. But how does it compare to Docker or Firecracker?

4
回复

@jn263 Firecracker's not a competitor, it's the engine underneath. CreateOS Sandbox runs on Firecracker microVMs, then adds the control plane around it: scheduling across hosts, an HTTP API, pause/resume with snapshot, per-sandbox networking and egress rules, SSH tunnels, disk mounts. Raw Firecracker gives you the VM primitive. We give you the fleet management layer teams actually need to run agent workloads at scale.

Docker's the real comparison point. Docker containers share the host kernel, isolation runs on namespaces and cgroups. A kernel exploit in one container can reach others on the same host. Each CreateOS sandbox gets its own kernel via KVM hardware virtualization, so the isolation boundary sits at the hypervisor, not the kernel. Boot still lands in the milliseconds, close enough to container speed that most workloads don't feel the difference, but you get VM-grade isolation for untrusted or agent-generated code.Here, what I can see is that the features on top are real enablers, like:

  • Bring your own S3

  • Sandbox

  • Mesh VPN

  • Editor support

  • Egress control

  • Asynchronous file sync

  • Synchronous file sync

  • etc.

1
回复

Hardware-isolated sandboxes are a strong direction for agent work. The practical question I always look for is whether the environment makes rollback, artifacts, network access, and final-state evidence obvious enough that a small team can trust the result without babysitting it.

3
回复

@krekeltronics That I consider application layer not the infra layer but you can use fork smartly and figure it out

0
回复

30ms startup is great, but for agent workloads I'd also want a teardown receipt: which egress rules fired, what credentials were present, and whether the writable layer was actually destroyed. Is that available per sandbox today?

3
回复

@new_user___2672025cf1bc18102609b53 Good question, and honestly the answer is "partially" today. Every destroy emits an audit event (`sandbox.destroy`, with cause user/admin/TTL/drain) and egress rules are logged when configured (`sandbox.egress.set`, allowlist only, no raw secrets). Credentials never round-trip through logs, disk creds are ECDH+ChaCha20-encrypted at rest and SSH keys are audited by fingerprint only, never raw.

If we want the application layer to take care of the forensics part, else you can always visit audit logs

0
回复

The 30ms cold start is genuinely impressive for hardware isolation, basically unheard of in that category. Loving that you threw in the Claude plugins and SDK examples right out of the gate too, makes it way less of a chore to actually try it.

1
回复

@kuzeyvv1l Wanted trying it to be zero-ceremony, not "read docs for an hour first." Glad that came through. Let us know what you build next — always want the second thing you try, not just the first.

0
回复

The 30ms cold start actually held up when I spun up a few sandboxes back to back, no weird warm-up lag. Loved having the SDK examples handy instead of digging through docs.

1
回复

@sultanrdm6 That's deliberate — every shape's pre-warmed so there's no "first one's slow" tax. Glad the SDK examples were there instead of you spelunking through docs mid-test.

0
回复

The 30ms cold start actually held up in my testing, which is wild for hardware isolated sandboxes. The claude plugin integration made spinning up agents feel almost frictionless.

1
回复

@azadxpsa That was the point of shipping the plugin day one — agent loops shouldn't need a side quest just to get a sandbox. Glad it landed that way for you.

0
回复

The 30ms cold start is genuinely impressive for hardware isolated sandboxes. Whoever tuned that cold boot path clearly obsessed over it.

1
回复

@tunahancangelir Correctly clocked. Prewarmed snapshots + a dedicated boot lane so new-VM spawns don't queue behind everything else — took a while to get the tail latency down, not just the median. Good to hear it holds up outside our own benchmarks.

0
回复

The ~30ms cold start is genuinely impressive, especially for agent loops where latency compounds. One thing that would save me a lot of time: a local dev mode that spins up a fake sandbox emulator so I can iterate on agent logic and exception handling without burning real compute credits during testing.

1
回复

@feyzaqk5q Real pain point, makes sense especially for exception-path iteration where you don't want real infra in the loop. 500 free credits take some of the sting out short-term, but a proper local emulator is a different, better answer. Adding it to the list.

0
回复

The 30ms cold start is genuinely impressive, especially the p90 timing. One thing that would help me as a builder is a built-in state diff or snapshot viewer in the CLI. Being able to compare filesystem and env changes between two sandbox runs without having to write custom diff scripts would make iterating on agent behavior much faster. Maybe something like `sandbox diff run-123 run-124` that outputs a clean summary.

1
回复

@cihanomakgzs0 Good ask, love the specificity. Nothing built into the CLI for that today — closest workaround is pulling files via the file API before/after and diffing yourself, which is exactly the friction you're describing. `sandbox diff run-123 run-124` is a clean shape for it. Noting this one, thanks for the concrete syntax idea.

0
回复

Super useful for a whole host of scenarios!

Curious if you plan to give more integration options for the plug-n-play style agents? Looks like currently we just would connect / control with Telegram?

1
回复

@inferhaven Today the real integration surface is the SDK (Go/Python/TS), the CLI, and a Claude Code plugin — Telegram's just been the flashiest demo, not the ceiling. More triggers/connectors are on the list. What would unlock the most for your setup — a specific chat platform, webhooks, something else?

1
回复

Spun up a sandbox in around 30ms like they claim, the CLI felt snappy and the SDK examples actually made sense for once.

1
回复

@emircansowj Appreciate the backhanded compliment to the rest of the industry's docs. Glad it actually clicked instead of fighting you.

0
回复

A built-in cost dashboard would be super helpful, especially showing compute time per sandbox session and monthly burn. With cold starts at 30ms there's a real risk of accidentally spinning up thousands of sandboxes during testing, and right now there's no easy way to see what's running or set spend alerts.

1
回复

@yarenaralp Fair, and honestly a real risk we think about too — 30ms cold starts make it trivially easy to fire off way more sandboxes than you meant to. Today you can list running sandboxes and check `/bandwidth` per sandbox, but there's no aggregated spend view or alerting yet. Logging this as a real gap, not just a nice-to-have.

0
回复

Spun up a sandbox in about 30 seconds and the hardware isolation gave me real peace of mind for running untrusted agent code. The CLI felt snappy and the SDK examples made it easy to wire into my existing flow without much fuss.

1
回复

@tugayabaylungn This is exactly the reaction we built for — untrusted code, real kernel boundary, no crossed fingers. Glad the SDK examples got you moving without a fight. Curious what you end up wiring it into.

0
回复

30ms p90 for hardware-isolated spin-up is a serious number — most sandbox stacks I've tried are an order of magnitude slower, and it changes what you can do per-task vs per-session. Curious about the lifecycle model: when an agent needs state to persist across runs (a working directory it comes back to tomorrow), do you snapshot/restore, or is the intended pattern ephemeral-always with external storage? Building long-running agents, that's the decision that shapes everything downstream.

0
回复

i run coding agents most of the day and the thing that always makes me nervous is what they can touch. hardware isolation with egress enforced inside the sandbox feels like the right way round. does the 30ms spin up hold once you're pulling real dependencies in, or is that a bare box number?

0
回复

A native MCP server for the sandbox so agents can spin up isolated environments on the fly without bolting on extra glue code. Would make the CLI and SDK even more plug and play for Claude and other MCP clients.

0
回复
#9
Skim
Free, open-source AI email client for Windows
164
一句话介绍:Skim 是一款基于本地优先理念、体积仅 5MB 的开源 Windows 邮件客户端,通过 Rust 极简内核与可控的本地 AI(自带 API Key)辅助回复/摘要,解决了用户对传统邮件客户端臃肿、启动慢、隐私侵犯的痛点。
Windows Email Open Source GitHub
邮件客户端 开源软件 Rust 本地优先 AI辅助 Tauri Windows 极简 隐私保护 离线可用
用户评论摘要:用户普遍称赞其极小的体积(5MB)和亚秒级启动速度。主要建议集中于:支持多账户(开发者已计划添加)、统一收件箱视图(开发者已响应并着手设计)、PGP/S/MIME加密签名、以及“稍后提醒”功能(开发者对此持谨慎态度)。也有部分用户认为缺少过滤和暂停功能显得功能裁剪过度。
AI 锐评

Skim 在“反 Outlook/Thunderbird 审美疲劳”的浪潮中精准切中了一个小众但忠实的需求:开发者群体的邮箱洁癖。它的真正价值并非在于“AI”,而在于“极致减法”带来的那种Windows生态中久违的清爽感。

先说亮点:用 Rust + Tauri 2 做到 5MB 安装包、瞬间冷启动,这本身就是对底层技术选型的自信,也是直接打击 Electron 类产品用户痛点的利器。BYOK(自带密钥)的AI模式,本质上是一种“功能合规套件”——它把数据隐私的道德责任甩给了用户自己,却换来了开发者自身的免责和软件的永久免费,这是极高明的商业策略,也是开源社区喜欢的交易方式。

然而,这种“极简”也是一柄双刃剑。从评论中可以看出,Skim 目前仍然缺乏多账户、统一收件箱、邮件加密、以及大多数主流邮件客户端标配的邮件过滤/规则系统。社区对“稍后提醒”的需求恰恰说明,用户需要的不是极简,而是“极简但聪明”。开发者在 GitHub 上偏哲学化的“反对”回应固然有个性,但若一味沉溺于自己的认知堡垒,很容易陷入“少数品味自嗨”的陷阱,最终被那些虽然臃肿但足够强大的客户端反向收割。

另外,依赖用户自行提供API Key才能使用AI,实际上将大部分轻度用户的体验门槛抬高了。AI 在这里更像一个“发烧友彩蛋”,而非核心卖点。如果 Skim 无法在未来快速补齐多账户管理和邮件过滤等关键基础功能,它可能会永远停留在“开发者的另一个玩具”的层面,而无法成为真正的生产力工具。噱头与实用之间,还得走好平衡木。

查看原始信息
Skim
Skim is a free, open-source (MIT) email client for Windows. ✂️ Minimal on purpose: no calendar, no rules engine, no bloat. Fewer buttons, less cognitive load. ⚡ Native Rust core (Tauri 2): ~5 MB installer, sub-second cold start, instant full-text search. 🔒 Local-first, offline-ready, zero telemetry. ✦ AI on your terms: bring your own Anthropic or OpenRouter key and Skim drafts replies in your voice and answers questions across your whole mailbox.
I built Skim because I wanted a Windows email client that doesn't suck that hard and couldn't find one. Now I wanna share it with you. Skim is open source and MIT-licensed, free forever. No subscription, no telemetry. Join as a contributor plz! 🙏 https://github.com/nikserg/skim Skim is small: a native Rust core (Tauri 2, no Electron), a ~5 MB installer, offline-first IMAP sync into local SQLite, instant full-text search, and a keyboard-first UI with a Ctrl+K palette. ✦ The AI part works on your terms: paste your own Anthropic or OpenRouter key to enable AI features: ✍️ it drafts replies in your voice (and can learn from your sent emails) 📥 summarizes your unread mail 🔍 answers questions across your whole mailbox, with cited sources Pretty basic, huh? Right. That's all I need, and I believe there are people who need the same. Requests go straight from your machine to the provider: no middleman, no markup, no subscription. Let me share a bit of the philosophy I try to follow while building Skim: ✂️ Minimalism. It's an email client and that's it. No integrations -> no bloat. Cognitive load matters! 🎯 Contextual actions. Skim tries to be the kind of software that thinks FOR you to save your cognitive resources. Buttons show up only when needed, the app always does the most obvious thing, and the defaults should be the best ones. ⚡ Resource efficiency. Skim is built to run in the background 24/7, so it stays as humble as possible and eats as few resources as it can. That's why the installer is under 5 MB 😄 So, with all that said: please, have fun with Skim! I'd love your feedback, especially from folks tired of heavy clients. What would make you switch? What's missing? How could I trick you into contributing? 🤘
5
回复

@nikita_zarubin I've migrated from Windows this year, but upvoted for branding only - great job, I like the anti-boring approach ;)

1
回复

@nikita_zarubin Congrats on the launch! LOVE the website - you clearly put a lot of thought and personality into it.

0
回复

I will take a look because anything has the be better than MS Outlook these days! Does it handle multiple email accounts?

1
回复

@siridley Hey, my thoughts exactly 😄

No multi-account support rn. I thought it wasn't needed, but just yesterday I found myself in a situation where I actually did!

Soooo I'll probably add multi-account support in the next release or so. There are usually 2-3 releases a day btw 😄

2
回复
No rules, no filter, no snooze… sounds like you cut away a bit too many features…
0
回复

love that it's local-first and tiny, but honestly a unified inbox view would be a game changer for me. like, one screen where i can see all my accounts side by side instead of toggling between them. would make it feel even more minimal and way faster to triage.

0
回复

@rem1219681 oh this one I love. You actually nudged me over the line - unified inbox is happening.

The plan: all your mailboxes merged into one view by default (not just the inbox - sent, archive, everything), Each message tagged with a little colored dot + letter so you always know which account it's from. And if you ever want the old one-account-at-a-time mode, it stays a toggle.


Wrote it all up here, design + open questions - would love your take:

https://github.com/nikserg/skim/issues/10


Thanks for the push 🙌

0
回复

Honestly the 5 MB installer is kind of wild for an email client, opened it up and was reading mail before my coffee was ready. The bring your own AI key approach feels right too, no weird data stuff happening behind the scenes.

0
回复

@kevserauj4 glad you liked it! 🫶

0
回复

Love that you kept it under 5 MB while still shipping full-text search and AI drafting. The local-first approach with bring-your-own-key feels rare these days.

0
回复

@hmeyrazbakmtsy glad you like it! Enjoy! 🫶

0
回复

Love how lightweight this feels, especially coming from bloated desktop clients. One thing that would make it a daily driver for me: add support for PGP or at least S/MIME signing and encryption. Local-first is great, but most of the work I get requires signed replies, so I keep having to fall back to Thunderbird for that.

0
回复

@sametgnelwoiv hey, glad you liked it!

You bring up an interesting topic. Crypto is tough. PGP is probably out of scope, but S/MIME not so out.

I've created an issue with my thoughts on that:

https://github.com/nikserg/skim/issues/9

Would really appreciate your participation in that! 🙏

0
回复

5mb installer and sub-second cold start on windows is honestly kind of wild, my current client takes forever to open. love that ai is opt-in too.

0
回复

@beratmcdecotaw hope you're enjoying it! 🫶

0
回复

been using it for a few days and honestly the instant full-text search is kind of wild, way faster than outlook. the minimal vibe is actually refreshing, feels like it respects my time

0
回复

@arifpekgila0cj Skim respects what Outlook obviously doesn't 😄

0
回复

Sub-second cold start from a 5 MB installer is genuinely impressive craft, especially with native search baked in. Really respect the discipline of cutting the calendar and rules engine instead of bolting them on.

0
回复

@rukiyevatad9da greet Claude that made it possible this way 😄 When code is cheap, discipline becomes expensive.

0
回复

Tried it on my old laptop and yeah, the cold start is genuinely instant which is wild for an email client. Love that the AI stuff is opt-in instead of shoved in my face.

0
回复

@sinemimirc3hz glad you liked it! Skim is dinosaur-laptop-friendly 🦖

0
回复

Love that you kept the scope tight enough to ship a 5MB native Rust client that still feels full-featured. Bringing your own AI key is a smart move, treats the model layer like infrastructure instead of locking people in.

0
回复

@satd7og my thoughts exactly 😄

0
回复

The 5MB Rust install is genuinely impressive, and it actually opens faster than Outlook does. Curious to see how the BYOK AI drafting holds up on longer threads.

0
回复

@eymenevf2 comparing to Outlook it opens instantly 😄

As for drafting - I personally use it all the time, it keeps context well as I can see. But I'm curious to learn your use cases and experience to make things better. Do not hesitate to open a Github issue or text me here.

0
回复

finally a windows email client that doesn't try to be a whole operating system, love that the rust core keeps it to like 5mb

0
回复

@kamilsatlkebkj Me too, man! 🤘 Enjoy!

0
回复

Love the minimal take, especially the Rust core keeping things snappy. One thing that would make it stickier for me: a quick "snooze until later today" or "remind me tomorrow" option right in the inbox. Most of my email backlog isn't urgent, it just needs to come back at the right moment, and that small loop would make Skim feel genuinely complete without adding much weight.

0
回复

@sevda9qob Ouch, that's a hard one — a highly debatable topic. As for now, I'm closer to not implementing snooze than to implementing it.

However, there's a "canonical" thread about snooze, where my argumentation against the feature lives:

https://github.com/nikserg/skim/issues/8

But. But! Please do argue.

0
回复

love how lean this is, the rust footprint is wild. one thing that would sell me even more is proper vim keybindings for the inbox and reading panes, j/k navigation feels like a natural fit for something this minimal.

0
回复

@yamurobancsyyp Yeah, j/k navigation works in inbox!

0
回复

Honestly the local-first angle is super appealing, and a tiny Rust installer sounds great. One thing that would seal the deal for me though, kind of a small ask, is per-account send aliases so I can manage work and personal replies without juggling separate logins. Would fit nicely with the minimal vibe too.

0
回复

@lyaszkvraknr2o hey, looks like a solid feature! Tbh, I don't use aliases myself, so I'd love to understand the workflow. I'd super appreciate if you'll share more details in Github issue that I created for your request:

https://github.com/nikserg/skim/issues/7

A few things that'd help me nail the use case (feel free to add comment in Github issue):

  1. What's your setup, one email provider (e.g. a Gmail with a couple of addresses attached), or genuinely separate ones like a work domain + a personal Gmail?

  2. Do work and personal mail land in the same inbox for you, or would you rather keep them separate but just share one "send" experience?

  3. What's missing in the current multi-account switcher for your case, is it the extra login, mixing the inboxes, or specifically picking who the reply comes from?

  4. When you reply, should the app auto-pick the right "From" based on which address the original email was sent to?

0
回复

finally an email client that respects my storage and my sanity, the 5mb install is unreal on windows. love that i can just plug in my own api key and skip the data-sucking defaults.

0
回复

@asyadizili Enjoy! 🫶

0
回复

love that this is BYOK instead of yet another subscription tier, that's rare for an email client. since the AI can answer questions across the whole mailbox, does it send full email bodies to the API per query or is there some local pre-filtering first so you're not shipping your entire inbox history every time you ask something?

0
回复

@omri_ben_shoham1 nice one! And no, it doesn't ship your inbox on every query. Retrieval is local first.

Under the hood the AI is a tool-calling agent. When you ask something, it searches your mailbox on your machine. What crosses the wire from that search isn't email bodies, it's compact rows: date, sender, subject, and a snippet capped at ~160 chars. That's the pre-filtering step you're hoping for: the full-text index narrows thousands of emails down to a handful before anything touches the network.

Full bodies go out only for the specific emails the model then chooses to open - and even those are truncated (roughly 6k chars per email, ~10k for a whole thread). Per query it's bounded to about a dozen reads, not your entire history. Trash and junk are excluded unless you explicitly ask for them.

0
回复

Do you have plans for Mac or Linux, or is Windows the main focus? The ~5MB installer and sub-second startup would be a pretty strong selling point for anyone tired of bloated Electron apps, so I'm wondering if expanding makes sense or if you want to keep it tight and focused.

0
回复

@talhakhalidmtk I'd rather focus on Windows, but not because it's against product philosophy, but because I use only Windows 😄 If you want to add a builder for another OS - feel free to contribute!

1
回复

Hey Nikita, congrats!
Any plans to release a Mac version?

0
回复

@luis_parker thanks! I don't use Mac, but there's no obstacles to have such version. It would require new building pipeline and some settings. I invite you to contribute 😄

1
回复

The Rust-backed cold start is genuinely instant, feels closer to opening a text editor than an email app. Glad the AI part is opt-in with your own key instead of some bundled subscription, keeps things lean.

0
回复

@nilferkofag6jt True! 🤘

0
回复

Curious how you handle Gmail OAuth secrets in a public repo though! Great launch, refreshing branding, and love every MIT license.

0
回复

@aidan_codefox Hey, I'm flattered you love the branding! It's something I'm secretly proud of. At first I went the usual route - you know, purple, polished, sterile shit - but then I thought, "Hey, it's free and open source, right? Fuck it, let's have fun," and ended up with this wonderful zine landing you're looking at.

Okay, flex time's over - to your question.

tl;dr: there are no secrets in the repo. The magic is called PKCE.

Brace for the AI-slop explanation, if you're curious:

  1. Gmail uses the installed-app flow (loopback + PKCE), where the "client secret" isn't actually a secret. Google's own docs say the client secret for installed apps "is obviously not treated as a secret" — a distributed desktop app can't keep one confidential, so security comes from PKCE (S256) and the 127.0.0.1 loopback redirect, not from a secret. There's no server in the middle; the token exchange goes straight from your machine to Google.

  2. The client ID/secret aren't committed — they're baked in at build time from env vars (SKIM_GOOGLE_CLIENT_ID / _SECRET), injected in CI from GitHub Actions secrets. So official installers carry them, but nothing sensitive sits in the source tree. Build from source and you just plug in your own Google Cloud project.

  3. Microsoft/Outlook is a true public client — PKCE only, no secret at all.

  4. Per-user tokens never touch disk in plaintext. Refresh tokens, app passwords, and the AI API key live in the Windows Credential Manager — never in the SQLite DB or any config file.

So "public repo" and "OAuth" coexist fine here: the only thing that would be dangerous to leak (per-user refresh tokens) never leaves your OS keychain, and the app-level identifiers are the kind Google/Microsoft explicitly design to ship inside native binaries. 🔒

1
回复
#10
Bolna Agent Studio
Build Voice AI Agent in 10 Minutes
145
一句话介绍:Bolna Agent Studio 允许企业仅需上传文档或回答几个问题,在10分钟内无需提示工程即可构建并部署可投入生产的语音AI代理,彻底解决了传统语音代理构建周期长、调试复杂、易出错的痛点。
SaaS Artificial Intelligence No-Code
语音AI代理 无代码开发 代理工作室 快速部署 企业级语音 多语言支持 自动化对话 生产级对话 对话式AI 全渠道语音
用户评论摘要:用户高度赞扬其易用性和多语言自然度(如印地语、西班牙语),与竞品对比显著节省时间。主要问题包括:如何保证代理独特性而不千篇一律?如何处理文档未覆盖的边界情况(如自动拒绝与转接)?医疗等敏感行业是否支持BAA与合作合规?有无沙盒测试环境?
AI 锐评

Bolna Agent Studio 的核心价值并非“低门槛”,而是“工业化的经验固化”。它用20万小时通话数据提炼出的模块化模板,实际上是把过去需要资深工程师手动调试的“隐性知识”——身份设定、边界处理、模型组合——变成了可复用的标准化积木,这才是它敢宣称“上线当天即可投产”的底气。

但真正的软肋藏在看似最强的部分:用模板组装会导致代理趋于同质化,尤其在竞争激烈的垂直领域(如客服、销售),企业最怕的就是和对手用同一个声音套路。更关键的是,对于医疗、金融等受强监管场景,合规是刚需而非选项。评论中关于“文档未覆盖内容导致模型自信生成错误答案”的质疑非常精准——如果Studio无法自动插入“不知道”或“转接”的硬性拦截,那它仍然只是一个高级对话引擎,而非真正的企业级代理。此外,目前缺乏明确的沙盒/红队测试环境的说明,让“上线后迭代”显得更像一种项目风险的转嫁。

一句话总结:这是一个将语音代理从“手工作坊”推向“预制菜”的关键一步,但它离“米其林后厨”还有很长的路要走——尤其是当客户需要定制风味和合规保险的时候。

查看原始信息
Bolna Agent Studio
Bolna's Agent Studio lets any business build and deploy a production-grade voice AI agent in minutes, no prompt engineering. Upload a doc or answer a few guided questions, and Studio assembles a call-ready agent from production-tested modules. Sign up and start calling on Day 0.

Hey Product Hunt 👋

We are the team behind Bolna, and we are here to make the whole process of building a Voice AI agent a whole lot easier and quicker than ever before.

Let’s be honest! The starting point to build a Voice AI agent has always been intimidating. Hours of effort gathering context, foolproof guardrails, detailed conversation flows, and yet something sifts through the whole orchestration, leading the agent to embarrassingly fail at deployment.

Agent Studio on Bolna takes all the above hindrances away between you and your perfect Voice AI agent by specifically putting every piece of the matrix together with one drag-and-drop brief. All you have to do is upload a contextual document and let it fill in the blanks to deliver a production-grade voice agent which can go live the same day as you sign up.

What makes Agent Studio handy for Voice AI agents:

  • No prompt engineering expertise to get started.

  • A modular structure in place including Identity, Conversation and Closing blocks.

  • Built on thousands of proven templates applied across sectors.

  • Handles all edge cases, instead of just linear scripts.

  • Auto-picks the best combination of models for each use case.

  • Warm endings and fallbacks for smooth user experiences.

  • Reviews every agent across quality benchmarks and fills in whatever is missing before shipping the agent.

This changes the way Voice AI agents were fundamentally built, taking weeks to curate the perfect prompt and deployment. Agent Studio makes your perfect Voice AI agent come to life with a single session, delivering a production-grade deployment that holds up through real calls for enterprise scale. The barrier to building and scaling with voice has now dropped to almost nothing.

Presenting this to the Product Hunt community is huge for us, with weeks of deliberation gone into each step of the process. We’d love for you to give Agent Studio a shot to build Voice AI agents. Go all out with your use cases on Agent Studio, and we are sure you’d adopt it as your core Voice AI stack! And if you end up loving it as much as our beta users have, we’d love your support here!

Every upvote and comment means a lot to the team.

So what's your excuse for not building a Voice AI Agent anymore?

Cheers,

Sonam from Bolna

P.S. We'll be around to answer questions throughout the day, so shoot all your questions right away!

11
回复

@sonambala What kind of use cases it can cover can I make a BFSI agent?

0
回复

Hey Product Hunt 👋

For the longest time, building a Voice AI agent at Bolna meant one of us sitting down and building it by hand.

I have spent tens of hours building agents, any my team, who is building agents for some of India's largest enterprises, has probably spent 100x that.

Every new agent meant hours of gathering context, writing guardrails, mapping conversation flows, and second-guessing every edge case. We "Did Things that Didn't Scale" until we just physically couldn't. There was too much of a gap between the quality and speed of agent building required and what we could provide (even as our FDE team grew from 2 to 20 in 3 months)

Then somewhere past 200K+ hours of calls, a pattern emerged. The same building blocks. The same edge cases. The same fixes we kept making by hand, over and over.

Agent Studio is us finally putting all of that back into the product itself.

You upload a context document. Drag and drop a brief. And it -

  1. Fills in the blanks using thousands of patterns proven across sectors

  2. Structures the agent into Identity, Conversation and Closing blocks

  3. Auto-picks the best combination of models for your use case

  4. Handles edge cases, warm endings and fallbacks, not just linear scripts

  5. Reviews itself against quality benchmarks and fills whatever's missing before shipping

Sign up in the morning, have a production-grade agent taking real calls by evening. No more prompting frustratiosn.

The part I keep coming back to - the hardest thing about Bolna's first two years was the manual work. Agent Studio is us handing that exact superpower to anyone who signs up.

The barrier to building with Voice has dropped to almost nothing.

We'd love for you to break it, push it, throw your weirdest use case at it. Give us your thoughts!

Super proud of @xan_ps and the entire team for building and shipping this!

So - what's stopping you from shipping a Voice AI agent today?

7
回复

Back on PH after a long time and there's a reason! While building out voice agents I was always looking for ways I could vibe-create an agent and just talk to it. While Claude solved for building apps for custom use cases by just explaining in English, I found an equivalent for the Voice AI sector here!

Loving it through and through! But have a few suggestions. How can I share those, team?

3
回复

@siddharth_batra Vibe-creating an agent is definitely one way to put it! XD
Would love to hear more of your experience with Agent Studio, why don't you write to us at hello@bolna.ai?

0
回复
Finally moving away from Claude projects which draft emojis in the prompt to contextualised agent which works :)
3
回复

@sarvagya_chhabra Super glad that you found it useful enough to drag you away from Claude. Would love to know what voice AI agents you have built using Agent Studio!

0
回复

insane !!!

it made life easier to write agents on the platform itself , now i can easily make agents for different usecases from different industries with Agent Studio within just 5 mins of uploading few documents and context

2
回复

@sanket_bhat_080802 Glad you found it easy to build agents on Agent Studio. Curious to know what use cases you are deploying voice agents for.

0
回复
Voice AI usually means wiring together ASR, LLM, and TTS yourself before you've even built the actual conversation. Studio hides all that; you describe the call and it's live.Best part is watching non-technical folks like ops, marketing ship a working agent
2
回复

@sharath_kuntanahal yeah, definitely! Agent Studio just makes it easy for anyone to build a voice agent on the go. You don't need to have any expertise whatsoever, just a brief and you're good to go.

0
回复

Finally got around to testing Bolna for an outbound campaign and was honestly surprised how natural the voice sounded in Hindi and Tamil. Setup took maybe an afternoon.

1
回复

@nuran37cs Thank you for giving it a shot! Yes, achieving human-like simulation in Indian languages is at the core of Bolna. Would love to know what use case your agent served, though!

0
回复

Finally moving away from long iterations because chatgpt just doesn't understand how my voice agent is supposed to work! Glad to know it is trained on real data.

1
回复

Honestly the multilingual support surprised me most, our test calls in Hindi and Spanish both sounded pretty natural compared to other voice AI tools I've tried. Setup was straightforward too.

1
回复

@farukkorkuygpl Thank you! Bolna is centred on ensuring human-like pacing for conversations in all languages that we offer. So glad you found it seamless to test.

0
回复

The async setup was surprisingly painless for handling multilingual flows, and the latency on outbound calls felt close to real human pacing.

1
回复

@simge5mt4 Thanks for giving it a shot! Low latency and human-like conversation are actually at the core for us at Bolna for Voice AI deployments. Curious to know what languages you tested out, though?

0
回复

the "handles all edge cases, not just linear scripts" line is the part I'd want to poke at before pointing real call volume at it. since Studio can get you live the same day you sign up, is there a staging/sandbox step where you can run a batch of adversarial or unusual calls against the generated agent before it starts taking live traffic, or is the expectation that you go live and iterate on real calls from day one?

1
回复

@galdayan At Bolna, we have thought this through. You can actually test it out by chatting with the agent to refine it before taking it live to your target audience. This helps you evaluate the quality and build of your agent before they ever touch any of your customers.

0
回复

Maitreya, the 200K hours of calls turning into reusable blocks is the most convincing part of this story: that is the kind of pattern library a hand-built agent never benefits from.

My question comes from my corner of the market. I build AI for healthcare, and voice is the obvious next interface for appointment reminders and patient intake, but every call recording there contains protected health information.

Do you support healthcare use cases today: things like a BAA, redaction of recordings, or region-locked storage? That answer decides whether teams like mine can even prototype on Bolna.

1
回复

Looks super easy to use! Congratulations on the launch.

1
回复

@kritikapathak Thank you! Hope you had fun making your voice agent.

0
回复

Everyone ends up building similar-sounding voice agents if they're using the same templates. How do you keep each one feeling unique?

0
回复

The auto-assemble-from-a-doc part is the real leap here, and it's also where I'd look hardest. When we generated agent flows from source material, the thing that never came for free was the negative space: a caller asks something the doc just doesn't cover, and the model improvises a confident policy answer instead of saying 'I don't know, let me transfer.' Does Studio insert those refusal and escalation boundaries automatically, or is that still the part you tune by hand after the agent is generated?

0
回复

A real-time analytics dashboard showing live sentiment and call drop-off rates during campaigns would help teams tweak prompts on the fly. Right now we mostly wait until after to see what worked.

0
回复

Spent a few minutes poking around and the multilingual demos actually handled code-switching better than I expected. Pricing transparency was a nice surprise too, most voice AI vendors make you book a call just to see numbers.

0
回复

@brahimo9zc So glad that Bolna was able to surprise you pleasantly. What languages did you test out your agents for?

0
回复

Would love to see a built-in call quality analytics dashboard that breaks down latency, transcription accuracy, and drop-off points per language. Right now we're piecing it together from logs and it makes it hard to spot where the model is struggling in Hindi versus English versus Tamil. A single view that surfaces the metrics that matter most would save our team hours every week and help us fine-tune without guesswork.

0
回复

A dashboard view showing live call metrics with filters by region, language, and intent would be super helpful when running thousands of concurrent calls. Right now I'm guessing it's hard to spot issues at scale without granular real-time visibility into performance breakdowns.

0
回复

Would love to see a built-in call quality analytics dashboard showing latency, transcription accuracy, and drop-off rates per language. Right now it's hard to benchmark performance across different regional deployments without piecing together data from multiple sources.

0
回复

Most of the thread is on edge cases and compliance. The linguistic complexity claim is the one I'd poke at, we run voice on the support side in a lot of languages.

On real calls people don't stay in one language. They switch mid sentence, and the worst place is digits. Someone speaking English will still say an order number or a postcode in their own language, almost every time. Same when they get annoyed, then a whole sentence flips.

If language is picked at call start and pinned there, that is where transcription falls apart, and it falls apart on the field you least want wrong.

So does the agent detect the switch inside the call and follow it, or is it locked once the call starts? And are digits handled separately from the rest..

0
回复

@jernej_jan_kocica You'll be pleasantly surprised with Bolna on this. We do have agents with LID who follow the user's language through the conversation with code-switching. The digits are handled accordingly as well.

0
回复
#11
Manifest
Turn any webpage into an action manifest for AI agents
144
一句话介绍:Manifest 将任意网页转化为AI代理可操作的JSON动作图谱,通过一个API调用返回可点击、可填充、可提交的元素及其依赖关系,解决传统浏览器代理因选择器脆弱而频繁失效的痛点。
API Developer Tools Artificial Intelligence
AI代理工具 网页自动化 DOM解析 动作图谱 浏览器Agent 元素依赖 Playwright LangChain集成 MCP服务 无代码选择器
用户评论摘要:用户高度认可“requires”字段解决元素依赖问题,避免代理无序失败。主要建议:增加结构变更检测(指纹/ETag)、缓存与快照差异对比、记录重放用于回归测试、动态内容与Shadow DOM支持待加强。
AI 锐评

Manifest切中了一个长期被忽视的刚需:AI代理在网页上不是“看”内容,而是“操作”控件。传统方案要么给浏览器(Playwright/Selenium),要么给内容(HTML/text),但没人给“操作语义”——哪些按钮可点、哪些字段必填、点击后出现什么。Manifest用“requires”字段编码元素间的隐式依赖,相当于为网页预编译了一张操作流程图,这正是Aria快照和选择器字符串从未解决的问题。

但冷静看,它的核心壁垒并非技术难仿,而是先发优势与场景绑定。当前版本依赖LLM+Playwright实时提取,延迟与成本是硬伤——每次调用都跑一次完整解析(尽管有缓存),这对高频场景(如监控类Agent)并不友好。评论中大量提到的“diff/缓存/回放”诉求,本质是用户希望它从“单次查询”进化为“持续追踪”,而这需要维护状态机与变更检测能力,与当前无状态API架构存在冲突。

另一个隐忧是生态位:Manifest目前是“代理的代理”,充当上游解析层。但LangChain、Playwright等框架完全可以在自身层级内置类似能力,一旦集成,独立产品生存空间会被挤压。创始人说“优先服务Agent开发者,再吸引网站接入”,但若网站不主动输出语义,Manifest始终在猜测闭门造车,准确率无法100%。

价值明确,但天花板也清晰。作为解决选择器脆弱的中间件,它在原型验证和小规模Agent工作流中可大幅降低编排痛苦,但想成为生产级别的标准化基础设施,还需要在缓存策略、变更检测、动态DOM覆盖上补齐能力,并回答一个终极问题:当AI Agent可以通过视觉直接理解页面时,这种从DOM到动作的显式映射是否仍是必要?

查看原始信息
Manifest
Manifest turns any webpage into a structured JSON map of what an AI agent can click, fill, and submit. One API call, no fragile selectors. Every action comes with resolved CSS/role locators, plus a requires field that encodes dependencies between elements (e.g. "select a plan before this button is clickable"). That's context aria snapshots don't give you. Python SDK, LangChain support, and an MCP server included. Built for anyone shipping browser agents.
Hey Product Hunt 👋 I'm Max, solo founder of Manifest (Omfang AB). I built this because every time I tried to get an AI agent to reliably interact with a webpage, I ended up hand-writing selectors that broke the moment a site's DOM changed. Browser automation tools solve "give me a browser," and content extraction tools solve "give me the content" — but nothing solved "tell me what's clickable, fillable, and submittable, and how those actions depend on each other." That's what Manifest does: one API call returns a structured JSON action manifest for any webpage, with resolved locators and a requires field that encodes cross-action dependencies (e.g. this submit button requires that field to be filled first). It's live now as REST API, Python SDK, LangChain integration, and an MCP server. You can try it with no signup here: demo.manifest.omfang.io/demo I'm building this solo and pre-revenue, so I'd genuinely love feedback, especially from anyone building browser agents who's hit the same selector-fragility problem. What would make this useful for what you're working on?
1
回复

That's the honest answer I'd rather have than an optimistic one — batch on structural change, don't re-call per field. The catch is knowing a structural change happened without paying for a full extraction to find out. Would a cheap fingerprint endpoint help here — one that returns just a hash of the interactive-elements set, so my agent compares cheaply and only triggers the full Playwright+Sonnet pass when that fingerprint moves? Even coarse, it'd let me avoid both reflexive re-calls and missing a new modal.

0
回复

The requires field encoding dependencies between elements is the interesting part — "select a plan before this button is clickable" is exactly the context accessibility trees drop, and it's where most browser agents burn their retries. Question: how do you handle drift? If the site ships a redesign, does the manifest regenerate per-call (so it's always fresh but you pay latency), or is there caching with some invalidation heuristic? That trade-off decides whether I'd trust it in a scheduled job.

0
回复

@max_nordstrom that's a genuinely honest answer, most tools would just quietly claim it handles all cases. the disabled/aria-disabled inference makes sense as the pragmatic v1 scope. for the async-JS-validation gap, would something like a lightweight 'confidence' flag on the requires field be feasible - even a rough signal telling the agent 'this dependency was inferred from static DOM, verify before relying on it' vs 'this one's a hard disabled attribute' - or would that just push the complexity onto the caller without really helping?

0
回复

i spend a lot of time watching agents guess their way around web pages and it is a mess, so this scratches a real itch. the requires field encoding dependencies is the clever bit. who publishes these first though, site owners or agent builders running it on other people's pages? feels like the chicken and egg question

0
回复

@terminal_candy 

Thanks — and yeah, that's the exact right question to ask, because it's the crux of whether this scales past a demo.

Short answer: agent builders first, out of necessity. Right now Manifest works on any page without the site owner doing anything — it's Playwright plus an LLM extraction step running against whatever's already rendered, so there's no cold-start problem on the consumption side. That's deliberate: the extraction layer had to work without any publisher buy-in or it never gets off the ground.

The site-owner side (a publisher SDK letting sites annotate their own semantics) is a later layer, not a launch requirement — and honestly it only makes sense once there's enough agent-side usage that a site owner has a reason to care about agent traffic at all. So the sequencing is: prove value to agent builders on the open web first, let that create the demand that eventually pulls publishers in, rather than trying to bootstrap both sides at once.

0
回复

the requires field for element dependencies is genuinely useful, something I've hacked together badly on past projects. excited to try it with an MCP server setup

0
回复

@hanmtanrkanwyt 

Thanks, Hanım — "hacked together badly" is a pretty universal experience with this problem, which is exactly why I wanted it baked into the response instead of something everyone has to rebuild themselves. Hope the MCP server setup makes it an easy first try — let me know how it goes.

0
回复

The dependency tracking between elements is genuinely useful, saved me from manually wiring up plan selection logic in my agent. Curious how it handles shadow DOM and dynamic SPAs though, that's usually where these tools fall apart for me.

0
回复

@emrekjvj 

Thanks — plan-selection logic is exactly the kind of thing requires is meant to save you from hand-wiring.

Fair question on shadow DOM and SPAs, and I'll be straight about it: extraction runs on a rendered Playwright page using the accessibility tree plus DOM queries, with a wait after load to let client-side rendering settle. That handles most SPA cases fine since it's working off the live rendered DOM, not raw HTML. Shadow DOM is the shakier case — standard querySelectorAll doesn't pierce shadow boundaries, so elements inside a closed (or even open, depending on how it's queried) shadow root can get missed. It's not a solved problem yet, more an honest gap in coverage right now.

0
回复

Finally a tool that captures the "why" behind clickable elements, not just the selectors. The requires dependency mapping is genuinely useful and saved me a ton of trial and error on a flaky login flow.

0
回复

@utkufqkz 

Thanks, Utku — that "why, not just the selector" framing is exactly what I was going for, since a bare selector tells you an element exists but not what has to happen before it's actually usable. Glad requires cut down the trial and error on a flaky login flow — that's a classic case where the ordering isn't obvious until something silently fails.

0
回复

honestly the requires field is the part that sold me, most selectors break for me when modals depend on prior actions. gave it a quick spin on a messy form and it nailed the dependency mapping first try.

0
回复

@alparslan5ths 

Thanks, Alparslan — that modal-depends-on-prior-action pattern is a really common way this stuff breaks, so glad requires caught it instead of you having to debug it after the fact. And nailing it on a messy form first try is a good sign — that's usually where the DOM signals get noisy enough to trip things up. Appreciate you giving it a real spin.

0
回复

the requires field for encoding element dependencies is genuinely clever, that's the kind of context aria snapshots miss and usually breaks agents. gonna test it on a gnarly multi-step checkout flow

0
回复

@diyars9zv 

Thanks — that's exactly the gap requires is meant to close, since raw aria snapshots give you what's on the page but not what order it has to happen in, and that's usually where agents quietly fail. A gnarly multi-step checkout is a good stress test for it — that's precisely the kind of flow where dependency ordering actually matters. Let me know how it holds up.

0
回复

The action graph idea is really clever, especially encoding those prerequisite relationships between elements. One thing that would help me a lot is a built-in diff or change-detection endpoint, so when a site updates its structure I can quickly see which locators broke and get suggestions for replacements instead of manually hunting through the JSON map. Would save a ton of debugging time in production.

0
回复

@rmeysaertuvujj 

Thanks — and that prerequisite encoding is the piece I spent the most time getting right, so glad it's landing.

The diff/change-detection ask is a real gap. Right now every /manifest call is a fresh snapshot with no memory of a prior state, so when a site's structure shifts, you find out by hunting through the JSON yourself rather than being told what broke. An endpoint that takes two manifests and surfaces the actual delta — locators that no longer resolve, actions that disappeared, requires that shifted — would turn that into a quick lookup instead of a manual diff. The "suggested replacement" part is the harder half (matching an old locator to its likely new equivalent isn't always clean), but even the plain diff alone would save the debugging time you're describing. Noting this as a real direction.

0
回复

Really cool approach, the requires field for dependencies is a smart solve. One thing that would make this even more useful for us is a built-in way to record and replay the exact sequence of agent actions against a stored map, so we can diff when a site changes and pinpoint exactly which locators broke. Would save a ton of debugging when a flow stops working after a deploy.

0
回复

@emelkorkune7b0 

Thanks — and that's a sharper version of a gap that's come up a few times here: right now every /manifest call is a stateless snapshot with no memory of a prior run, so there's no built-in way to say "here's the sequence that worked, tell me what changed."

A record-and-replay layer — capturing the sequence of manifests an agent actually used for a flow, then diffing a fresh run against that stored baseline — would turn "the flow broke after a deploy" into a direct answer (this locator no longer resolves, this action disappeared) instead of a manual hunt through logs. That's genuinely more useful than a plain diff endpoint since it's tied to the actual path your agent took, not just any two arbitrary snapshots. Noting it as a real direction, closer to a CI/regression feature than a runtime tweak.

0
回复

the requires field is honestly the standout for me, it captures context i've been missing when agents just blow past dependencies and fail silently. one api call returning resolved locators plus that semantic info is way cleaner than juggling aria snapshots.

0
回复

@aligndoduz2fn 
Thanks — "fails silently" is exactly the failure mode requires is meant to prevent, since an agent blowing past a dependency doesn't even know it did anything wrong until something downstream breaks. Glad the single call is landing cleaner than juggling raw aria snapshots yourself — that consolidation (semantics plus resolved locators in one response) was the whole point of not making you stitch it together on your end.

0
回复

The MCP server integration is a nice touch for getting started quickly. One thing that would really help in production is some kind of caching layer for repeated calls to the same URL, since right now it seems like every API request would re-scan the whole page. Even just an optional ETag-style header so we can skip re-mapping when the page structure hasn't changed would save a ton of tokens and latency when agents are looping through tasks on stable sites.

0
回复
@erafettinrqqxz Thanks, Şerafettin — good news, most of this is already in place: manifests are cached (TTL-based) under the hood, so repeat calls on the same URL within that window skip the full re-scan rather than paying for it every time. The ETag angle is the sharper idea though — right now the cache is time-based, not change-based, so you’re either trusting the TTL or eating a re-scan even if nothing moved. A way to signal “structure hasn’t changed, skip re-mapping” would be more precise than a flat TTL, especially for stable sites where an agent’s looping through tasks for a while. Noting that as a real improvement over what’s there today.
0
回复

Finally something that gets rid of the brittle XPath mess my scraper has been held together with. The requires field is a nice touch, saved me from another broken dropdown flow.

0
回复

@emin6zgn 

Thanks — that XPath brittleness is exactly the kind of thing I wanted this to replace, since a selector that snaps the moment a site tweaks its layout is a maintenance tax nobody should have to keep paying. Glad requires caught a dropdown flow before it broke on you instead of after.

0
回复

Really useful for browser agent work, the requires field alone is a nice touch. One thing I'd love is a built-in way to test the generated map against an actual page session, so you can catch mismatches between the JSON output and what the DOM actually looks like before your agent runs.

0
回复

@can4ojy 
Thanks Can! Good ask, and a real gap. Right now the manifest is generated once and handed over as-is, so there's no built-in check that it still matches the live DOM by the time your agent uses it, if the page shifted between the call and the run, you find out mid-task instead of before. A validation step that replays the manifest against a fresh session and flags mismatches (locator no longer resolves, action missing, etc.) before the agent commits to it would catch exactly that class of bug. Noting it as a real direction.

0
回复

finally something that gets at the real pain point. the requires field encoding dependencies is genuinely clever, going to wire this into our langchain setup this week and see if it holds up on a few gnarly sites.

0
回复

@minekld0 

Thanks, Mine — glad it's landing on the actual pain point rather than the surface-level stuff. The LangChain integration should make wiring it in pretty straightforward, but I'd genuinely like to hear how it holds up on the gnarly sites — that's exactly the kind of real-world stress test that surfaces where requires breaks down. Let me know how it goes.

0
回复

the requires field is genuinely useful, been burned by agents clicking submit before the form was complete way too many times. one thing id love to see is a way to mark elements that trigger async loads, like a dropdown that fetches options after you click it. right now my agents try to interact before the data is there

0
回复

@ecebedk3yvp 

Thanks, Ece — that "clicking submit too early" failure is exactly the class of bug requires is meant to prevent.

The async-load case is a good catch though, and a bit different from ordering: it's not that the dropdown depends on another action, it's that the dropdown's own options aren't populated yet when the manifest is captured, so your agent has correct ordering but stale data. Flagging elements that trigger a fetch-on-interact (and maybe hinting the agent should wait/re-resolve after clicking) would close that gap. Noting it as its own case rather than folding it into requires.

0
回复

honestly the requires field sounds super useful, would love to see some kind of caching layer for repeat calls on the same URL so agents aren't paying the extraction cost every single time

0
回复

@emircandemtyvr 

Thanks, Emircan — good news, that part's already in place: manifests are cached (TTL-based) under the hood, so repeat calls on the same URL within that window skip the full extraction cost instead of re-running it every time. Appreciate the ask lining up with something already shipped.

0
回复

Tried it on a gnarly multi-step checkout flow and the dependency graph actually caught a hidden "select plan before continue" rule I had been missing. The resolved CSS plus role locators side-by-side saved me a ton of selector soup.

0
回复

@masalyxsz 

Thanks, Masal — that's a great real-world case, and exactly the kind of hidden rule requires is meant to surface, since "select plan before continue" is the sort of thing that's obvious to a human eye but easy to miss when you're just working off a flat element list. Glad the CSS/role locator pairing cut down the selector soup too — having both was meant to give you a fallback when one breaks rather than betting everything on a single brittle selector.

0
回复

The dependency graph is a really smart touch since most selector tools miss those relationships entirely. One thing I'd love to see is a built-in diff endpoint that compares two Manifest snapshots over time, so when a site silently changes its checkout flow the agent can detect the break before it fails mid-task.

0
回复

@mahiry1rs 

Thanks, Mahir — and that's a sharp distinction from what most selector tools give you, since they don't even try to encode the relationships in the first place.

The diff-endpoint idea is a good one: right now every /manifest call is just a fresh snapshot with no memory of what came before, so a silently changed checkout flow only shows up when the agent fails mid-task, not before. An endpoint that takes two manifests and returns what actually changed — actions added/removed, requires shifted, locators broken — would let an agent (or a CI check) catch that drift proactively instead of finding out the hard way in production.

Noting it as a real direction, not just a nice-to-have.

0
回复

finally something that gets the dependency thing right, you know? the requires field saving me from writing all that state-checking logic is honestly a huge time sink off my plate

0
回复

@beyzaokhan 

Thanks — that state-checking logic is such a silent time sink until you actually have to write it, and then it's suddenly half your integration code. Glad requires is taking that off your plate instead of just being another thing you have to reverse-engineer yourself.

0
回复

The requires field for encoding element dependencies is genuinely clever. It's the kind of context that usually takes teams weeks of brittle DOM mapping to figure out, baked right into the response.

0
回复

@uurk6oe 
Thanks, Uğur — that's exactly the gap I was aiming at. Teams end up hand-mapping that dependency logic per site, and it's brittle the moment the DOM shifts, so baking it into the response instead of leaving it as tribal knowledge in someone's scraper script was the whole point. Appreciate you naming why it matters, not just that it exists.

0
回复

the requires field is honestly such a nice touch, most tools i've tried just hand you a flat list and you have to figure out ordering yourself. this feels way more usable out of the box.

0
回复

@fadimea2ds 
Thanks, Fadime — that ordering problem is exactly what I built requires to solve. A flat list of clickable things looks useful until you actually try to drive it and realize you have no idea what needs to happen first, so you end up guessing or brute-forcing the sequence. Glad it's landing as usable out of the box rather than something you have to fight with.

0
回复

Having a built-in way to record and replay a full user flow during integration testing would be huge. Something like a "golden path" capture from one successful Manifest call so my agent team can replay it on staging and catch drift before it bites production.

0
回复

@tansuorbacogbk 

That's a good use case — a different ask from a runtime helper, this is more of a regression net.

Right now every /manifest call is a stateless snapshot — there's no notion of "this is the golden run for checkout" that later calls get compared against. Building that would mean capturing a manifest (or a sequence of them across a flow) as a baseline, then giving you a diff against staging: which actions disappeared, which requires changed, which locators broke. That's genuinely useful for catching silent drift before an agent hits it in prod instead of after.

It's a bigger lift than a simple SDK convenience — closer to a testing/CI feature — but it's a real gap. Noting it as its own thing.

0
回复

Love that the requires field captures element dependencies, that's a real pain point I keep hitting. One thing that would save me a ton of time: a built-in wait helper that knows to poll until those required dependencies are satisfied before resolving the click action. Right now I still end up writing retry logic around it.

0
回复

@kamilkiei 

Thanks Kamil — and yeah, that's a fair ask. Right now requires tells you what's blocking an action, but the client doesn't do anything with that information for you — you're stuck writing the poll loop yourself, which sounds exactly like what's biting you.

A helper that takes an action id, watches its requires list, and resolves once those are satisfied (with sane backoff/timeout defaults) is a pretty natural extension of what's already in the SDK — it's sitting right on top of blocked_actions(), which already exists for exactly this check. Feels like the kind of thing that should just be client.wait_for(action_id) rather than something every user re-implements.

Adding this to the list — appreciate the specific pain point, it's a lot more useful than a vague "nice to have."

0
回复

the requires field is such a smart touch, encoding dependencies directly into the map saves so much back-and-forth between agent and page. clean SDK and MCP server setup too.

0
回复

@ferhattkge 

Thanks, Ferhat — that back-and-forth is exactly what I wanted to cut out. Without the dependency info baked in, the agent's stuck poking at the page to figure out what's actually available, which burns cycles it shouldn't need to. Glad the SDK and MCP setup are landing clean too — appreciate you calling that out specifically.

0
回复

honestly this looks super useful for shipping agents, the requires field alone solves a headache ive dealt with for months. one thing id love to see is some kind of caching or snapshot diffing so you dont have to re-scan a full page every time the agent loops, basically only re-resolve the parts that actually changed.

0
回复

@sleymanbayrwcs 

Thanks, glad the requires field is hitting a real pain point — that was the part I spent the most time getting right.

On caching: there's already a manifest cache under the hood (Postgres, TTL-based) so repeated calls to the same URL within that window don't trigger a full re-scan. But you're pointing at something sharper — diffing within a session so an agent looping on the same page only pays for the parts that actually changed, rather than a flat TTL that's blind to state changes. That's a better model than what's there today, especially for longer agent loops where the page mutates between actions.

Filing this away as a real direction — appreciate you naming the actual mechanism instead of just "make it faster."

0
回复

the requires field is the clever part here, most of these tools stop at "here's a button" and skip the ordering constraints entirely. how do you actually detect those dependencies though - is it driving the page and observing which elements become clickable after which actions, or some static analysis of the DOM/JS? asking because sites that gate steps behind async validation (like a plan check that hits an API before enabling the button) seem like they'd be hard to catch without actually executing the flow.

0
回复

@omri_ben_shoham1 

Right now it's static DOM signal inference, not driving the page — that's the honest limitation. When Playwright pulls the page, I capture things like disabled/aria-disabled attributes, which fields sit in the same form as a disabled element, and which required fields map to which submit button. Those signals get handed to the LLM extraction step along with the accessibility snapshot, and it infers requires from there (e.g. "this submit button is in a form with these three required fields" → requires: [field_a, field_b, field_c]).

So it catches the common cases — disabled-until-valid buttons, forms with required fields — but you've correctly identified the gap: anything gated behind async JS validation (a plan check hitting an API before enabling a button) is invisible to this approach if the element doesn't expose it via disabled/aria-disabled in the initial DOM snapshot. No custom validation logic execution, so it can miss dependencies that only manifest through JS side effects rather than DOM state.

The more accurate version would mean actually driving the flow — filling fields, watching what enables/changes — but that's a much heavier (and slower/costlier) operation than a single-page snapshot, so it's a real tradeoff, not just an oversight. Right now requires is explicitly best-effort: good enough to unblock naive agents on the common patterns, not a guarantee.

0
回复

The requires field is a really nice touch, caught me off guard how often other tools just hand you a list of clickable things without telling you what's gated. Quick to wire up with the Python SDK too.

0
回复
@erife149505 Thanks! That was the exact gap I kept hitting — a “clickable” list is only half the picture if you don’t know what has to happen first to actually use it. Glad the SDK setup was smooth too, that was a priority for launch.
0
回复

Congrats on the launch. The action-manifest idea is useful because browser agents usually fail on hidden state, auth changes, or brittle selectors. How do you encode uncertainty or required preconditions when a page changes after the manifest is generated, and can developers inspect why an action was skipped?

0
回复
@yaroslav_stelmakh Thanks, appreciate it! Good questions, breaking them apart: Preconditions/uncertainty: the requires field is exactly this — each action carries the dependency graph of what needs to happen first (e.g. “select shipping method” requires “address form submitted”). It’s declarative rather than a live probability score, so it tells an agent what has to be true, not how confident we are that it still is. Staleness after generation: manifests are cached with a 6-hour TTL, so if the page changes within that window (auth state, dynamic content, etc.) there’s a real risk of drift — that’s the honest limitation right now. Selectors are resolved against DOM structure/role/name rather than raw CSS paths, which makes them more resilient to minor markup changes, but a full state change (e.g. login → logged-in) will still require regenerating. Inspecting skipped actions: this is the sharpest part of your question, and it’s not something I’ve built yet — there’s no explicit “why was this excluded” trace surfaced to the developer today. That’s a genuinely good feature request (audit log of extraction decisions), and I’m going to add it to the roadmap. Happy to go deeper on any of these if useful.
0
回复
#12
Topolines
Generate Topologic contours
140
一句话介绍:Topolines 通过选取真实地理位置并自动抓取SRTM等高精度海拔数据,生成可直接用于Figma或打印的矢量等高线地图,解决了设计师无需打开重型3D软件即可获取真实地形轮廓的痛点。
Design Tools
地形制图 等高线生成 矢量地图 GIS工具 设计素材 Figma插件 SRTM数据 打印输出 地图可视化 海拔分析
用户评论摘要:用户普遍认可其基于真实海拔数据的实用性,但核心建议集中在功能拓展上:期望添加等高线海拔标注、保存自定义渐变预设、支持CSV/用户高程数据导入。界面流畅与SVG导出干净获一致好评。
AI 锐评

Topolines切中了一个小而精准的缝隙市场——介于专业GIS软件与设计工具之间的“地形矢量素材生产”。其真正价值不在于技术壁垒(SRTM数据本身公开),而在于将复杂的地学数据处理压缩为了一个“画框-生成-导出”的极简三步骤,成功将专业能力降维交付给设计师群体。产品克制地选择了“不画图,只生成真实地形”的立场,避免了沦为花哨的涂鸦玩具,这是其核心差异化优势。

但短板同样明显:目前仍是一个“单次转换”的工具,而非“地形资产管理”的工作流。评论中反复出现的“海拔标注”“预设保存”“导入数据”等诉求,实质上都在要求从一次性工具进化为可复用的设计组件库。若只停留在当前形态,用户粘性有限——设计师完成一次地形下载后,下次使用可能已是数月后。真正的护城河应是沉淀“个性化地形参数预设”与“区域高程数据集缓存”,并接入Figma插件生态形成即开即用闭环。另外,10m GeoTIFF的数据覆盖率不够透明,如果用户频繁遇到不可用区域,付费意愿将迅速衰减。整体而言,这是一款“聪明的工具”,但若要成为“必要的工具”,需要从数据消费层向上游设计与协作层再进一步。

查看原始信息
Topolines
Draw a zone, generate crisp topographic contour lines, and export print-ready SVG & HD PNG. Gradients, elevation taper and Figma-ready output.

Hey Product Hunt! 👋

Topolines turns any real place on Earth into a clean topographic contour map. You pick a location on the map, it pulls the actual elevation data for that spot (SRTM 30m, or 10m GeoTIFF for finer terrain), and generates the contour lines from it. So the ridge you export is a real ridge, not a shape you drew.

I built it to solve a recurring friction in my own design work: getting vector topographic maps without opening heavy 3D software or rebuilding them by hand in a vector tool.

From there you can tweak contour density, line weights, the elevation taper and colors, then export as clean SVG or HD PNG, ready for Figma, web, or print.

Try it on somewhere you know well, that's when it clicks. I'm around all day for questions and feature requests 🙌

6
回复

@alexis_oules Refreshing to see something on here that is not just AI Slop and solves an actual usecase. It's not in my wheelhouse but the demo video looks really cool! :) Good luck!!

0
回复


Seeing a few comments about drawing or sketching zones, so a quick clarification for anyone landing here: there's no freehand drawing in TopoLines. You pick a real location on a map and the contours are generated from actual elevation data (SRTM 30m, or 10m GeoTIFF on paid). The terrain you export is a real place, not a shape you drew. Happy to answer anything else.

1
回复

I am completely obsessed with this app. I've been hiking in the LA area for 3 years now with a really weak angle. I'm able to assess depth over distance swiftly and decide if I can make it down or up super steep mountainsides. Both me and my ankle thank you 🫡

0
回复

The contour generation actually looks smooth and respects the drawn zone boundaries properly, not just a sloppy clip mask. That gradient-to-line opacity taper is the kind of detail that makes it feel made for real cartography work, not just a tech demo.

0
回复

using real SRTM elevation data instead of just letting you doodle contour-looking shapes is what makes this actually useful for real design work, not just a pretty toy. how fine does the 10m GeoTIFF option get in practice - is that resolution available everywhere on Earth, or only for certain regions where that survey data exists?

0
回复

the gradient control on the elevation taper is genuinely thoughtful, gives you that subtle handmade cartography feel instead of looking like a flat computer readout. love how clean the SVG export is too.

0
回复

@coztturk23123 That's exactly the feeling I was after. Old survey maps have a weight that a flat computer readout loses. For me it's the taper plus the elevation gradient working together that gets you there, one thins the lines with height, the other shifts the colour. Thanks for noticing.

0
回复

Adding a quick elevation labeling option directly on the contour lines would save a ton of time, especially for detailed terrain work where you need to identify specific altitudes at a glance without opening a separate legend.

0
回复

@ervahnkw It's on the list. The tricky part is placing labels along a curving path with a clean break behind them, done badly it reads worse than no labels at all. Not next week, but I want it.

0
回复

It would be great if you could add an option to import an existing SVG or image so we can trace over real terrain data instead of having to draw zones from scratch every time.

0
回复

@ferdivppa You draw a zone, but only to select the area. The contours come from real elevation data for that spot (SRTM 30m free, 10m GeoTIFF on paid), so there's nothing to trace over, what you get is already real terrain.

0
回复

Love how clean the export looks and the Figma output is a huge plus for my workflow. One thing that would save me a ton of time is letting me save custom gradient presets tied to elevation ranges, so I can quickly reuse the same color scheme across multiple terrain projects without re-tuning each one.

0
回复

@fahrettinnfsk Saving presets is very doable and honestly should already exist. Quick question so I build the right version: do you reuse the same elevation ranges across projects, or re-anchor them per terrain? A 300m hill and a 3000m range need different stops, and that changes whether presets store absolute values or relative ones.

0
回复

Love how the elevation taper keeps the contours feeling organic instead of mechanical. The Figma-ready SVG output is a really thoughtful detail for designers.

0
回复

@abdurrahmah5cb Thanks! Organic vs mechanical is the exact line I was trying to walk.

0
回复

Honestly the gradient and elevation taper feature is super useful, way better than other tools I've tried. Exported SVG dropped cleanly into Figma without any cleanup.

0
回复

@veyselclyw Thanks Veysel! The clean Figma drop-in was a big goal — the SVG is generated server-side with proper path grouping, so there's nothing to ungroup or flatten on your end. Glad it's holding up in a real workflow.

0
回复

Drew a quick test zone and the contour output came back surprisingly clean, with smooth gradients and crisp lines ready to drop straight into Figma. The SVG export saved me a ton of cleanup time.

0
回复

Would love a way to import a simple list of elevation points or a CSV so I don't have to draw zones manually when I already have survey data. That would speed up my workflow a ton.

0
回复

@cafermqyf Interesting one. Right now everything comes from public elevation datasets rather than user data, so a CSV import would be a different pipeline. Curious what your survey data looks like, is it dense enough to interpolate contours from, or more like spot heights?

0
回复

The contour generation is impressively clean for such a quick draw, and the elevation taper really makes the depth read well. Exporting straight to a print-ready SVG saved me a ton of cleanup.

0
回复

The elevation taper setting is such a thoughtful touch, it keeps the contours from feeling flat at the edges. Clean Figma export is the cherry on top.

0
回复

Love that the elevation taper actually shows up in the exported SVG, makes the contour map feel tactile rather than just flat geometry.

0
回复

@hacerisdf Yep, it's computed server side before export, so the editor preview and the SVG are the same thing. No baked raster.

0
回复

Drawing a freehand zone and watching the contours snap out so cleanly was a nice surprise. The HD PNG export is already perfect for dropping straight into a deck.

0
回复

Honestly kind of surprised how clean the contours came out, even when I scribbled a messy zone. The Figma-ready export saved me a bunch of fiddling.

0
回复

The zone drawing UX is genuinely satisfying, feels closer to sketching than configuring software. Love that the contour output drops straight into Figma without a cleanup pass.

0
回复

The SVG export is genuinely crisp and the elevation taper control feels like something a cartographer would appreciate, not just a design toy.

0
回复

@uurtqst Appreciate it. The taper came from staring at a lot of old topo sheets trying to figure out why they read so much better than equal weight lines.

0
回复

Drew a quick mountain ridge and the contour output looked surprisingly clean for a browser tool. The Figma-ready SVG was the real treat, dropped straight into my layout without any cleanup.

0
回复

@sezerstenknpj Thanks! Just to be precise about what's happening: you draw the zone to pick the area, but the contours inside it come from real elevation data, so that ridge is an actual mountain rather than a shape. Glad the SVG dropped in clean.

0
回复

Drew a quick zone on a hillside photo and the contours came out way cleaner than I expected, especially with the elevation taper. Exporting straight into Figma saved me a solid hour of fiddling.

0
回复

Drew a quick zone around a coastal area and the contour lines came out clean and print-ready without any fiddling. Love that the SVG drops straight into Figma without weird cleanup.

0
回复

Honestly the contour lines come out way cleaner than I expected for something this quick. Drawing a zone and getting print-ready SVG basically instantly feels like cheating, you know

0
回复

Smooth workflow from sketching a zone to getting clean contour lines, and the export quality actually holds up when I dropped the SVG into Figma. The elevation taper feature is a nice touch I didn't know I needed.

0
回复

The zone-based drawing flow feels really intuitive, and the elevation taper on the contour output looks beautifully clean. Love that it spits out print-ready SVG without any fiddling.

0
回复

Would love a way to import a simple point list with elevation values so I don't have to draw zones by hand every time, basically just paste lat/long/alt and get contours out

0
回复

The contour smoothing looks genuinely sharp on the preview, and the gradient-to-taper blend feels really considered. Love that the SVG output drops straight into Figma without weird cleanup work.

0
回复

The contour lines came out surprisingly crisp on the first try, even on a quick freehand zone. Love that the SVG dropped straight into Figma without me having to clean anything up.

0
回复

the fact that the contour taper responds to your drawn gradient rather than just snapping to preset levels is a really thoughtful touch. the svg export stays clean too, no weird node cleanup needed when I dropped it into figma.

0
回复

@ahinyu95 Thanks for the kind words on the export. Small correction, the taper doesn't follow the gradient, it's just an on/off that thins the lines with elevation. The gradient is separate.

0
回复
#13
OpenChatCut
Open-source AI agent video editor with a real timeline
133
一句话介绍:OpenChatCut是一款开源、本地优先的AI视频编辑器,核心在于让AI代理(如Codex、Claude)在真实的多轨道时间线上直接编辑剪辑、转场、字幕和特效,解决现有AI视频工具“黑箱式生成、导出即死、无法微调”的痛点。
Productivity Open Source GitHub Video
开源AI视频编辑器 本地优先 多轨道时间线 代理工作流 MCP协议 AGPL许可 视频剪辑工具 AI创作工具 可编辑导出 透明AI
用户评论摘要:用户普遍赞赏“真实时间线”和“本地优先”,认为这解决了AI工具“导出即封死”的痛点。主要问题包括:4K代理生成支持不足、需简化非技术用户的分享预览流程、以及希望支持实时预览代理编辑过程。有用户询问与ChatCut的关系及音频编辑深度。
AI 锐评

OpenChatCut的“真实时间线”概念,在遍地AI视频工具的今天,的确切中要害。当前多数AI视频工具(包括ChatCut、Runway等)本质是“AI生成器”,用户获得的是一个几乎无法反编译的成品。OpenChatCut试图反向操作:让AI作为助手在标准非编时间线上工作,保留人类编辑者的最终控制权和迭代能力——这才是“AI辅助创作”而非“AI替代创作”的正确姿势。

然而,其价值能否落地存疑。AGPL许可证是一把双刃剑:它向开发者社区释放了信任信号,但对普通用户(尤其是依赖GUI的商业剪辑师)几乎无感。真正的挑战在于:MCP代理(Codex/Claude)如何理解并操作一个复杂的、有嵌套关系的多轨道时间线?如果代理生成的剪辑逻辑混乱、关键帧错位,用户最终需要花更多时间修正AI的“创造”,那本地优先和开源的优势将荡然无存。

更关键的是“用户原始赞同数仅133”,这在Product Hunt上算中等偏下的声量,说明其“技术性叙事”未能打动足够多的泛视频创作者。团队画了一个“AI+专业时间线”的饼,但能否将“可编辑性”从一句口号落地为流畅的用户体验,并解决导入性能(如4K代理)、导出质量和协作分享等硬伤,才是决定其是“颠覆性工具”还是“技术demo”的分水岭。一句话:方向对了,但距离“好用”还有至少两个大版本的代码要写。

查看原始信息
OpenChatCut
OpenChatCut is an open-source ChatCut alternative: local-first AI video editing where Codex, Claude, and MCP agents work on a real multitrack timeline. Free, editable exports, AGPL.
Hey Product Hunt 👋 Launching OpenChatCut — an open-source, local-first AI video editor where agents edit a real multitrack timeline (not a one-shot black box). • Built-in agent + external agents via MCP (Codex / Claude Code) • Real timeline: clips, transitions, captions, effects, exports • Free & open source (AGPL) — independent, not affiliated with ChatCut 🌐 https://openchatcut.com ⬇️ https://openchatcut.com/downloadhttps://github.com/0xsline/OpenC... Early build note: macOS may warn about an unsigned app — expected for now. Would love feedback on install UX and agent → timeline quality. Happy to answer anything 🙏
2
回复

@0xsline This is timely I was just wrestling with AI video tools that paywall you after a couple of clips, so a local-first open-source option is exactly what I needed. Love that it plugs into Claude Code via MCP. Curious how the agent handles multi-clip edits right now stitching a screen recording + b-roll + captions? Excited to try it. Congrats on the launch 🚀

2
回复

Love the open-source approach and the multi-agent timeline idea. One thing that would help adoption a lot is adding a simple share/export-to-web preset that bakes the project into a viewable link, so non-technical folks can preview cuts without installing anything. That alone would make sharing work-in-progress much easier.

1
回复

@yeimoyranshdr Great suggestion, Yeşim. A simple share-to-web workflow would make reviewing work-in-progress much easier, especially for non-technical collaborators. We’re definitely exploring this direction.

0
回复

Local-first editing with a real multitrack timeline is a really thoughtful call here, makes the whole thing feel like an actual editor instead of a toy. Appreciate the AGPL stance too.

1
回复

@ersinnxtr Thank you, Ersin. We wanted it to feel like a real editor from day one—not just a chat interface around a rendered video.

0
回复

finally an editor that doesn't fight me on the timeline, and the local-first angle means i can actually trust it with rough cuts. genuinely impressed it stays free under AGPL.

1
回复

@nuraypzdp Really appreciate that, Nuray. Making the timeline feel predictable and trustworthy is exactly what we wanted. And yes, staying open and free under AGPL is an important part of the project.

0
回复

finally a chat driven editor that respects the actual timeline instead of pretending it does, love that it stays local first

1
回复

@eypmerteelw7z Thanks, Eyüp! That distinction is exactly what we care about: the chat should operate on a real, structured timeline—not pretend to. Local-first is a big part of making that workflow trustworthy.

0
回复

Love that the timeline stays multitrack and editable on export, not flattened into a render. Hooking Codex and Claude into real video editing through MCP feels like the first time agents can touch a timeline without breaking it.

1
回复

@takbug32958 Really appreciate that! Keeping the timeline multitrack and editable through export is one of our core principles. The Codex/Claude + MCP workflow is designed to let agents make structured timeline edits without flattening everything into a render.

0
回复

been wanting a real open source editor like this for a while, the local first angle is huge honestly. one thing that would make me actually switch though is proper proxy generation for 4k footage, my laptop struggles with raw timeline scrubbing and most editors punt on this

1
回复

@anlcengi7zjn Thanks, Anıl — totally agree. Proper 4K proxy generation is high on our list. A local-first editor should make high-resolution footage smooth to work with, not force users to scrub raw files. We’re exploring automatic proxy generation with flexible resolution and codec options.

0
回复

the timeline preview when agents are working on it would be a game changer, like being able to see the rough cuts they're proposing in real time without having to open each export. would make the back and forth so much faster when you're collaborating with the agents

1
回复

@seherb62222 Totally agree — being able to watch rough cuts take shape directly on the timeline would make collaborating with agents much faster. Real-time progress and previewing are definitely areas we want to improve. Thanks for the thoughtful feedback!

0
回复

finally something that doesn't lock my edits behind a paywall or a cloud upload, the local-first setup with claude actually tweaking the timeline felt surprisingly natural and the agpl license is a huge plus.

1
回复

@saliha489828 That’s exactly what we hoped the experience would feel like: your project stays local and under your control, while the cloud helps with the heavy lifting. Really glad the workflow felt natural to you — and thank you for appreciating the AGPL choice!

0
回复
The thing that puts me off most AI video tools is you get one finished clip and can't change anything after. A real timeline you can still edit and undo, makes way more sense to me. Congrats on the launch!
1
回复

@etiennegarcia Exactly — the goal is to keep AI-generated work editable rather than turning it into a dead-end export. You should always be able to refine, rearrange, or undo what the agent did. Thanks so much for the support!

1
回复

just curious whats your relations with chatcut lol

1
回复

@cruise_chen OpenChatCut is an independent, open-source alternative to ChatCut. We’re not affiliated with the commercial ChatCut product—the name describes the category and comparison, not an official relationship.

0
回复

I like that the agent edits a real, reversible timeline instead of handing back a flattened export. Keeping captions, music, transitions, and keyframes visible makes this feel much easier to trust.

0
回复

The 'edits proposed for review before applied' loop on a real timeline is what makes this feel usable rather than a black box. When an external agent works over MCP (say Claude Code), does it get the full timeline state as context each turn or just a diff of the current edit, and where do those proposed-but-unaccepted edits live — in the project file on disk I could version in git, or app-internal state until I accept? I'm trying to gauge whether the whole editing history stays reproducible outside the app.

0
回复

The real-timeline part is what sells me — chat-only editors feel like a black box, but coming from a DAW world, being able to drop into a multitrack timeline and hand-nudge things after the agent's pass is exactly the workflow I trust. Quick one: how deep does the audio side go — multitrack audio with volume keyframes/ducking, or mostly video-first for now? Local-first + AGPL is a great call btw.

0
回复

i cut short product demos for social every week and the timeline is always where ai tools fall apart. they generate clips but you can't nudge anything after. agents on a real multitrack timeline over mcp is the missing piece. can claude code drive it today or do i need your chat ui?

0
回复

Love the local-first approach and the real multitrack timeline finally being usable with AI agents. One thing that would really help me as an editor: a way to version the timeline edits the agents make, so I can diff or roll back specific agent actions when something goes sideways.

0
回复

The "real multitrack timeline" part is what makes this stand out to me — most AI video tools spit out a black-box render you can't touch, so having the agent edit an actual editable timeline (and local-first) is a much saner workflow. Love that it's AGPL too.

Quick question: since it works with Claude/Codex/MCP agents — do I bring my own API key/model, or is there a local-model option for people who want it fully offline?

0
回复

The part I'd poke at is how the agent addresses clips through MCP. Every time we've handed an agent a stateful document as tools, state drift is the thing that bites: it reasons over a snapshot it read a couple calls back, so 'trim the third clip to 5s' hits the wrong clip once the timeline moved underneath it. Are the edits keyed to stable clip IDs it can re-reference, or to positions and timecodes that shift when tracks change?

0
回复

Local-first multitrack editing with agent support sounds genuinely useful. One thing that would make a huge difference for me is a simple undo history that spans across agent-driven edits, so I can roll back a whole AI cleanup pass without losing my manual tweaks from before.

0
回复

Love the local-first approach and editable exports being free. One thing that would help a lot is proxy media generation so 4K timelines stay snappy on lower end machines. Would also be great if the AI agents could remember project wide style choices between sessions so you don't have to re explain your color grade each time.

0
回复

honestly the multitrack timeline with MCP agents sounds super cool, but one thing i'd really want is a way to save and share edit "recipes" or agent prompts so i can reuse the same workflow across projects. basically a little library for your favorite prompt chains that you can drop in whenever.

0
回复

finally a chat based editor that doesn't lock my timeline away, ran a quick cut with claude and the export actually stayed editable which is more than i expected from something free.

0
回复

A collab mode would be huge, like letting two people hop on the same timeline remotely with comments pinned to specific clips. Since the agents already touch the project files, feels like a natural next step for a local-first tool.

0
回复

Would love to see a simple preset system for common edits like "remove filler words" or "add jump cuts" so I don't have to prompt the agents for the same workflow every time. Could be a small dropdown that bundles the right tool calls together.

0
回复

The local-first approach with real multitrack timeline support is such an underrated detail, makes it feel like an actual editor instead of a chat toy.

0
回复

Love the local-first angle and the multitrack timeline with real agent integration. One thing that would make a huge difference for me: a built-in proxy/media cache so 4K footage doesn't bog down the timeline when the agents are running analysis in the background. Would make the whole editing flow way smoother on mid-range hardware.

0
回复

Love the local-first approach and editable exports being free. One thing that would really help adoption is a built-in shareable project link feature where collaborators can open the same timeline in a read-only browser view without installing anything, perfect for getting quick feedback from clients or teammates before they commit to the setup.

0
回复

Love the local-first approach and the multitrack timeline is exactly what I was missing. One thing that would really help: add a visual diff or changelog view when the AI agents make edits, so I can quickly review what Codex or Claude changed before approving. Right now trusting invisible agents on a timeline feels a bit opaque.

0
回复

finally something that doesnt lock your timeline behind a subscription, ran a quick cut with claude and it actually respected my existing tracks instead of nuking them

0
回复

@canertaanux3r That’s exactly the experience we’re aiming for — your timeline should stay yours, and AI should work with your existing tracks instead of flattening or rewriting them. Glad Claude handled the quick cut cleanly. Thanks for giving it a try!

0
回复
#14
DualStream
Simultaneous desktop and mobile streaming, the easy way.
126
一句话介绍:DualStream是一款面向混合内容创作者的直播软件,通过单一引擎实现桌面端与移动端同步推流,并利用中继服务器提供断线保护,解决了多平台、多格式直播时设置复杂、依赖插件、容易崩溃的痛点。
Video Streaming User Experience Streaming Services
直播软件 多平台推流 桌面移动同步 断线保护 中继服务器 原生提醒 音频源控制 VTuber支持 即时回放 创作者工具
用户评论摘要:用户普遍认可双端同步推流和断线保护的实际价值,并提出了具体需求:跨平台统一聊天审核与禁言规则、组建共享版主面板、以及内置带同步标记的多轨录音功能以优化后期剪辑流程。
AI 锐评

DualStream的出现,本质上是对直播行业近十年技术惰性的一次精准打击。当OBS凭借其开源地位成为事实上的行业标准,整个上下游的创新都被困在了“给OBS做插件”的浅层游戏中——开发者借此捞金,用户则被迫忍受日益臃肿、安全漏洞频发的系统。DualStream的价值不是做出了一个“更好的OBS”,而是自建了一套完整的中继架构,将多平台分发从用户端本地的插件拼凑,转移到了云端,并以此为基础构建了原生断线保护。这一架构思维上的降维打击,直接解决了核心场景的刚需痛点,而非在OBS的篱笆外修修补补。

从产品逻辑看,其野心不止于替代推流工具,而是定义下一代创作者工作流。原生集成VTuber控制、多格式即时回放、分轨录音与后期标记,这些功能在OBS中需要极其复杂的插件组合才能实现,而DualStream将其作为系统级的原子能力设计。这种“一体化”思路有效降低了创作者的认知负荷与管理成本,尤其对于单兵作战或小团队,是实实在在的提效。

但风险也同样明显。作为三人团队,能否在功能迭代中持续保持稳定性是最大挑战。评论区中对“断线保护”的具体机制尚有疑问,证明该核心技术还未完全成熟。此外,9.99美元的月费门槛在功能和免费版OBS之间创造了明确的性价比博弈。如果DualStream无法在体验上提供碾压式的边际价值,或者团队因为资源不足而频繁出现BUG,用户的切换到头来只会变成一次昂贵的试错。产品方向正确,但护城河仍需时间与稳定来铸造。

查看原始信息
DualStream
Beautifully powerful streaming. Astonishingly easy. DualStream brings Desktop and Mobile audiences your content at the same time, from a single patent-pending engine. DualStream Relay carries your broadcast to every platform from our servers — and if your connection blinks, disconnection protection keeps your stream alive. Native Alerts, Per-Source audio control, VTuber support, instant replay clips, cinema grade effects, pop-out multi-platform chat and so much more is all built in.

Hey Product Hunt! 👋 MD here, founder of DualStream, with my co-founder John (Tangent on Twitch where he's been Partner for over 13 years). Quick story on how this got built, because the origin actually matters:

I have spent 20+ years building digital products. For a big chunk of that time I've watch streamers as a sort of background environment while I code or design, I love chat and being a lurker. For most of that time, I watched streamers I know describe their setups in horror: plugin duct tape, multi-streaming nightmares, software so complex they'd lose entire weekends to it. The same stories, year after year.

A while back I reached out to Tangent - I'd been a chatter of his since the days of No Man's Sky drops - and asked if he'd be willing to test out some prototypes I was playing with about a better way to multi-stream. He said yes, and then walked me through everything streamers actually deal with. The list was long. We started talking to other streamers. The list got longer. Their wish lists were specific, repeated, and almost never met by the tools on the market.

Those conversations became DualStream.

A note before the pitch: OBS is one of the greatest open-source projects of the last 20 years and the foundation that taught most of the streaming world how this all works, we have nothing but love and respect for OBS and the dedicated team behind it. DualStream isn't here to replace it. We're here to be the next thing - for creators who want a different path. Every creative field got its intuitive layer (Squarespace, Canva, CapCut). Streaming was the last one waiting.


DualStream is one studio that streams to your Desktop and Mobile audiences at the same time - same sources, arranged for each scene, live everywhere in one click. The DualStream Relay - one upload from your PC, live on every platform from our servers. And if your connection drops mid-broadcast, Relay keeps every platform live on your standby screen and picks back up the moment you reconnect. Native Alerts, Per-Source audio control, VTuber support, instant replay clips (horizontal and vertical), cinema grade effects, pop-out multi-platform chat and so much more is all built in.


The app is free to download - build everything, test offline, go live to one channel. Premium is $9.99/month for access to the DualStream Relay, multi-format recording, water-mark free clips and streaming to more then one channel at a time.


🎁 For Product Hunt: use code PH-3M-HI6NIX at checkout and your first 3 months of Premium are free (or just go to dualstream.gg/ph, its applied automatically that way).

We're a three-person team shipping every two weeks, and we genuinely build what users ask for. We'll all be here all day in the comments. Ask us anything, about the product, the build, the streaming category, the design process, the streamer feedback that shaped it, whatever. Genuinely grateful for your time!


// MD & Tangent

Oh! Fun little P.S. ~ We also just released our first Open Source project for DualStream: An MCP server! This server lets AI assistants like Claude operate your DualStream studio: switch scenes, save instant-replay clips, build entire scene layouts from scratch, restyle your alerts and widgets, and react to what happens on stream. Everything runs locally on your machine, with your explicit approval, and every change renders on both your desktop and mobile canvases. Check it out here.

7
回复

Greetings Product Hunt! I am John, co-founder of DualStream! I am a full time live streamer on Twitch with 13 years of Partnered, full time experience on the platform with 230,000 followers under the online handle 'Tangent'. I started streaming as a passion hobby once I discovered such things existed and I was instantly hooked. Over a decade later I am still here and have experienced every up and down a creator can imagine!

In that time the content creation space has evolved and changed (And grown exponentially!) in ways many creators might not expect! One place that has never changed though? The reliance on open source OBS for operating what 99% of the streaming world utilizes.

My contributions to DualStream are in the features I feel that content creators of all sizes not just want, but absolutely need in todays ruthlessly competitive streaming space. MD can handle the technical side of things, but if you are a content creator; a dreamer just starting out with a hobby or an experienced content creator hearing a streamer's perspective is, read on my friends!

OBS is responsible for giving me and many of my peers the careers that we have had and enabled us to turn hobbies into careers. OBS was never intended for live streaming, but with a little creativity and understanding of its (confusing) settings, we have been able to fit it to what we need.

But that is one of the big challenges of OBS! It is robust in the hands of somebody who is technically experienced with the software, but to a new user it is a daunting trial and error experiment. Multi streaming has been the recent hot topic people have been trying to crack in the last few years. If you want to branch your stream from Twitch to Youtube, Tiktok, Kick, etc AND have both a vertical feed for phones and a horizontal for PC users, you have to jump through many hoops to make this possible.

OBS is capable of doing these things, but requires the use and understanding of plugins and mods, all of which are designed by good (usually, there have been security issues) intentioned users. The specific installation instructions and settings required to use many of these plugins results in frustrations and radical changes to the OBS UI. On top of that, it bloats the OBS software and makes it perform in unexpected ways depending on your PC setup. You shouldn't have to have custom docks installed and complicated modifications to your streaming software to branch your stream out in 2026.

When Twitch removed their no compete clause on multi streaming, the wild west once more opened up to creators. The push to crack the easy multi streaming code was the hot topic. For OBS, this meant bigger bloat, more plugins, worse PC performance, many crashes, and a lot of experimentation. For a new streamer you can afford to make mistakes and adapt to the multitude of glitches and crashes if you mess it up, but if you are a full time caster where your mortgage payments are reliant on you keeping your content schedule steady and professional, this can be harrowing!

Remember where I mentioned OBS has 99% of the market for users? Well, all of the innovations for multi streaming (And everything else) began flowing through OBS. The industry decided that the best way to cash in on this, was to bloat OBS even more, and design entire companies and software around utilizing the already strained OBS systems simply because it was so prevalent in the space. Virtually nobody has innovated streaming software from the ground up, too concerned about leaving the security of the confusing, bloated, and increasingly expanding software just because the assumption was that everything must be tied to OBS. The companies that did try to make their own software charge ridiculously high prices for essentially what OBS does with a somewhat more simple approach.

This is where we come in! MD has been a long time viewer of mine in my twitch channel, and witnessed me having some extreme problems with OBS. Crashes, disconnects, encoding problems, multi stream failures. At one point Twitch kept closing my connection to OBS mid cast due to a problem with OBS that would require them to patch it out, leaving me helpless for months.

He approached me and offered to build me a private connection to Twitch, my own personal streaming software because he was very kind and wanted to help me. He built me a very simple, but extremely effective piece of software that proved to be incredibly stable. He asked me if there was any features that I would like to have on it, and this turned into multi hour brainstorming sessions about the industry and tools that I need for my content.

This would eventually become DualStream. I was so floored by MD's ability to implement my requests (And I have a lot of them!) so quickly and so simply that the discussion started to steer in the direction of "If this is helping me so much, it has to be able to help others too."

In no time we had a complex streaming software that was so SIMPLE to use that required ZERO plugins or modifications and sidesteps virtually all of the pain and agony associated with trying to multi stream in multiple formats. I have been using DualStream exclusively for thousands of hours and proudly use it every day for my live streams. I entrust my financial wellbeing as a full time creator in DualStream to put my money where my mouth is. Without wanting to sound cocky, I have had more financial success in the last year of utilizing DualStream on my casts than previous years with OBS, even with the bumps in the road of developing the software and live fire testing it on my live streams.

But it isn't just the raw ability to reliably connect. It is all the fun features for interacting with audiences, manipulating your feed with special effects and on screen novelties that normally require browser sources and many plugins to use. We even have a fully customizable on screen Alerts that offer more customization that Streamlabs and Streamelements without the need of coding knowledge (And it is less taxing than using browser sources.. Though we have browser sources available as well of course!). A lot of streamers don't realize that browser sources are TAXING on your system and the reason every streaming software utilizes them is because the streaming world is chained to OBS and everyone is too scares to take the chance on innovating and those that do see even small creators as cash cows and will charge very high prices knowing that there are no better options.

So that's it! DualStream is a from the ground up approach to live streaming, designed by creators aimed specifically for creators, not merely by people trying to cash in on creators. I have not abandoned streaming to focus on DualStream, I am actively streaming 60+ hours per week and feel that if I am to properly guide DualStream in its future creator features that I will need to be on the front lines of content creation still.

If you have any questions about live streaming, the successes and failures (and risks!), or anything in the creation space as well as DualStream, please feel free to ask! Thank you for your time! - John / Tangent

twitch channel for verification: https://www.twitch.tv/tangent

4
回复

Simultaneous desktop + mobile streaming is a great use case for hybrid content creators. Does it sync audio/video timing automatically, or is manual alignment still needed?

3
回复

@ark_y_k Hey there! Audio is synced by default, and DualStream allows you to set every audio source manually! If you are audio savvy you will find a lot of flexibility in how you can be selective with your sources. If you want a more simple approach, you can do catch all audio with desktop audio as well.

One more great feature is that every audio track can be selected to be excluded from VOD's so that when fans are watching your videos you won't get hit with DMCA's for playing copyright music because it will not save or appear in the recording, enabling you to enjoy music with your live audience and not risk your videos being taken down after your platform of choice saves them for casual viewing!

1
回复

honestly this looks really solid, the simultaneous desktop and mobile push is a great angle. one thing i'd love to see is built in collaborative moderation for chats across all linked platforms, basically a shared mod queue so my team can handle trolls without jumping between five tabs. would make running a bigger stream way less chaotic

2
回复

@semihajmuo I'm actually in the middle of designing collab/shared control for mods, its been a goalpost feature for a while and now with the core systems in place, it's on the nearer term roadmap :) Appreciate the note!

1
回复


How does the disconnection protection actually work? Does it buffer on your servers or keep something local? That's genuinely useful if it can hold a stream alive for even 30 seconds, but I'm curious what happens if someone loses internet completely.

2
回复

@talhakhalidmtk Great question! Right now what happens when internet drops is we play a "Standby" slate - a looping video. You can define what media (image, gif, video) is played or use our default. Your stream actually never goes down, we maintain your connection to the platforms via our relay system and the moment your connection is back, we start broadcasting your stream again. I'm actively working on improving this system to include components like live chat and clips playback automatically, so all the stream would lose while you're disconnected is your actual live A/V, but the rest could remain up while you sort out your internet.

1
回复

finally something that handles desktop and mobile at the same time without me juggling two setups. honestly the relay feature saved me when my wifi dropped mid stream

2
回复

Something different for a change, everyone else seem to be trying to clone each other :0)

Good luck with the launch you guys

1
回复

@albattran This comment made my day, thanks so much!!

0
回复

the fact that everything from chat overlays to VTuber support lives natively in one window instead of stacking extra apps is honestly kind of a game changer, you can tell the team actually streams themselves

1
回复

Pushed to twitch and tiktok at the same time with no extra setup, that alone sold me. The chat pop-out keeping everything in one view while streaming is genuinely nice.

1
回复

The multi-platform chat pop-out is great, but it would be super helpful if there were a way to set per-platform chat moderation rules so I don't have to ban the same spammer three times across Twitch, YouTube, and Kick.

1
回复

@fcalapverd79575 Love that idea! Our entire cross-platform lives on a relay I built, so building in custom rules like that is totally possible and would make a great feature, we'll add it to the list! Appreciate you!

1
回复

One thing that would make DualStream a no-brainer for me is built-in multitrack recording with sync markers. Being able to pull a local high-quality archive after every stream and drop the markers straight into Premiere or DaVinci would save a ton of post-production time.

0
回复

@melekstoluqf6g We have multi track recording built in! Right now the multi-track is for audio stems, but video is on the horizon. Outputs as .mkv files so premiere/davinici will love it :)

0
回复

Finally a tool that doesn't make me choose between desktop and mobile viewers. The relay server thing is basically magic - my stream stayed up when my internet died mid-game, which honestly never happens with other setups.

0
回复

The dual desktop and mobile stream actually looked crisp on both without me tinkering with bitrates, which never happens for me. Also appreciate the relay kicking in when my wifi stuttered mid broadcast.

0
回复

Tried it out for a quick stream and honestly the dual desktop and mobile push just worked, no fiddling. The relay keeping the stream alive when my wifi hiccupped was kind of a lifesaver.

0
回复

The way the per-source audio control is baked right into the same panel as the stream health indicator is a really thoughtful touch. It feels like the team actually streamed before building this.

0
回复

@azizstolufz4r Appreciate that! We have really obsessed over the UX of DualStream, glad to hear you're digging it.

0
回复

Pushed a test stream to desktop and mobile at the same time and it just worked, no fiddling with OBS scenes or restream plugins. The disconnection protection saved me when my wifi hiccuped mid-broadcast, which honestly sold me on it alone.

0
回复

The relay setup genuinely saved me when my internet flickered mid-stream, no awkward goodbye screens. Love that I can finally chat with both desktop and mobile viewers without juggling windows.

0
回复

The cross-platform streaming idea is solid, especially with that disconnection protection built into Relay. One thing that would really level it up is a built-in collaborative editor where co-hosts can queue clips, soundboard cues, or scene changes in real time without needing a third-party tool like Streamdeck. It would make multi-person streams feel way smoother.

0
回复

@gkhanyurdam5oq We have streamdeck integration and it is growing in features every day! We also have an incredible portal system that allows you to connect to another DualStream user and pipe their camera and audio directly into YOUR stream! It will make everything from podcasts to collabs incredibly easy.

A collaborative co host itself is something we have been considering as well, so we are of a similar mind!

0
回复

@gkhanyurdam5oq Great feedback and idea Gökhan! We have a beta system we call Portal built into DualStream that allows users to send their A/V to other DualStream users via a shortcode, sort of like joining someones co-op game. Adding in collab edit tools would be a great addition, we'll be looking into this!

0
回复
#15
Routebase
Catch API drift before your customers do
126
一句话介绍:Routebase是一款API规范漂移检测工具,通过将设计稿、文档、模拟、测试和监控统一为单一事实源,帮助团队在客户发现之前捕获API接口与实现之间的不一致。
API SaaS Developer Tools
API规范管理 漂移检测 开发工具 文档同步 MCP集成 开发者体验 OpenAPI CI/CD 代码Agent 通知告警
用户评论摘要:用户普遍认可其漂移检测实效,多人反馈“一改即查”到数周未发现的过期端点。核心诉求集中于:优先级最高的GitHub Action CI门禁、Slack/Teams即时告警、直观的规格与消费端并排差异对比。另有人提出漏检出站Webhook载荷的问题,团队已将其纳入规划。
AI 锐评

Routebase切中了一个虽不性感但极具破坏力的“慢性病”——API规范与实现的持续熵增。当团队规模扩大、Agent编码成为常态后,这份“细微不匹配”的成本会从少数人的核对痛苦,指数级放大为整个交付管线的连锁错误。其产品逻辑非常清晰:不是再建一个文档工具,而是作为规范生态的“中间校验层”,通过MCP接口让AI Agent直接引用最新规范,巧妙绕过了传统文档更新的“人肉维护”陷阱。

从评论反馈看,用户对“检出问题”给出了高分,但对“如何让检出结果驱动行动”提出了强烈甚至有些苛求的期望。Slack/Teams告警、CI门禁、并排diff——这些都不是锦上添花,而是从“发现问题”到“闭环解决”的刚需。目前产品在“发现”环节能力突出,但在“响应与集成”环节还显单薄。关键在于,若只停留在仪表盘式的被动告警,团队依然要依赖个人责任心来处理漂移,这恰恰是脆弱的。

最值得玩味的是用户对出站Webhook的质疑。这暴露了当前产品专注于OpenAPI的REST范式,而对于现代微服务中大量存在的异步事件、Webhook回调缺乏覆盖。如果Routebase只停留在REST的“单点”漂移检测,而无法横向扩展到AsyncAPI等更广泛的契约类型,其“单一事实源”的野心就会大打折扣——在事件驱动的架构里,没有覆盖到的部分恰恰是最容易产生隐秘故障的地方。

团队在评论区的快速响应和规划展示了极佳的执行力,但真正考验价值的时刻,在于能否在功能堆积之前,将“核心检测引擎”的覆盖边界明确,并围绕“开发工作流自动化”而非“人工巡检”重构体验。解决痛点容易,打造一个让团队“无感”的防熵系统,才是这个产品从“好用”走向“必需”的关键一跃。

查看原始信息
Routebase
Your API lives in your design tool, your docs, your mocks, your tests, your monitoring — and every copy drifts. Routebase makes them one living source: change the spec, publish, and it catches what drifted. (Your AI agents read the same source over MCP.)

12 hours in — quick update: This thread has been better product feedback than most user interviews. The three most-requested things today — a turnkey GitHub Action for CI gating, first-class Slack/Teams drift alerts, and a side-by-side spec↔consumer diff — all went onto the roadmap today. Keep it coming. And my question from this morning still stands: what's the worst "the docs lied to me" moment you've had? 👇

0
回复

the request/response drift coverage in this thread is thorough. one direction I didn't see covered: outbound webhooks. those break consumers just as often as REST responses do, but there's no request to inspect, just whatever payload you decide to send on an event. is webhook payload schema part of the same spec/drift pipeline, or is that a separate problem you're leaving to the consumer to validate on their end?

0
回复

@galdayan Great question, and you've spotted a real gap. Today the answer is honest: outbound webhook/event payloads are not in the spec/drift pipeline yet. Everything — mocks, tests, docs, drift — is built on request/response operations parsed from OpenAPI paths; we don't currently ingest the 3.1 webhooks/callbacks object or AsyncAPI, so event payloads aren't modeled at all. That's not a deliberate "leave it to the consumer" stance — it's just an unmodeled surface. You've convinced me it belongs on the roadmap as its own coverage axis (model the event contract → same doc/mock/test/drift guarantees as REST). Thanks for pushing on it. 🙏

0
回复

Finally tried this on a messy internal API setup and the drift detection caught two stale endpoints we totally missed. The MCP piece feels like the real unlock for our coding agents.

0
回复

@sraaf2l Thank you! Two stale endpoints you'd totally missed is exactly the catch we're after — and hearing MCP feels like the unlock for your coding agents means a lot, that's precisely the bet we made. 🙌

0
回复

honestly the drift detection part is what pulled me in, super useful. one thing though, it would be amazing if you could add some kind of slack or teams alert when drift actually gets caught in prod, so the right person gets pinged right away instead of someone having to go check the dashboard. would save a lot of "wait when did this break" moments.

0
回复

@demettozkavevg Thank you! You're pointing right at something we're actively working on. The plumbing is there — live monitors already raise a dedicated "schema drift" alert into our notification pipeline, and alert policies can fire to a webhook, so a Slack/Teams incoming webhook can already catch it. What we still owe you is the polished part: first-class Slack/Teams alerts (nicely formatted, right person pinged) and drift that pages by default instead of just logging quietly. It's on the roadmap — really useful nudge. 🙏

0
回复

@demettozkavevg Quick update before the day ends: this went from "on the roadmap" to an actual scoped ticket this afternoon — first-class Slack/Teams alerts wired into alert policies, so the right person gets pinged instead of a channel dump. Your comment tipped the scale 🙏

0
回复

finally something that tackles the spec drift problem head on. hooked up our openapi and it immediately flagged two endpoints in our docs that were out of sync, which honestly saved me a headache before our next release.

0
回复

@zehraasnf Thank you! Catching two out-of-sync doc endpoints before a release is exactly the headache we want to spare you — glad it paid off right away. 🙌

0
回复

love that the MCP angle means my agents actually pick up the same spec updates instead of working off last week's version — that detail alone saves a ton of "wait why did this break" moments

0
回复

@leventgzaygdkz Thank you! Killing the "wait, why did this break" moments is exactly why MCP pulls live spec updates instead of a stale copy — glad that detail lands for you. 🙌

0
回复

Finally tried routebase on a messy node project and it caught three routes that drifted between our swagger doc and tests, exactly the kind of thing that always sneaks past review. The MCP angle is genuinely useful too since our agents can pull the same source.

0
回复

@metehaneniknjo Thank you! Routes quietly drifting past review is exactly the leak we built this to close — glad it surfaced three on a real project.

0
回复

The drift-detection angle is genuinely useful, most teams I know are drowning in stale API docs that nobody trusts anymore. Love that it hooks into MCP so agents stay in sync too.

0
回复

@tolgavpkb Thank you! "Docs nobody trusts anymore" is exactly the problem we set out to kill — and yeah, keeping agents in sync via MCP was a big part of the point. Really appreciate it. 🙌

0
回复

Drift detection sounds great, but it would help a lot if Routebase could show a side-by-side diff view between the published spec and each consumer like the docs or mocks so I can spot what changed at a glance instead of just getting a list of drift alerts.

0
回复

@sevilgwkc Great suggestion. Quick tip in the meantime: click "Review changes" on any drifted mock/doc/test and you'll get a field-level breakdown (path + old → new value), so you can already see exactly what shifted rather than just the alert. A true side-by-side spec-vs-consumer panel isn't there yet though — I like it, noting it down. 🙏

0
回复

Finally something that ends the API doc drift nightmare for me. I tweaked a schema on Monday and it actually caught three mocks and a test that had gone stale without me noticing for weeks.

0
回复

@fatmaburkal9ee Thank you! Three mocks and a test going stale without anyone noticing is the exact nightmare we built this to end — glad one schema tweak surfaced all of it for you. 🙌

0
回复

finally a way to stop the endless api drift between figma and our actual endpoints, the mcp piece is slick and the diff view caught two stale params we'd missed for weeks

0
回复

@zekiyehwuy Thank you! Catching two stale params you'd missed for weeks is exactly the kind of thing we want the diff view + MCP to surface. Really glad it's paying off for your team. 🙌

0
回复

A diff view inside the publishing step would be huge so I can see exactly which endpoint fields changed before pushing the update live. Right now it just says "drift detected" but I have to hunt to figure out what actually moved.

0
回复

@didem2rjr you're right that it's harder to see than it should be. the field-level diff actually is in the publish step — which endpoint, which field, old vs new — but today it's tucked behind a "View Full Diff" click, and the "drift detected" notification you're seeing is way too bare (it should link you straight to what moved). surfacing the exact fields right where you first land is exactly where we're taking it. thanks for calling it out 🙏

0
回复

The drift detection actually works. I tweaked one endpoint and it flagged a mock and a test that were both out of sync within seconds, which is exactly the kind of thing I usually find weeks later. Nice that AI agents can read the same source too.

0
回复

@alperenkkkbb2y thank you 🙏 catching it in seconds instead of weeks later is exactly the point

0
回复

honestly the drift catching is what got me, like i updated one endpoint and it flagged two stale mocks i forgot about. saves me from those awkward "works on my end" moments.

0
回复

@diyarkmdwjr thanks so much — killing the "works on my end" moment is exactly what we're here for 🙏

0
回复

Finally something that tackles the copy drift problem head on. Tested it on a small endpoints project and watching it flag a stale response example in the docs was genuinely satisfying.

0
回复

@etinkocako7yqn Thanks so much — really glad it clicked for you! 🙌

0
回复

Finally a tool that gets at the actual pain of API drift. I poked around and the MCP piece is genuinely useful since my agents were always pulling from stale docs.

0
回复

@muhammedjbxo thanks so much — really glad it's clicking, especially the MCP side 🙏

0
回复

Honestly the MCP angle is what got me, I hooked our docs up and it actually flagged two stale endpoints we had been ignoring for months. Kind of wish the publish step was faster but the drift catching alone makes it worth using.

0
回复

@srazgencilswsc thank you, this means a lot 🙏 stoked the drift catching and MCP are landing for you.

0
回复

love that it actually catches drift across mocks, tests, and docs instead of just generating them separately. the MCP piece is smart too, basically lets agents stay in sync without a ton of glue code.

0
回复

@nerminaydao7wn exactly the distinction we care about — anyone can scaffold a mock, a test and a doc from a spec once, but they rot the moment the spec moves. keeping them honest against the source is the whole point. and yeah, MCP means your agents read the same source of truth we do, so there's no glue layer to babysit. thanks for getting what we're going for 🙏

0
回复

Would love to see a visual diff view for breaking changes before publishing the spec update, so we can quickly show backend and frontend teams what exactly shifted instead of just flagging drift.

0
回复

@azizayerrm9y this one already ships — sounds like we just buried it 😅 the publish wizard's first step is exactly that: a color-coded side-by-side diff of your current vs. last-published version, breaking changes flagged red with migration hints and a suggested version bump, all before you confirm. separate from the drift flagging. curious whether it covers what you need once you try it!

0
回复

spent an afternoon hooking it up to our openapi spec and it actually caught two stale routes in our docs i didn't know were drifting. the MCP bit for our agents is genuinely handy too.

0
回复

@ebrarieklitgfp This made my day — catching silently drifting docs against the real spec is the whole reason the OpenAPI sync exists. And the MCP integration being genuinely useful for your agents is exactly the bet we made. Would love to hear how it holds up as you connect more of your spec.

0
回复

Finally something that actually tackles the spec drift problem instead of just complaining about it. The MCP piece is clever too, now my agents read the same source as my docs.

0
回复

@esilape7l Thank you, honestly means a lot 🙏

0
回复

Adding a built-in code snippet generator per language/framework would be huge. Right now catching drift is great but if Routebase could also output a ready to paste fetch client for React, Python, Go etc. right from the spec tab it would save tons of time.

0
回复

@selahattinjtpw Really appreciate this — and good news: a chunk of it already ships today. Every endpoint in Routebase gives you ready-to-paste request snippets in 8 flavors — JavaScript (fetch), Python, Go, cURL, HTTPie, C#, Java, Ruby — with an example body generated straight from the schema. You get them both in the spec editor and rendered inline in the published docs/playground, so it's a copy-click away.

Where you're pointing at something we don't have yet is the step up from a single-request snippet to a real client/SDK — one typed client over the whole spec (centralized base URL + auth, a method per operation), plus a proper React hook flavor rather than a bare fetch call. That's a genuinely useful gap and we're taking it into the backlog. Thanks for the nudge 🙏

0
回复

Really cool concept. One thing that would make this indispensable for my team is if Routebase could auto-generate a changelog or diff summary whenever a spec is updated, so we can quickly see what broke or changed without manually comparing versions.

0
回复

@mzeyyenxgtv Thanks so much — glad it resonates! 🙌

Good news: this already exists. When you cut a new spec version, Routebase auto-generates a changelog from the diff against the previous version — added/removed/modified endpoints and schemas as clean Markdown. And every change is tagged breaking / non-breaking / deprecated with a migration hint, so "what broke" is answered for you. You can also diff any two versions on demand.

One choice: we generate it on publish rather than on every keystroke, to keep it meaningful. Would a live diff while editing a draft be useful for your team too? Happy to give you a walkthrough!

0
回复

How do you handle versioning when an API spec changes? Like if v1 and v2 need to coexist, does Routebase track both or do you have to manage that separately in your tool?

0
回复

@talhakhalidmtk Great question — coexisting versions is exactly what we built for.

Routebase tracks both. Each version gets its own isolated snapshot of endpoints, schemas and folders, so editing v2 never touches v1. On top of that:

  • Aliases — pin v1/stable to a semver like 1.4.2 so consumers don't chase exact numbers.

  • Consumer strategy — configure per spec how versions are exposed (URL path /v1/users, header, query param, or content negotiation); the mock server routes to the matching version and the docs portal generates samples to match.

  • Deprecation & sunset — mark v1 deprecated with a sunset date, and it emits the right headers + reminders so you can retire it gracefully.

Plus a diff between any two versions for an auto-generated changelog. All in one place — no juggling v1/v2 in a separate tool.

1
回复

Congrats on the launch. Catching API drift early is valuable only if teams can triage it quickly. Do you show the exact contract change, affected route, sample failing request, and confidence level so developers can separate real breakage from noise?

0
回复

@yaroslav_stelmakh Exactly the right lens — triage speed is the whole game. Here's the honest breakdown of what a drift item gives you today:

  • Affected route: yes — every drift finding is tied to the specific endpoint the monitor covers, so the route is explicit, not inferred.

  • Exact change: yes, structured per field — the JSON path of the field, the kind of change (missing field / type mismatch / unexpected extra field / format mismatch), and the expected type. So you see which field drifted and how, not just "something changed."

  • Sample: we capture the failing check's actual response — a body sample, status code, response headers, and timing — so you can eyeball the real payload that broke the contract. What we don't do yet is replay the outbound request for you; the request is defined by the monitor/environment config rather than snapshotted per check.

  • Confidence: we don't emit a numeric confidence score today. What we do is classify each deviation by severity — a missing required field or type mismatch is an error (real breakage), an unexpected extra field is a warning, a format nuance is info. That severity is what separates "page someone" from "noise," and it's what drives the alert threshold.

Fair to say the raw signal is structured and route-precise; the piece we're actively investing in is a tighter triage surface around it (a dedicated drift view rather than a dashboard card). Appreciate the sharp question.

0
回复

@denny_riedl Congrats on the launch! The always-on monitoring detail is interesting! I'm curious how it handles auth gated endpoints for that continuous check.. Does it need live credentials handed over to hit protected routes, or is monitoring scoped to public endpoints only?

0
回复

@clement_avq Great question — and no, it's not limited to public endpoints.

Monitors reuse the same environment auth you've already configured for testing that API, so there's no separate handing-over of credentials. Basic, Bearer, API key, OAuth2 and JWT are all supported — for OAuth2 we refresh the token automatically at check time, so a monitor keeps hitting a protected route without you babysitting it.

On the credential side: secrets are stored field-encrypted (not plaintext), and you can also just reference secret environment variables ({{token}}) instead of pasting a raw value. So you get continuous checks on protected routes without exposing anything.

1
回复

Does the drift detection also run in the direction nobody announces? Publishing a spec change and firing contract tests covers the drift you know about, but the deploy that reshapes a response without anyone touching the spec is the copy I've watched drift first. Whether the monitoring side watches the live API continuously or needs a publish to wake it up would be my deciding question as a buyer.

0
回复

@vollos Yes — that split is exactly the point, because publish-triggered contract tests only ever catch the drift you started yourself.

The unannounced direction is handled by Monitoring, not the spec pipeline. You bind a monitor to an endpoint and set schema validation to Warn or Strict. It then hits the live URL on an interval (as tight as every 30s), validates the actual response body against that endpoint's response schema, and raises a Schema Drift alert when the shape diverges — no publish, nobody touching the spec. The deploy that quietly reshapes a 200 is precisely what it's there to catch.

Two honest boundaries so you can decide cleanly:

  • It's opt-in per endpoint — a monitor with nothing bound only watches uptime/status/latency; the schema check turns on once you bind the endpoint and set the mode.

  • It diffs the response body against the JSON Schema on a successful check. Header contracts, request bodies, and undocumented new endpoints aren't part of that automatic diff today.

So the monitoring side watches continuously once armed — it doesn't wait for a publish to wake up.

0
回复

Congrats on the launch! I've run into API drift multiple times, so the centralized API spec makes sense. I'm curious about the last mile: how does Routebase keep the spec aligned with the actual frontend and backend implementations? Can CI block incompatible changes, or does Routebase only detect drift after deployment?

0
回复

@mateuszkonik 

Thanks — drift is the exact itch we built this for. Routebase is spec-first: the OpenAPI spec is the contract, and mocks, docs, tests and monitors are all derived from it. That gives you two directions of alignment.

Against your spec history: on publish we diff against the last version and classify breaking changes — endpoint removed, field made required, response type changed. Any derived artifact that's now stale gets flagged with a one-click "sync from spec."

Against your running API: contract test suites hit your real endpoints and monitors validate live responses against the schema — that's how you catch the implementation drifting from the contract.

On CI — great point, and honestly the missing piece today. The breaking-change diff and contract tests are already available headlessly (MCP + REST), so you can gate a build on them, but there's no drop-in GitHub Action yet. We're putting a turnkey CI check on the roadmap right now and shipping it as soon as possible.

…that's how you catch the implementation drifting from the contract. We don't reverse-engineer your API from source, though — the spec stays the source of truth.

0
回复

Hey Product Hunt 👋 I'm Denny, founder of Routebase.

I'm a CTO, and I built this because of what I kept watching teams go through —
not one dramatic outage, but the same slow friction on every project. The backend changes an
endpoint, and the frontend finds out from a failed integration instead of a conversation. The
docs get updated "afterwards", which means never. The frontend sits blocked, waiting for an
endpoint that's "almost done". And the API counts as "tested" because there are unit tests
around the handler — but nobody has actually checked that what runs in production still matches
the contract everyone coded against.

None of that is one team's fault. It's structural: design, docs, mocks, tests, monitoring — each
tool quietly keeps its own copy of the API, and the copies drift the moment you ship a change.
The pain happens between the tools, so no single point tool can fix it.

Routebase is the one living source they all derive from. You design the API once; your docs, doc
portals, mock server, contract tests, and monitoring come from that same spec. The frontend codes
against the mock from day one instead of waiting for the backend. And when you change the spec
and publish, Routebase catches what drifted. My favorite moment in the demo: I add one required
field to a schema, publish, and a contract test runs against the live API and goes red — a
breaking change caught before a single consumer saw it, and before the frontend lost a day to it.

(And the part I quietly love: your AI agents read that same source over MCP — so the agent answers
from the real API, instead of becoming one more drifting copy of it.)

Routebase v1 is live today: API design, auto-generated docs, contract testing, mock servers,
monitoring, and governance — one source of truth for the whole lifecycle.

I'd genuinely love your feedback — especially if you've lived the backend↔frontend standoff.
What's the worst "the docs lied to me" moment you've had? I'll be here all day. 🙏

0
回复
#16
tterm
A terminal, a real browser, and Claude Code under one roof
125
一句话介绍:tterm是为macOS开发者打造的AI编码驾驶舱,集成真实浏览器和Claude Code,通过逐块审查diff和自重构功能,解决多项目管理时AI代码审查效率低和工具碎片化的痛点。
Productivity Artificial Intelligence Vibe coding
macOS工具 AI编码助手 终端替代 代码审查 diff逐块审查 Claude集成 自重构应用 多项目管理 无账号/无遥测 原生Chromium浏览器
用户评论摘要:用户称赞无账号设计、自重构和逐块diff审查体验。核心建议包括:增加会话状态/项目布局保存、浏览器权限可控(隔离Agent驱动的登录会话)、热重构失败时自动回滚、为提交添加草稿笔记面板。多数功能反馈积极,但苹果架构局限受关注。
AI 锐评

tterm真正的价值不在于“又一个AI编码助手”,而在于它对开发者工作流中一个被忽视的断裂点做了手术:评审AI生成的代码。当前主流方案要么是全盘接受VSCode插件的diff视图,要么是繁琐的PR页面审查,而tterm用“逐块确认+回滚”的交互设计,将AI从“生成者”降格为“协作者”,把人的决策权重重新拉回核心。自重构功能看似炫技,实则是工具链元层次的闭环——开发者用AI构造AI工具本身,这种“吃狗粮”的模式一旦成熟,将极大降低垂直工具迭代的摩擦。

但风险同样明显:嵌入式Chromium直接读取Chrome配置文件的做法,在安全性上存在黑盒隐患。即便开发者承诺只复制不写入,同一进程内Agent驱动已登录浏览器的设计,本质上是将用户数字身份的管理权移交给了Claude的上下文窗口。除非严格实现Origin白名单或独立的Agent沙箱环境,否则这将成为企业级用户最大的采用障碍。另外,“无编辑器”的坚持虽然激进,却可能扼杀审查后的快速微调场景——用户不得不在tterm和编辑器间频繁切换。如果后续版本不能提供“一键打开文件到编辑器”的桥接,这种极简主义反而会演化成新的摩擦点。整体而言,tterm是一个有远见的alpha级产品,但距离成为日常依赖,还需在安全沙箱和状态持久化上补课。

查看原始信息
tterm
tterm is a macOS cockpit for working with Claude Code. Every project is a row of explorer, Claude, and file viewer. Stack as many as you're juggling. A real Chromium browser lives inside, and there's no editor on purpose: you review diffs hunk by hunk and commit. It even rebuilds itself. Ask the Claude pane for a feature and the running app hot-reloads with it. No account, no telemetry. Free for hobbyists, on Apple silicon.

Appreciate you owning the tradeoff so plainly, most people would hand-wave it. Since you're clearly keeping the signed-in mode (I would too, that's where the useful work is), the thing that softened it for us was splitting read from write: the agent browses freely, but anything that changes state, a push, a delete, a purchase, waits on a one-tap confirm. Keeps most of the flow autonomous while the irreversible part stays yours. Would that fit tterm's diff-review philosophy, given you already gate commits hunk by hunk?

1
回复

@dipankar_sarkar definitely would, human in the loop is much needed. Claude will often (but not always) end tasks prematurely in order to hand off the sensitive task to the user (don't ask me how I know this ;)).

0
回复

The layout is genuinely nice, basically treating each project like a little row you can stack while juggling multiple. Kind of wild that you can ask Claude in the pane to add something and the app just rebuilds itself around you.

1
回复

@melahatdnda0pp haha thank you!

0
回复

the self-rebuilding feature is wild, asked the claude pane for a dark mode toggle and watched it hot reload right in front of me. also appreciate the no-account stance, just downloaded and started using it.

1
回复

@kadir1626034 aww man you made my day :)

0
回复

No account and no telemetry is the right call, but the embedded Chromium carrying my Chrome bookmarks and cookies is the part I would want scoped before running this daily. Does the browser pane read my live Chrome profile directory, or a copy made at first launch? And since the Claude pane sits in the same app, can the agent drive that browser against sessions I am already authenticated into, or is agent-initiated navigation confined to a separate profile or an origin allowlist?

1
回复

@hi_i_am_mimo tterm reads copies of Chrome's Bookmarks, History, and Cookies files from disk once, decrypts cookies with Chrome's own Safe Storage key from the macOS Keychain (which fires a visible one-time consent prompt), and writes everything only into tterm's own browser profile under ~/Library/Application Support/tterm. Chrome's profile directory is never written and never watched :)

Secondly, yes, the claude pane in the base version can drive the embedded browser, including pages you're signed into. Admittedly, this is still something I'm working on making better. It only runs when you explicitly send it a message, and everything it reads from pages is treated as untrusted data in its instructions. There's no origin allowlist or separate agent profile yet. That's a fair ask and I'd like to add it.

Thanks for the great questions!

0
回复

The self-rebuilding feature is genuinely wild. One thing I'd love: a way to bookmark or save the exact stack layout across projects so I can pick up right where I left off when I jump back in.

1
回复

@asiyekmeogipbc multiple people have had this suggestion, will definitely be adding it in! Thanks for the feedback :)

0
回复

Love the no-editor, diff hunk by hunk flow, that's a clever take on reviewing AI generated code. One thing I'd find useful is a way to save session state per project so when you reopen a workspace the explorer tree, open files, and even scroll position come back exactly where you left them. Right now it sounds like every launch starts from scratch and that would save a lot of clicking.

1
回复

@aysunasralq0vp this (minus the scroll position) is the default at present! will add the missing piece :)

0
回复

stackable project rows with the explorer, Claude, and diff viewer side by side genuinely sped up how I review changes across multiple repos at once. Asking Claude for a tweak and watching the app hot-reload with it still feels a bit like magic.

1
回复

@ayekapszm5vk my feelings exactly!

0
回复

finally a mac app that gets out of the way for claude code, and the hunk-by-hunk diff review is genuinely faster than clicking through vscode. weirdly charmed by the self-rebuild feature too

1
回复

@doukanywb5 thank you for the kind words :)

0
回复

Finally a Claude wrapper that doesn't try to be an editor. Stacking multiple projects side by side with the diff review built in is exactly what I needed when juggling a few repos at once.

1
回复

@eyllggerciyevg haha or a cozy youtube video while claude goes at it!

0
回复

as a real Chromium browser is bundled in, would love to see a way to save reusable diff-view layouts or presets per project type, makes flipping between code reviews and Claude chats way faster

1
回复

@hiranurerkasap love the suggestion, gonna def prompt that into my tterm (and you can too!) :)

0
回复

The hot-reload from asking Claude for a feature is genuinely clever, and committing without ever opening an editor is a real UX call, not just an omission.

1
回复

@irmakmemik Thank you!

0
回复

I would test if it wasn't mac only, looks cool

1
回复

@inferhaven haha I might just port this for you and you only!

1
回复

The diff-by-hunk flow is genuinely clever, I'd love a small "drafted explanation" panel where I can leave notes for my future self before committing. Some commits have weird "why I did this" context that gets lost once the message is final, and writing it inline in the commit box always feels rushed. A scratchpad that auto-attaches to the commit hash would fit right in with the project-row layout you already have.

1
回复

@nazkudak Love this, and this is something you can totally bake into your iteration of tterm! My fav part of all the suggestions i'm getting is that every single person can just bake in their suggestions into their tterm. Regardless, will add this to my list to add to the base version!

0
回复

the self-rebuilding part is the thing I'd want to stress test before trusting it daily. if you ask the Claude pane for a feature and the hot-reload lands on something broken, does the app roll back to the last working build automatically, or are you stuck debugging your own cockpit while it's down? that failure mode seems like the actual risk of a tool that rewrites itself while you're using it.

1
回复

@omri_ben_shoham1 in my experience, I haven't had any issues that I haven't been able to resolve with a subsequent prompt. the base version does not have checkpointing, but I will be sure to add that - thanks for pointing out!

0
回复

Splitting panes per project with the diff review flow feels natural on a mac, and the self-rebuild from a Claude prompt is a wild idea that somehow just works.

1
回复

@zzet1032068 Thank you!

0
回复

Putting terminal, browser, and Claude Code in one place makes sense if the handoffs stay visible. For shipping work, the valuable layer is checkpoints: what changed, what was verified, what failed, and where the human needs to make the next call.

1
回复

@krekeltronics agreed! presently, the diff review is meant to be that checkpoint - nothing agent writes lands unless you've reviewed it hunk by hunk. is there anything you'd like to see added in addition to this?

0
回复

So grateful to everyone supporting this launch :). Totally unexpected for me that people would see this (and potentially like this, it's something I made for myself!

1
回复
Hey Product Hunt! I built tterm because I was living in three windows: a terminal running Claude Code, a browser to check what it built, and an editor I barely typed in anymore. tterm folds them into one. Every project is a row of explorer, Claude, and file viewer. A real Chromium browser lives inside, with your Chrome bookmarks and cookies along for the ride. There's no editor on purpose: you review diffs hunk by hunk and commit. The fun part is that tterm builds itself. Most features shipped by asking the Claude pane running inside tterm, and the app hot-reloads. Even the landing page is in on it: the ocean is the word OCEAN riding a live wave simulation, and it reacts when you scroll. Free, no account, no telemetry. I'd love feedback!
0
回复

That breakdown is exactly what I wanted — copy-once into tterm's own profile with a Keychain consent prompt is the clean answer. The part I'd still gate is the agent driving authenticated sessions: treating page text as untrusted data protects the model's reasoning, but a prompt-injected page could still steer it to navigate and submit somewhere I'm already logged in. Even before a full origin allowlist, would a read-only mode on authenticated origins (view but not click/submit) be feasible as a stopgap?

0
回复

the browser-session sharing thread here is a great read. curious about the other direction: if I've got two or three project rows stacked for different repos, is each row's Claude pane fully isolated from the others, or could a prompt in one row's context ever end up seeing something from a different project's files or conversation history?

0
回复

@tanay good to know a follow-up prompt usually gets you unstuck. checkpointing would be a nice safety net for the cases where the rebuild goes sideways in a way that's hard to describe back to the model though - worth adding even as a manual 'save state' button before you ask it to touch the cockpit itself.

0
回复

Ha, I've seen that exact behavior, the model bailing right before the irreversible step and going 'you take it from here.' It's a nice instinct but I wouldn't build a guarantee on it, since the same model that stops 90% of the time will confidently push the button the other 10. What worked for us was moving the decision out of the model entirely: the confirm gate fires on the action type, any write, delete, or spend, so it's the same whether Claude flagged it or not. Then the model bailing early is a bonus, not the safety net.

0
回复

The bundled Chromium carrying your Chrome cookies is the detail I keep circling. I run agents against logged-in Chromium too, and each one gets its own throwaway profile on purpose: if the Claude pane can drive a browser already logged into your email or GitHub, a wrong click there executes as you, against your real accounts. That is a much bigger blast radius than a bad hot-reload. Is the in-app browser sharing your everyday Chrome session, or is it a separate profile the agent is scoped to?

0
回复

@dipankar_sarkar They are on a separate profile, however, working with the user's real accounts is optional (and in my case, intentional). I have tterm do a bunch of real work which requires me to be signed in. Unfortunately that comes with the caveat of the larger blast radius :(

0
回复

Combining terminal, browser, and Claude Code in one window sounds like it could cut a lot of context-switching. How does it handle screen real estate when all three are active — tabs, splits, or something else?

0
回复

@ark_y_k On the base version, I use splits - but you're free to configure it however you like!

0
回复
#17
PieceKeeper
Track your music repertoire and practice
122
一句话介绍:PieceKeeper 是一款帮助音乐人管理曲目库、制定个性化练习计划、追踪练习数据并提供乐理训练的工具,有效解决音乐人因练习不规律导致的曲目生疏和遗忘问题。
Music Classical Music
音乐练习 曲目管理 练习计划 学习追踪 乐理训练 钢琴 吉他 音乐教育 效率工具 ProductHunt
用户评论摘要:用户普遍认可曲目管理和练习计划功能,认为解决了曲目遗忘和练习不规律问题。高频建议包括:内置节拍器、音/视频录制与反馈、教师同步与协作、AI节奏分析、建立歌单功能。乐理训练模块被评价为基础。
AI 锐评

PieceKeeper 精准切中了业余音乐爱好者一个极为普遍但被长期忽视的痛点——“学了就忘,忘了就废”。其核心价值并非提供教学,而是作为“音乐练习的CRM系统”,将模糊的“练习”行为数据化、流程化、可追溯。从产品设计和用户反馈看,这套“曲目管理+定时练习+数据洞察”的闭环逻辑,确实能有效解决练习缺乏系统性的问题,这是其真正值得关注的点。

然而,产品的短板同样明显。首先,其乐理训练模块被普遍认为“基础”,在当前AI乐理应用层出不穷的市场中,这一功能显得缺乏竞争力,更像是一个为了提升产品完整性而添加的附属品。其次,所有用户提出的核心建议——如节拍器、录音反馈、AI分析、教师协作——目前均在“待办清单”上。这些功能并非锦上添花,而是专业练习工具的基础设施。若迟迟不落地,产品将停留在“记录本”层面,难以形成真正的练习护城河。创始人孤身一人,功能优先级和开发速度将直接决定产品命运。

PieceKeeper 的护城河在于其“曲目+日程+数据”的整合,而非某个单一功能的强大。未来若能优先补足节拍器和录音功能,并与教师/乐队协作功能打通,它有望成为音乐练习领域的“Notion”或“Strava”——一个围绕曲目展开的、拥有社区和协作属性的练习平台。否则,它极有可能被更具AI属性的、或内置了强大创作工具的应用所边缘化。总的来说,方向正确,但紧迫感不足。

查看原始信息
PieceKeeper
Never lose track of a piece again. PieceKeeper helps musicians manage their repertoire, run focused practice sessions based on a customized practice schedule, get insights and trends about their practice, and learn music theory through interactive drills and exercises aimed at improving both reading and listening skills.

Hey Product Hunt 👋 I'm Joey, solo founder of PieceKeeper.

The idea for this grew out of a personal problem. As a casual pianist, I'd learned tons of pieces over the years, but forgotten just as many.

Once you learn a piece, it needs regular practice or it goes rusty and eventually slips away entirely. I'd lost track of so much I wished I'd kept due to irregular practice and trying to store what I knew all in my head.

So I built a tool to track the pieces I know. Not just their names but their details too - reference links, sheet music PDFs, learning status, difficulty, and more. I also decided to go beyond just piano, so that anyone who plays an instrument can use PieceKeeper too.

Then my idea expanded: once your repertoire is stored, PieceKeeper schedules practice sessions around your preferences, rotating through everything so nothing falls through the cracks.


Then the idea expanded again. Music theory matters for all kinds of reasons, and as a casual pianist, mine was pretty weak. So I added interactive exercises and drills to build reading and listening skills across all the main categories with the option to fold theory practice into your regular sessions, so everything lives in one place.

So, putting it all together, what does PieceKeeper do? -

Repertoire Organization: save every piece with its sheet music, reference videos, difficulty, instrument, and status
Timed Practice Sessions: practice from a rotating checklist built from your repertoire, tailored to your schedule and practice habits.
Practice Insights: view your practice history, lifetime stats, per-piece stats, and session notes in a digital journal
Music Theory Trainer: choose from 46 interactive drills (along with flashcards) across 8 categories that train your music reading and listening skills.

I'm super excited to finally share this with everyone and would love all your feedback! You can try it for free and start practicing in 5 minutes here: https://getpiecekeeper.com/

Thanks to everyone for taking the time to read and check this out, and happy playing!
- Joey

2
回复

@joey_wang378 very cool idea! I used to play some instruments (at a very amateurish level) and would love to come back to it one day. But I always thought: how is it possible to keep all the learned pieces in my head? Good to know there is a ready solution :)

1
回复

I love this idea! I just started getting back to playing guitar and spent 2 hours yesterday sorting my tabs on UG, which honestly has a really old school interface. Do you plan adding social features along the lines of sharing content, etc?

1
回复

@foobar_beer thanks! I feel the pain of organization as a pianist with my references spread out too, so I did focus a lot on trying make sure the UI for this was clean but comprehensive.

Social features are a big item on my to-do list, I have 3 main aspects in mind - 1) teacher-student syncing and feedback, 2) connecting with friends, seeing what they're learning/practicing, and 3) connecting with bandmates/groups for regular practice or prepping for gigs. So lots of possibilities! Thanks for your feedback!

0
回复

The repertoire-decay thing is painfully real — you grind on a piece for weeks and it just quietly rusts because nothing nudges you back to it. Really like that the scheduler rotates through everything instead of relying on your memory: does the rotation weight by difficulty, or by how long since you last played a piece? Solo founder here too, respect for shipping something this thought-through lol.

1
回复

@lennoxbeflying the rotation logic is based on the order you set for your repertoire, and it cycles through it. But the order can be changed at any time, and you can practice any piece you want during sessions regardless of which are 'assigned' for that day. Adding a rotation based on difficulty is an interesting idea and could be worth implementing depending on how people feel! Thanks for your input!

0
回复

Adding a built-in metronome with variable time signatures and accent patterns would make practice sessions way smoother, no need to bounce between apps when you're running through a tricky section at 120 BPM.

1
回复

A practice mode that records a short audio clip of you playing through a piece and then gives feedback on timing or note accuracy would be really helpful. Hearing your own run-throughs back is honestly how I've improved the most, so tying that into the practice schedule would be a nice touch.

1
回复

@miray141473 for sure I agree, audio recordings are a priority on my to-do list for next big features. Having it give feedback on your recordings is a very interesting idea too, probably would use some AI integration or custom model for the actual evaluation. That part would not be trivial to implement but a feature like that would be really cool! Thanks for your feedback!

0
回复

Tried it out for a few minutes and the practice insights are actually useful, you can see where your time is going instead of guessing. The theory drills feel a bit basic but the repertoire management is solid.

1
回复

@glerobcu thanks for checking it out! The repertoire aspects are the main focus of the app and I agree that the theory drills are rudimentary for now, there’s a lot of room for them to grow and evolve over time and into more specialized training. Thanks again for your insight!

0
回复

As a guitarist juggling too many tabs at once, the practice scheduler sounds like a lifesaver. One idea though: add a metronome sync option where the app can tap into your recorded play and flag tempo drift in real time, so you can actually hear where you rush or drag instead of just seeing the data later.

1
回复

@aysimayasanfkz that’s a great idea! It would provide a real basis into how well your rhythm matches the piece so you don’t have to rely on just listening to your recordings . Thanks for the feedback!

0
回复

Finally something for musicians that isn't bloated with features I don't need. The practice schedule setup was surprisingly quick to personalize, and honestly the trend insights on my sessions are kind of eye-opening after just a few days.

1
回复

As a piano player who juggles a lot of rep, the one thing I'd love is a way to share practice schedules or progress with my teacher directly through the app, maybe even let them leave timestamped notes on specific sections of a piece.

1
回复

@abdulkadirgyt6 yes that is great idea! Having teacher - student syncing and them being able to give specific feedback on your practice would be awesome. This is definitely high on my to-do list. Thanks!

0
回复

The practice schedule feature actually made me sit down and play instead of just scrolling through sheet music. Nice seeing the trends chart fill up after a few days too.

1
回复

The practice insights are surprisingly motivating, seeing my consistency charted week over week actually made me want to keep going. The theory drills feel less tedious than other apps I've tried, more like quick puzzles than homework.

1
回复

Would love to see a setlist builder for gig nights, basically something where you can drag and drop pieces, auto-calculate total runtime with breaks, and share a clean printable version with the band.

1
回复

@seraptc6r that’s a great idea! The app so far is mainly oriented around regular practice, but integrating and scheduling upcoming gig prep is definitely doable. Thanks for the feedback!

0
回复

Love the practice insights angle, that part's really missing from most apps I have tried. One thing that would make this way more useful for me is a backing track or metronome built right into the practice session view, so I am not juggling tabs between PieceKeeper and a separate app while working through a piece.

1
回复

hey @farukbykkaioq6 a metronome is definitely on my to-do list! The pain of having to switch around different sources while practicing is real! Thanks

0
回复

A metronome built right into the practice session view would be super helpful, maybe with the option to gradually speed up difficult sections. Also being able to share a piece note or recording with a teacher or bandmate directly through the app would make collaboration way easier.

1
回复

@mcahitdemi69gf definitely both great ideas, and both are on my todo list of future features to add! Thanks for your feedback

0
回复

Finally something for tracking repertoire that actually makes sense, love how the practice schedule ties into insights.

1
回复

As a drummer with a huge library of charts, the practice tracker looks solid, but I'd love a "loop and slow down" tool built right in for the tricky passages I keep stumbling on. Something like slowing playback to 50% speed while looping a measure would save me from bouncing between apps constantly.

1
回复

@zgrpayc06lc Thanks for the insight! I agree that audio integration, recording, and playback options as accompaniments during practice are probably the best next features to add to the site. And being able to customize them for specific instruments like the drums. Something I’m definitely keeping in mind, thanks for the feedback!

0
回复

Love how you tied the practice schedule directly to repertoire tracking instead of just dumping it all into one long list. Makes it way easier to actually stick with a piece until it’s performance-ready.

1
回复

@gamzekalkmazo thank you and you’re correct, I definitely felt like there was no reason to keep the two parts “living” in separate places/apps. I had kept written lists of pieces I knew before, but never recorded the details of how I’d been actually playing them.

0
回复

the practice session timer with the customizable schedule is such a thoughtful touch, makes it feel like a real studio companion rather than just another tracker.

1
回复

@ensar89210 yeah I definitely wanted it to encapsulate the idea of being a practice companion and not just a set of static to-do lists. Thanks for checking it out!

0
回复

the practice insights sound genuinely useful, but honestly what would seal the deal for me is a metronome built right into each session. like if i open a piece to practice it should just be there ready to go, with the option to set the tempo to whatever i'm working on that day.

1
回复

hey @pakize177947 a metronome is a great idea! I’ve been thinking of adding a collection of relevant tools that could easily integrate with practice sessions. Thanks for the suggestion!

0
回复

the practice insights feature is honestly such a smart move, like finally a way to actually see where your time is going instead of just guessing

1
回复

@neclavuys thanks! Whenever people would ask me how often I practice I’d usually throw out a ballpark number, but I really had no idea how accurate it was. But having statistics around individual pieces and lifetime aggregates gives me a true picture into my own consistency, or lack thereof!

0
回复

the rusty-piece problem is so real, I have at least three pieces I could play cold two years ago and now can't get through the first line. question on the rotation logic: does it weight which piece comes up next based on how long it's been since you last touched it or how you self-rated it (rusty vs solid), or is it more of a straight round robin through everything in your repertoire regardless of how recently you played it?

1
回复

@galdayan right now it’s structured with a round robin logic, but practice session checklists are not absolute. You can ignore a piece on the list and mark ones that aren’t on the list as practiced too. So it’s designed to be flexible and the practice schedule itself can be changed anytime. However adding a “practice oldest first” rotation option may be a good idea too! Thanks

0
回复

honestly the practice insights sound really useful, one thing i'd love to see is a way to record short audio clips of myself playing during a session so i can hear my progress over time. kind of hard to tell if i'm actually improving just from stats alone

0
回复

@hayriyezeytun I definitely agree! Audio integration, recording, and playback is a high priority on my todo list of future features to add! Thank you for the feedback

0
回复
#18
Tidy
Your Mac tidies itself: screenshots, installers, Downloads
122
一句话介绍:Tidy 是一款驻留在 Mac 菜单栏的轻量工具,通过本地 AI 自动为新截图重命名、卸载 DMG 并清空安装器、按规则分类下载文件,彻底终结桌面和下载文件夹的日常混乱。
Mac Productivity Artificial Intelligence
Mac 工具 桌面清理 文件自动整理 下载分类 截图重命名 本地 AI 菜单栏工具 效率工具 隐私优先
用户评论摘要:用户高度认可“仅处理新文件”和“永不删除”的安全设计,并普遍赞扬本机截图 OCR 重命名的实用性和 AI 规则的自然语言设定。主要建议集中在:添加预览/预演模式(看到文件即将移动到哪)、增强撤销功能以支持整批操作回滚、以及提供更详尽的可审查操作日志。部分用户担忧 iCloud 同步中占位符文件的处理,开发者已给出分层次防护解释,并承诺在后续版本测试。
AI 锐评

Tidy 的价值不在于“自动化”,而在于“自动化”的信任门槛被彻底击穿。它的聪明之处,是默认把自己限制在“只能触碰新文件、永不删除”这个极小作用域里,然后用本地 OC R和 Apple Intelligence 优雅地解决三个高频但足够恼人的微场景——截图命名、DMG 清理、下载分类。这个产品定位精准得像手术刀,它承认自己不是全能的文件总管,而是那个只扫门前雪的邻居。

但这份克制既是优点也是瓶颈。当前功能虽干净,但扩展性有限:用户反复要求的“预演模式”和“批量撤销”都指向同一个事实——在现有机制下,仍存在用户不敢完全放心的模糊地带。尤其是 AI 规则被描述为“严格可检查的规则+AI翻译”时,其本质依然是基于扩展名的分类器,距离真正理解文件形态的“智能”还有距离;开发者对此的诚实(“AI 不是法官”)值得尊敬,但也意味着短期内的价值天花板明显。

定价策略是另一亮点。$9 的买断制在订阅泛滥的当下形成强烈反差,它巧妙地将“无服务器成本”转化为用户买单的坦率理由。这对于老练的 Mac 用户而言,既是价格信号,也是信任信号。

值得警惕的是,Tidy 目前的护城河很浅——类似的 Dock 清理器、截图重命名脚本不胜枚举。它的真正壁垒不在功能,而在“第一次使用就对你上瘾”的产品体验。OCR 重命名里那种“看似轻微但直接提升幸福感”的考量,才是稀缺品味。如果后续迭代能守住“不膨胀、不入侵、不智能失控”的原则,同时提供更强的可视化控制,Tidy 完全有机会从“改掉坏习惯的临时工具”变成“Mac 工作效率的默认守门人”。否则,它很容易沦为又一个过目即忘的菜单栏图标。

查看原始信息
Tidy
Tidy lives in your menu bar and handles three chores you keep doing by hand. It archives new Desktop screenshots, renaming them from their content with on-device OCR — "Screenshot 14.22.10.png" becomes "invoice-march.png". It ejects thedmg you mounted and trashes the installer. It sorts new Downloads by type, with rules you write in plain English. Apple Intelligence builds them on-device, so nothing leaves your Mac. Only touches new files. Never deletes. $9 once, no subscription.
Hi Product Hunt 👋 I built Tidy because three things kept eating my Desktop: screenshots, .dmg installers, and a Downloads folder I'd given up on. Funny thing happened while testing the checkout. I bought my own app, downloaded the .dmg, dragged it to Applications, opened it — then went looking for the installer to clean up. It was already in the Trash. Tidy had ejected the disk and trashed its own installer. Nobody told it to. That's just what it does. Two rules I set for myself: · Only touches NEW files — never your backlog. · Never deletes. Archives or trashes, always reversible. The plain-English rules run on Apple Intelligence, on-device — nothing leaves your Mac. That's also why it's $9 once, not a subscription: my marginal cost per user is zero. Happy to answer anything.
1
回复

I literally set aside time each week to do this myself so this is so cool!! Does it let you set custom rules beyond strict logic or is there a way to use the AI to set rules with more context that I can describe to it? Also would love to see a log of everything that's been changed every now and then so that I can clear my trash with confidence. Great stuff!

0
回复

@neuroartist Thank you — you just described the target user perfectly (past me, every Sunday).

On the AI: you can describe rules with as much context as you like — "invoices and receipts as PDFs go to Finance" — and the on-device model turns that into a rule. The honest nuance: what it produces is still a strict, inspectable rule (extensions → folder), by design. The AI is the translator, not the judge — no fuzzy per-file decisions moving your files on vibes. Content-aware sorting is a tempting future direction, but only if it keeps the "you can see exactly why" property.

And you're the second person today asking for a full reviewable history — it's on the 1.1 list, and your exact use case ("what did Tidy put in the Trash before I empty it") is what I'll design it around. Today the panel shows recent activity + one-click undo; 1.1 turns that into a proper log.

0
回复

the "only touches new files, never deletes" design is the right call for something running unsupervised. one edge case I'm curious about: I have Desktop & Downloads Folder sync turned on through iCloud, so files sometimes sit as not-yet-downloaded placeholders for a few seconds. does Tidy wait for a file to actually finish downloading locally before it acts, or is there a chance it grabs a placeholder mid-sync

0
回复

@galdayan Great edge case — the most technical question of the launch, and it deserves a straight answer.

Three layers protect you here:

1. Old-style ".icloud" placeholder stubs are dot-hidden files with an .icloud extension — Tidy skips them twice over (hidden files are excluded from every scan, and the extension matches no rule).

2. Modern "dataless" files (real name, content on demand): reading one — like the OCR does — makes macOS materialize it first, and moving one out of a synced folder does the same. If it can't be materialized (offline), the move simply errors: FileManager moves are atomic, so a file is never left half-moved. Tidy logs it and catches it on a later pass.

3. Tidy also refuses to touch 0-byte files or anything modified in the last 45 seconds — the general guard against in-flight transfers.

Honest footnote: I haven't run a dedicated iCloud sync test matrix yet — it's now on the QA list, and I'd genuinely love a report if you ever see it misbehave. The design intent is exactly what you'd hope: when in doubt, do nothing and try later.

0
回复

The screenshot renaming with on-device OCR is genuinely useful, finally my Desktop isn't a graveyard of "Screenshot 14.22.10.png" files. Love that it only handles new files and never nukes anything, plus nine bucks once feels right for something this focused.

0
回复

@ezgiirru "A graveyard of Screenshot 14.22.10.png files" — that's the most accurate description of my old Desktop I've ever read.

And thank you for saying the price feels right. $9-once only works because the app is this narrow: no servers, no per-user costs, nothing to subscribe you to. A focused tool gets to cost what a lunch costs — once.

0
回复

the on-device OCR renaming screenshots based on their actual content is such a clever touch. feels like the kind of small detail that makes you trust the rest of the app to do the right thing.

0
回复

@berkevgo Thank you! You can judge a restaurant by its bathroom — small, unseen care predicts the rest. That was the bet with Tidy: an app that touches your files has to earn trust in the smallest details first, because that's where people decide whether to believe the big promises. Really glad the OCR naming carried that weight for you.

0
回复

The OCR rename for screenshots is genuinely useful, turned a messy pile into something I could actually search. The plain English rules for Downloads feel like the right idea too, though I wish I could preview where a file would land before it moves.

0
回复

@salimwpnh "Something I could actually search" — that's the quiet superpower of the OCR names, glad you found it. Spotlight suddenly works on your screenshots.

And you're the third person today asking to see where a file will land before it moves — that officially makes dry-run/preview the most-requested feature of this launch, and it's locked in for 1.1. Current thinking: a pending list showing "file → destination" so tuning a rule takes seconds instead of trial and error. This thread is doing my product design for me and I'm not complaining.

0
回复

finally something that doesnt try to be a whole file manager. the ocr renaming actually got my messy desktop down to zero icons in a couple days.

0
回复

@abdulsametykzn Zero icons — that's the dream state, thanks for sharing it! And the "doesn't try to be a whole file manager" part is very deliberate: Tidy does three chores and refuses to grow a dashboard. Scope is a feature. The moment a utility needs its own onboarding, it's become the mess it was supposed to clean.

0
回复

Love the plain English rules and the on-device approach. One idea: let me exclude folders from the auto-sort, since I keep a Downloads subfolder for client assets that I want untouched. A simple right-click "ignore this folder" option would save me from writing a rule for every exception.

0
回复

@resuljfoq Good news — your client-assets folder is already untouchable, by design. Tidy only ever looks at loose files at the top level of Downloads: it never recurses into subfolders, and it never moves folders themselves. So everything inside your subfolder is invisible to the sorter — no rule or exception needed. (Someone earlier in the thread had the exact same worry about a WIP folder — clearly a thing many of us do!)

An explicit "ignore this folder" control is already on the 1.1 list though, and I like the right-click idea for it. If Tidy ever gains optional deeper sorting, that toggle ships first.

0
回复

Screenshot auto-renaming on Mac is one of those tiny things I never bothered to fix and now wish I had sooner. The "never deletes, only touches new files" promise is what sold me.

0
回复

@glenakbeleokoa Thank you! Screenshots are the perfect example of an annoyance too small to fix but too frequent to ignore — 30 seconds of squinting at thumbnails, 40 times a week. It never feels worth automating until someone automates it for you.

And you're in good company: "never deletes" is officially the most-quoted line of this launch. I'm taking the hint for the homepage :)

0
回复

the fact that it only touches new files and never deletes anything is such a thoughtful call, basically removes all the anxiety of letting a tool mess with your stuff.

0
回复

@birsensfl7 Thank you! Honest origin story: I've never trusted "cleanup" tools myself — so I built the one I'd trust. The rule I gave myself was simple: the worst possible bug should be "a file moved somewhere visible", never "a file gone". If that's the ceiling of what can go wrong, you can actually relax.

Anxiety was the real competitor here, more than any other app.

0
回复

Love the plain English rules for Downloads, that's exactly the kind of thing I'd actually use. One thing I'd love is the ability to undo a sort run if a rule misfires, since sorting is the kind of action that feels safer when you can roll it back.

0
回复

@egemenucaf Undo is already in there — every move lands in the Activity log and one click reverses it, and the undo stack even survives restarts. But your ask has a fair nuance: today it's per-file, so rolling back a whole sorting pass takes a few clicks. A true "undo this run" — one click to revert everything a pass just did — is a clean upgrade, and it joins the dry-run mode others requested today on the 1.1 list.

Between a preview before and a run-level undo after, a misfiring rule becomes a non-event. That's exactly where this is heading.

0
回复

honestly the plain english rules thing is what got me, i typed "put pdfs from clients into Work/Clients" and it just worked. apple intelligence on-device feels like the right call too

0
回复

@ceydausoa This comment genuinely made my day — you're the first person to confirm the AI rules working out in the wild, on a Mac I've never seen, with a phrasing I never tested. That "it just worked" is Apple's on-device model doing the inference, plus a strict validator double-checking its output before the rule goes live.

And yes, on-device was non-negotiable: a tool that reads your file names has no business sending them anywhere.

If a phrasing ever misfires, tell me — I tune the prompt with real examples.

0
回复

Love the "never deletes" promise and the plain-English rules idea. One thought: a quick "undo" or staging area where I can preview what Tidy is about to do before it actually renames or moves things would make me way more comfortable letting it run unsupervised on my Downloads folder.

0
回复

@azadtekeli Good news on half of that: Undo already exists! Every action lands in the Activity log, and there's a one-click Undo right in the panel — it even survives app restarts. And since Tidy never deletes, Undo always has something to restore from.

The preview/staging half: you're the second person in this thread to ask for exactly that, so a dry-run mode officially went on the 1.1 list an hour ago. Clearly the right instinct — if the tool shows its plan, you don't have to trust it blindly. That's the direction Tidy is heading.

0
回复

Would love a "dry run" mode before it actually moves files - maybe a weekly notification showing what it would have done so I can catch any rule tweaks before things get shuffled around.

0
回复

@zge6ydq That's a genuinely good idea — added to the 1.1 list (this thread is writing my roadmap for me, and I love it).

Today the model is the reverse: act → show → undo. Every move lands in the Activity log with filename and destination, and one click reverses it. Since nothing is ever deleted, a "wrong" move is a 10-second fix rather than a lost file.

But I get the appeal of seeing the plan before anything moves at all — especially in week one while you're tuning rules. A dry-run toggle fits Tidy's philosophy perfectly: trust through transparency. Thanks for this one.

0
回复

the screenshots renamed via on-device OCR is such a thoughtful touch. honestly didn't know i needed it until i saw it. love that it only touches new files too, feels safe to leave running.

0
回复

@berilyldrtfy2q Thank you! That feature came straight from my own pain: squinting at a wall of "Screenshot 14.22.10.png" thumbnails trying to find one bank receipt. Naming them by what they actually say felt obvious in hindsight — the best tweaks usually do.

And "feels safe to leave running" honestly made my day — that was the entire design brief in five words.

0
回复

the "never deletes" line buried at the end is such a thoughtful trust signal, made me actually relax while reading. and the plain english rules for downloads feel like the right level of magic, on device and inspectable instead of some opaque sorter.

0
回复

@ferdibatszcb Thank you — and you've spotted something real: you're the second person today to say the "never deletes" part is what made them relax. It might deserve a promotion from footnote to headline.

And "the right level of magic" is exactly what I was going for: the AI writes the rule, but the rule is a thing you can see — extensions, folder, minimum age, all visible and editable before it ever runs. Magic that shows its work, not a black box that moves your files on vibes.

0
回复

Love that it only touches new files and never deletes, took that worry off my mind right away. The plain English rules for Downloads feel like the kind of thing I actually want to set up once and forget about.

0
回复

@bekirtrkneux6c Thank you! "Set up once and forget" is literally the design goal — there's a section on the site with almost exactly that title. My success metric is honestly a bit backwards: the less you open Tidy, the better it's working. The panel is there for the day you're curious about what it did — the other 364 days, it should just quietly earn its $9.

0
回复

Hi! Does it leave a log of some sort? Meaning, in the case of "hey, where is my file?" can the user know quickly what happened to it?

0
回复

@luis_parker Yes! The menu bar panel has a live Activity log — every action is recorded with the filename and timestamp: "invoice.pdf → Documents", "Screenshot archived and renamed: …", "Installer to Trash: …". So "where did it go?" is answered at a glance for recent actions, and there's a one-click Undo right next to it.

And as a deeper safety net: since Tidy never deletes, a file is always in one of three predictable places — its rule's folder inside Downloads, the screenshots archive (one click to open from the panel), or the Trash. Never gone.

A full searchable history is a fair ask though — adding it to the 1.1 list. Thanks!

0
回复

Auto-organizing Downloads and screenshots is one of those "small thing, huge relief" features. Does it let you set custom rules per file type, or is the sorting logic fixed?

0
回复

@ark_y_k Fully custom! Sorting is rule-based: each rule is a set of extensions → a destination folder, and you can edit everything — add your own rules, change the folders, set a minimum age (e.g. "only move installers older than 1 day"), group into monthly subfolders, or toggle any rule off.

It ships with sensible defaults (Documents, Images, Archives, Installers…) so it works with zero setup, but nothing is fixed.

And the part I'm most proud of: you can just type a rule in plain English — "PDF invoices to Finance" — and Apple Intelligence builds it on-device, inferring the extensions, folder and icon. No syntax to learn, and nothing leaves your Mac.

0
回复

Love how it only touches new files and never deletes, that alone makes me trust it more than most utilities. One idea: let me exclude specific folders from the Downloads sorter, since I keep a WIP folder inside Downloads where client assets live that I don't want shuffled into Images or Documents at midnight.

0
回复

@aleynal0ij Thank you — that trust is exactly what I optimized for.

And good news about your WIP folder: it's already safe, by design. The sorter only looks at loose files sitting at the top level of Downloads — it never looks inside folders, and it never moves folders (the only exception is .app bundles, which macOS treats as single files). So your client assets are invisible to Tidy, at midnight and always.

That said, an explicit exclude list is a fair idea for power users — adding it to the 1.1 list. Thanks for the thoughtful suggestion!

0
回复
Leftover .dmg installers piling up in Downloads is such a specific little annoyance. Good that it just cleans up after itself. Congrats on the launch!
0
回复

@etiennegarcia Thanks Etienne! That was the original itch — the day I started building Tidy I counted 14 forgotten .dmg files in my Downloads. And the fun part: it doesn't discriminate. Seconds after I installed it for the first time, it ejected and trashed its own installer. It practices what it preaches :)

1
回复
@sebadiaz Haha! The self-trashing installer is a great detail and proving it on first contact already earns some trust. Congrats again!
0
回复
#19
Universal Dictation on Stream
Private Voice Ring for dictation across iOS & Mac
119
一句话介绍:Universal Dictation on Stream 是一款跨iOS和Mac的私密语音环,通过按压通话实现无需切换应用的即时语音转文字,解决用户在多设备、多应用间碎片化记录和思考断点问题。
Wearables Artificial Intelligence Audio
语音转文字 跨设备协同 隐私优先 快捷记录 思考外化 生产力工具 无AI人格 离线模式提议 免提模式提议 标签收藏提议
用户评论摘要:用户普遍认可按压说话的流畅体验,喜欢其无AI人格的纯粹记录功能。主要建议包括:增加锁定录音的免提模式、支持离线或弱网使用、提供语音笔记收藏或分类功能、以及添加应用内静音选项以屏蔽干扰。
AI 锐评

Universal Dictation on Stream 的巧妙之处在于它精准地切中了一个被忽视的痛点:思维的“瞬间泄漏”。当我们在多个App之间切换时,那些稍纵即逝的灵感或未成型的思考往往在点击和跳转中被冲散。Stream放弃了传统语音助手试图与人对话的“人格化”负担,转而成为一根纯粹的“思维导管”——通过统一的按压交互,让声音无缝穿越iOS和Mac的生态壁垒,直接进入任何目标应用的文本字段。

其更深层的价值在于重新定义了“记录”这一行为。它不追求AI的过度解读或建议,而是对用户思维的绝对忠实复制,这本质上是一种对注意力的极大尊重。用户评论中提及的“自然”、“不打扰”、“像脑部的延伸”恰恰印证了其设计哲学的成功。然而,风险同样存在:当前的“按压-说话-释放”逻辑虽优雅,但对于马拉松式的头脑风暴或双手被占用的场景(如驾驶、烹饪)则存在天然短板,评论中对手势锁定和离线模式的呼声正是对此的隐性佐证。如果Stream不能快速进化出智能的背景噪音屏蔽、长录音无缝断点续传以及更灵活的手势体系(如双击锁定、摇动结束),它很可能止步于“极简主义者”的利器,而无法成为真正的通用生产力平台。在AI竞相“做加法”的当下,Stream的减法策略十分犀利,但市场的挑剔在于:用户既爱你的极简,又会在急需时痛恨你的“不够”。

查看原始信息
Universal Dictation on Stream
Introducing Universal Dictation on Stream. Push-to-talk across iOS and Mac: Instantly. No app switching, no reconnection. We built Notes so nothing gets lost, and Chat so you can think out loud. With Dictation, your voice goes anywhere. Notes, Chat, Dictation—in one voice ring
Last November we announced Stream, the private voice ring for everything on your mind. Stream lets you capture thoughts into notes and talk through ideas aloud, wherever you are. Over the past year we've been developing the next phase of Stream; Universal Dictation. Just push-to-text in any app across iOS and Mac. Your voice seamlessly moves between applications, without having to swipe between app screens or reconnect devices. It's the quietest, most portable, and most ergonomic way to dictate anywhere. Stream is Notes, Chat, and Dictation in one private voice ring. No subscription required, available at sandbar.com
2
回复

Pressing and just talking feels surprisingly natural, and I like that it nudges me with questions instead of doing all the thinking itself.

0
回复

The press and hold flow works great, but a small "lock recording" toggle would be amazing for longer brainstorming sessions when my hands get tired or I want to gesture while talking. Maybe a double tap to switch into hands free mode that auto ends on silence.

0
回复

the press-and-release thing feels pretty natural, almost like texting but with my voice. kind of nice that it just listens without trying to be its own personality, just acts like an extension of my brain.

0
回复

The press-and-release flow feels really considered, like they obsessed over that tiny moment between speaking and the app catching your thought. love that it's positioned as an extension of you rather than another chatbot with a personality.

0
回复

honestly this looks pretty cool, the press and release flow sounds smooth. one thing i'd love is a quick way to mark certain voice notes as favorites or pin them, so i don't have to scroll back through everything to find the important ones.

0
回复

The push-to-talk idea with earbuds is genuinely nice for quick dictation while walking around. Curious how it handles longer thinking prompts without getting cut off mid sentence.

0
回复

Have been wanting something like this for ages, the press-and-hold capture sounds delightful. One thing that would make it really stick for me is a solid offline mode so my voice notes and dictation work on flights and in spotty wifi, then sync up smoothly once I'm back online.

0
回复

The press-and-release interaction feels really considered, especially how it works the same whether you're dictating into a chat app or capturing a quick note. Nice that it skips the personality theatrics and just acts like an extension of your own thinking.

0
回复

Pressed, spoke, released, and it actually captured my rambling note cleanly without me having to repeat myself. Love that it stays out of the way instead of trying to be its own personality.

0
回复

the press and release flow feels really well thought out, basically just tap and talk without any extra steps getting in the way.

0
回复

The "press, speak, release" flow sounds really nice for capturing thoughts on the fly. One thing I'd love to see is a simple way to chain voice notes into longer threads or projects, so I can keep related ideas together instead of having a flat list of snippets. Maybe a quick "add to..." command after recording.

0
回复

The press-and-hold capture looks really clean. One thing that would help me is per voice mute options for background apps during dictation, since Telegram notifications or podcast audio keep bleeding into my transcripts. A quick toggle in settings to silence everything except Stream during capture would be huge.

0
回复

The press-and-hold interaction is a really nice touch, makes it feel less like opening another app and more like a natural extension of how you already think out loud.

0
回复

Pressing a button and just talking into my notes feels surprisingly natural. Love that it asks thoughtful questions back instead of trying to act like its own person.

0
回复

The push-to-talk flow here feels really considered, like you can just press, speak, release and it disappears into whatever you were doing without breaking your focus.

0
回复

Caught myself reaching for it after just a quick try, the press-and-release flow feels almost invisible. Kinda nice that it asks me questions back instead of just dumping answers.

0
回复

the "self extension" framing really lands for me. it sidesteps the whole AI ego trap and just feels like a quieter, more useful tool.

0
回复

Love how the press-and-release flow feels, super low friction for quick capture. One thing that would make it a daily driver for me is a way to bookmark or pin a recurring voice command, like a morning check-in or end-of-day brain dump, so I don't have to re-set the context each time.

0
回复

the 'self extension, not its own identity' choice is the rare one. most voice ai forces a persona. does dropping it change how people talk to it?

0
回复

Would love to see a widget for quick voice capture without even opening the app, basically like a one-tap floating mic button I can drop over other apps.

0
回复
#20
Diffsmith
Comment on your AI agent's code & collaborate on changes
117
一句话介绍:Diffsmith是一个面向AI编程助手的本地代码审查工作室,让开发者直接在未提交的代码差异行上批注评论,并通过MCP协议将反馈无缝回传给AI代理,实现高效的“人审AI代码”闭环。
Developer Tools Artificial Intelligence Vibe coding
AI代码审查 代码差异评论 MCP协议 本地Git工具 AI编程协作 Claude Code兼容 Cursor集成 离线编辑器 按行锚定批注 Code Review工具
用户评论摘要:用户普遍认可其“按行批注+MCP回传”的实用性,认为解决了与AI代理的沟通痛点。主要建议包括:增加并排原文件视图、评论间快速跳转快捷键、按代理/文件/逻辑块分组筛选、支持评论对遗漏代码的批注、以及评论状态在切换分支后的持久化处理。
AI 锐评

Diffsmith精准地切入了“AI生成代码后,人类如何有效反馈”这一日益尖锐的痛点。它没有去重复造一个全能IDE,而是用“本地离线、按行锚定、MCP回传”这三个极简支点,撬动了人机协作的最后一块拼图——单维度的对话反馈。

从产品逻辑看,它做得聪明且克制。100%离线、一次性买断,直击开发者对隐私和数据归属的敏感点。MCP路径则是真正的灵魂,它让“审查-反馈-修改”从单向的复制粘贴变成了双向异步的数据流,定义了AI编程协作的新范式:Agent不再是黑箱生产者,而是可对话的协作者。

然而,批评点同样精准。评论中“无法评论遗漏代码”一针见血——真正的代码审查不仅仅是看AI写了什么,更是看它没写什么。这是Diffsmith当前最大的功能短板,也是从“差异查看器”跃升为“真正审查工作室”的必经之路。此外,缺乏文件上下文并排视图、评论分组、键盘导航等细节,暴露出产品目前仍处于“能用但不够顺手”的阶段。如果它只是作为阅读AI代码差异的副屏幕,价值将很快被IDE原生插件吞噬。真正的护城河在于,能否将“按行评论”发展为“按逻辑块、按遗漏、按规则”的智能审查网络,并深化MCP协议,让AI在收到反馈后能更智能地定位和修补,而非仅仅展示回复。

一句话总结:它切中了真需求,方向极其正确,但距离“卓越的审查体验”还有至少20%的功能打磨和50%的能力升维。若不快速迭代,很容易沦为“漂亮的IDE插件替代品”。

查看原始信息
Diffsmith
Diffsmith is a code review studio for local git changes made by Claude Code, Cursor, Codex, Copilot, or any other AI coding agent. Open a repo, read the uncommitted diff, and click any line to leave a comment anchored to that exact spot. Then hand your notes straight back to the agent. Supports MCP for viewing you're agent's replies to your comments directly at the line of code it relates to. Diffsmith is 100% offline and a one time purchase unlocks lifetime access.

Would love a side-by-side mode that shows the original file alongside the diff, since right now I'm flipping back to my editor a lot to remember the surrounding context before leaving a comment on a single line.

1
回复

the archiving-on-rewrite answer above covers the case where the agent edits a commented line. curious about the other trigger though: if the agent just runs `git commit` mid-session, does Diffsmith diff against the new HEAD and archive everything the same way, or does it lose track since it was anchored to the previously-uncommitted state?

0
回复

@galdayan after a commit, all comments are archived as currently Diffsmith only shows uncommitted changes.

0
回复

finally a way to actually talk back to my AI agent line by line instead of copy pasting diff hunks into chat. The MCP reply viewer at the exact line of code is genuinely useful.

0
回复

The MCP path is what makes this a loop rather than a one-way review, since the agent's replies land back at the line they belong to. Where does that comment state actually live: inside the repo as something I would gitignore, or in app-local storage outside the working tree? I ask because I want to know what happens to anchored comments when I switch branches or stash mid-review, and whether the MCP server is something the agent polls for new notes or you push to it when I finish a pass.

0
回复

One thing I'd love to see is a keyboard shortcut to jump between comments during review, something like cmd+] and cmd+[ to cycle through them. Right now clicking around to find each note breaks my flow when I'm going through a big diff.

0
回复

Finally a clean way to review what my AI agent just churned out without scrolling through terminal output. Clicking a line to leave a note and bouncing it back through MCP feels exactly like the workflow I didn't know I needed.

0
回复

Would love to see a way to filter or group comments by the agent that generated the diff. When I'm reviewing changes from both Claude Code and Cursor in the same session, it gets hard to track which feedback needs to go back to which tool. A simple dropdown to view only comments for a specific agent's changes would save me a lot of mental overhead.

0
回复

A diff view that groups changes by file with a quick-jump sidebar would save a lot of scrolling on larger branches, especially when I'm bouncing between several modified files at once.

0
回复

Would love to see a way to group comments into review rounds so I can do a first pass, send notes to the agent, then do a follow-up review on just the changed lines from the next commit without losing my original feedback.

0
回复

A diff view that groups related changed lines into collapsible logical blocks (function, class, block) would be huge when reviewing large AI-generated changes, instead of scrolling through every single line individually. Would love to see something like that next.

0
回复

honestly this looks really useful for reviewing ai-generated code, the per-line commenting with mcp replies is a clever touch. one thing that would help me a lot is a way to filter the diff to only show files changed by the agent versus my own edits, so i can focus the review where it matters most. sort of a split view toggle.

0
回复

The line-anchor model nails what the agent wrote, but reviewing Claude Code and Codex diffs my highest-leverage notes are almost always about what it didn't write: 'you skipped the error path here', 'no test for the empty input'. There's no changed line to click on for those. How do you attach a comment to an omission, or to a file the agent should have touched but left alone? That's where a lot of my review time actually goes.

0
回复

Love how cleanly it anchors comments right to the diff line, makes giving feedback to the agent way less painful. The offline approach is a nice touch too.

0
回复

One thing I'd love to see is a way to group related comments into threads so I can have a back and forth on the same block of code without the chat getting messy.

0
回复

This is awesome, much needed visibility upgrade for reviewing what your coding assistant is ACTUALLY doing, love it.

0
回复

honestly the line-anchored comments and pushing feedback back to the agent sounds super useful, but it would be even better if i could batch related comments into a single thread so i don't have to send five separate messages to the agent when i'm reviewing one function.

0
回复

finally something that doesn't nuke my local diffs when I'm reviewing what cursor spit out, the line-anchored comments are kind of genius honestly

0
回复

honestly this looks really useful for my workflow since i use cursor and claude code daily. one thing i'd love is a way to batch similar comments together, like if i want to flag the same issue across multiple files, it gets tedious clicking line by line for repetitive notes. a "apply comment template to selection" or something similar would save a lot of time when reviewing bigger changes

0
回复

the line-anchored comments thing is genuinely clever, like you can basically hold a real conversation with the agent right on the code instead of copy-pasting snippets back and forth. also appreciate that it's fully offline, no telemetry nonsense, just open the repo and go

0
回复

AI code review is most useful when it is tied to release risk, not just style feedback. I would rather see fewer comments with clear repro, affected behavior, and test evidence than a long list of plausible-sounding notes the team has to triage again.

0
回复

Finally, a way to review what my AI agent actually wrote without scrolling through terminal output. The line-anchored comments feeding straight back through MCP is genuinely clever.

0
回复

One-time purchase for a tool that sits in the hot loop of agent coding is a refreshing call. The mechanic I want to understand: when the agent answers a comment by rewriting the whole block, where does the anchor go, does the comment follow the moved logic or die with the old lines? That resolution step decides whether review rounds stay readable after round three.

0
回复

@vollos Great question! In that situation, comments can’t stay with the new lines so they get archived into a “Past” tab in the sidebar (you can click it there to see the original lines it was anchored too).

Agents can add a comment to the newly changed code if they want to (and in my experience they often do because they understand that the original comment will be archived)

1
回复