Product Hunt 每日热榜 2026-07-28

PH热榜 | 2026-07-28

#1
Prefactor
Evaluate your AI Agents in real-time
552
一句话介绍:Prefactor 是一款实时评估生产环境中 AI Agent 运行质量的工具,帮助工程团队在 Agent 执行过程中即时捕捉质量退化、数据漂移和风险,并支持自动拦截或人工干预,解决了“Agent 在测试中通过、上线后失控”的核心痛点。
SaaS Developer Tools Artificial Intelligence
AI Agent 监控 实时评估 生产环境 质量退化 数据漂移 风险管理 人工干预 LLM-as-Judge 可观测性 DevOps
用户评论摘要:用户主要关注:1)实时评分是否覆盖全量流量(确认支持100%评估,成本可控);2)如何定义“质量”(团队可自定义评估标准,支持技术指标与业务结果);3)能否验证外部副作用(如页面是否实际更新,支持通过 emitted span 进行真实结果校验);4)漂移检测的实际价值(多位用户表示曾因模型无声变化导致业务受损);5)对“Agent 谎报成功”问题的真实需求(用户反馈约1/3任务看似成功实际未执行)。
AI 锐评

Prefactor 切入的是一个极具痛感却尚未被充分解决的“最后一公里”问题:AI Agent 的“生产事故黑盒”。其核心价值不在于另一个观测仪表盘,而在于“观测-评估-行动”的闭环——在 Agent 运行的毫秒级窗口内,将评分直接转化为拦截或人工决策信号,从而让“上线后听天由命”的团队真正拥有止损能力。

产品设计层面,技术路径清晰:通过 SDK 或 OpenTelemetry 实现运行时全量追踪,支持 LLM-as-Judge 与自定义规则混合评估,并允许将外部系统状态(如 GitHub、数据库)作为评估锚点。这解决了两个关键矛盾:一是“全量评估 vs 成本”的权衡(零 Token 成本的基础追踪 + 按需引入 LLM 评估);二是“内部状态 vs 外部事实”的不对称(Agent 声称成功 ≠ 真实结果生效)。

但风险同样不容忽视。首先,“实时干预”对延迟和可靠性要求极高,若出现误判或漏判,可能比无监控更糟糕——因为在信任工具后,团队可能会放松手动检查。其次,用户对“质量定义”的疑虑揭示了产品的边界:Prefactor 提供的是评估框架与执行引擎,而非开箱即用的行业标准模型,这意味着初期部署需要团队投入定义成本。最后,竞对威胁明显——LangSmith、Arize 等工具同样在抢占可观测性市场,而 Prefactor 的差异化在于“闭环行动”而非“只展示问题”。

一句话总结:Prefactor 解决了 Agent 从“能跑”到“跑好”的监管鸿沟,但其能否成为 AI 基础设施的标配,取决于能否在“控制”与“误报”之间找到平衡,并快速让用户感受到“止损”带来的 ROI,而非“又多了一套需要维护的规则系统”。

查看原始信息
Prefactor
Most agents pass their evals and fail in production. Prefactor is the evaluation layer that closes the gap. We score every agent run in real time, surface quality regressions and drift as they happen, and show engineering teams exactly how their agents are performing at scale. Built for the teams shipping agents to customers.

Hi Product Hunt - Matt here, co-founder of Prefactor with Simon.

We've been heads-down on this one for a while, so finally getting to show you is a real thrill.

Let me start with the question this whole thing is built around: do you actually know what your agents are doing in production right now?

For most teams we talk to, the honest answer is "…not really." Your evals pass, everything's green, you ship - and then every real run vanishes into a black box. Quality quietly drifts. Risk creeps in. Costs climb. And eventually someone asks which of your agents are still doing their job, and the room goes quiet.

That silence is the entire reason Prefactor exists. Gartner reckons 40%+ of agentic AI projects get scrapped by 2027 - and honestly, this gap is a big part of why.

So we built the thing we kept wishing we had: Prefactor scores every run in production the moment it happens - quality, drift, risk - then wires those scores straight into action. A failing agent gets caught live, not charted three days later.

How it actually works

  1. prefactor init - one command connects your workspace and discovers your agents across your runtimes. First traced run in under 5 minutes.

  2. Drop in the SDK (TypeScript or Python) - native for LangChain, Claude, Vercel AI, OpenClaw and LiveKit. Every call becomes a span, streaming in live with cost and data risk attached.

  3. Run the evals you define on every run - LLM-as-judge, technical checks, qualitative metrics. Custom spans pull context from GitHub, Linear, Jira or your database, so every eval is grounded in what actually happened, not a guess.

  4. Act. Hold, approve or block the second a run crosses a line - automatically at runtime, or routed to a human. Every decision logged and enforced through the SDK or API.

The payoff: you get to ship agents like real software — versioned, staged, promoted through dev → staging → prod only when evals pass, with instant rollback when they don't.

Why it's different

Most tools observe and score, then hand you the problem. Prefactor closes the loop - observe, evaluate, act, all inside the same run. A risky agent gets caught, not just charted.

My favourite bit of feedback so far: a customer with 40 agents in production and, in their words, "no honest way to say which ones were still doing their job." We gave them that answer - and the brake pedal for when one wasn't.

Who it's for

Engineering teams shipping agents to real customers, on any stack. Agent frameworks work natively; everything else plugs in through OpenTelemetry or the core SDK.

A little something for the PH community

Sign up today and you get 1,000,000 free agent steps. Every span counts, so that's a serious amount of live production evaluation on us. Valid until Friday 11:59pm PT, once you've set up your first agent.

Our ask

If you're running agents in production, tell us how you keep tabs on them today - even if the honest answer is "we're mostly hoping." We'll be in the comments all day and we'd genuinely love to hear what's working, what's breaking, and what you'd want a tool like this to do next.

Get started free at prefactor.tech. First 25,000 spans a month free, no card needed.

@simon_russell1 @joshgillies @ethan_lee8 @rheu @joeys and I will be here all day.

Big thanks to @rohanrecommends @rohanchaubey4 hunting us.

28
回复

@simon_russell1  @joshgillies  @ethan_lee8  @rheu  @joeys  @rohanrecommends  @rohanchaubey4  @matt_doughty Love this. I wouldn't leave my toddler home alone so why would I deploy an agent and not checkup on it. You get it. Cut through the agentic AI hype and solve the issues so we can actually get ROI on AI.

7
回复

@simon_russell1  @joshgillies  @ethan_lee8  @rheu  @joeys  @rohanrecommends  @rohanchaubey4  @matt_doughty Congrats on the lauch! I would like to know through what process your product identifies issues with the proxy program? I am very curious about the underlying technical logic.

0
回复
0
回复

Genuinely curious what "scoring every run" looks like at scale, is that sampled or are you actually evaluating 100% of production traffic? That distinction matters a lot for cost and for trust in the numbers.

9
回复

@yolanda_c_schneider All traffic. The 100% element is sitting alongside the agents, seeing everything they do and being able to articulate what they are actually doing. Eg, what tools are they calling. Are the actions they are making write actions vs read actions. How is the pattern of the agent different from what you know to be the case. If an agent takes 3 turns to resolve a problem, why is it taking 8 (insert HITL).

The beauty of our approach is that the cost is 0 until you introduce an LLM as judge, but by using the no token approach it means you can actually target the problematic runs/ spans/ actions rather than running random samples which could catch nothing when there is something.

3
回复

@yolanda_c_schneider right now, we score every run. This may change in the future, and if you have a preference we'd love to hear your thoughts on the matter! From a cost perspective, the data that goes into the scoring is entirely up to you. We score what you send us about your production agents, and generally the rule of thumb is something like, trace the parts of your agent required to answer the question of "did it do good job".

1
回复

@yolanda_c_schneider There can be different approaches depending on the application. Obviously running full LLM-based evals on a high volume chatbot might not make sense. But that's when being able to filter down the conversations you do choose to run deep analysis over makes sense. You can use heuristics to do that, then use that to guide things.

3
回复

Drift is the one nobody talks about until it's already cost them a bad week. Real detection here feels like the right instinct.

9
回复

@ado_audu Absolutely does. When a model can change underneath you without a notification, how can you possibly contain drift. Being able to sit inside the agent is absolute gold dust.

0
回复

@ado_audu this guy gets it! How are you monitoring this drift in production currently, and whats preventing you from adopting something like Prefactor?

0
回复

@ado_audu Thanks Michael! Drift is the one that hurts because there is no incident, no alert, no obvious moment where it broke.

Just a slow slide that someone notices a month later in a customer complaint.

Have you experienced this too?

0
回复

Congrats Ethan and team!

9
回复

@benln appreciate the support 🙏

0
回复

@benln Thanks Ben!

1
回复
@benln thanks Ben! 😁
1
回复

The line "most agents pass their evals and fail in production" is going to resonate with more teams than you probably realize. I've said some version of that sentence out loud in at least three postmortems this year.

8
回复

@billy_boy Hey Billy, thanks for comment. You're right. I find that most companies don't even realise the difference.

I can tell someone what their evaluation approach is before they've even opened their mouth because they are so standardised. Sampled runs, LLM as judge and golden dataset. All are point in time.

They have their value in dev for sure. But in prod, you need a suite of tools which reduce the need for tokens and non-deterministic outcomes. Being in the agent run/ live, allows us to do that.

Would love to get your feedback on the tool and learn if it would have solved some of those PMs earlier in the year

1
回复

@billy_boy It's such a big problem to solve -- so many projects failing, and it really doesn't have to be that way. We're trying to put an end to those post-mortems!

2
回复

@billy_boy Three postmortems is a rough year.

That sentence came out of the same place for us too, watching teams ship something that scored well on a fixed eval set and then quietly degrade over weeks with nobody able to prove it.

What was the failure mode in those postmortems btw, was it drift or something breaking outright?

1
回复

I spent months building internal scripts to catch exactly this kind of drift. Would've saved me a lot of late nights to just plug into something like this instead.

8
回复

@malka_parveen many such cases unfortunately... Which is the reason we have built Prefactor, to save your engineering team the time and burden of building, and maintaining a platform like this. We should compare notes at some point, would love to know if we missed anything you'd consider essential for keeping your agents on the straight and narrow.

1
回复

@malka_parveen Thanks Malka! Would love for you to try it out and get your feedback!

What sort of agents have you built?

1
回复

Hi everyone, Joseph here. I’m a designer at Prefactor.

Something I’m particularly interested in that’s easy to overlook: the moment Prefactor flags something mid-run, a human has to look at a screen and decide whether to step in or let it ride. That decision is only as good as the interface it happens on.

Agents generate an enormous amount of data, and most tools just show you all of it. My job is making sure that when something's going wrong, you can tell in seconds, not after scrolling through a wall of spans. Real-time control needs real-time legibility.

If you try Prefactor, I’d love to know: did you know where to look, or did you have to dig?

8
回复

Hey guys Ethan here, I'm part of the team at Prefactor.

My main focus is on the go-to-market side which means I pretty much spend my week on calls with teams and engineers running agents in production.

I usually hear the same stuff all the time.

For example: The agent worked in testing and POC, it went live, but now nobody can answer a basic question: is it doing what we told it to?

That’s where we’d come in, Prefactor gives those teams visibility into what their agents are actually doing in production, so accountability doesn't stop the moment it's in production.

If you've got an agent running live right now, what does your monitoring actually look like? Keen to know if anyone here has actually solved this properly and if so how.

Cheers!

8
回复

Hi folks, I'm Simon - CTO and other co-founder of Prefactor. The real challenge for companies now is not building agents, but managing the infrastructure and process around them. A lot of the lessons learnt from deploying traditional software at scale still apply, but agents pose new and unfamiliar challenges. It's a combination of traditional devops and HR. The approaches to risk and quality that have worked in the past need to be updated.

Prefactor makes this approachable for any engineering team. Closing the loop on turning feedback into improvements to the agent; ensuring consistency over model and prompt updates; simple ways to understand and contain risk; ways to monitor and control the actions of agents across frameworks and deployment environments.

8
回复

Friends, Josh here. AI Engineer @ Prefactor. I built the agent instance tracking that scores quality and flags data risk while your agent runs, plus the SDKs and docs that get teams from zero to live without guessing.


Here's the thing I keep coming back to. You can't have confidence in something you can't see. An agent in production is a bucking bronco; it'll throw you the moment you stop paying attention. I've talked to too many teams who deployed, watched it work for a week, then realised they had no idea what it was actually doing. No visibility, no guardrails, just hope.


That's the piece Prefactor owns: the quality and data risk scoring that runs in-process, flagging misbehaviour mid-gallop rather than in a trace you read after the damage is done.


If you've got agents live right now, how do you actually know they're behaving? Or are you just holding on and hoping?

7
回复

@joshgilliesan agent in production is a bucking bronco is uncomfortably accurate. our version of just holding on and hoping" was checking reply-sentiment after the fact by the time we noticed an outreach sequence had drifted (wrong tone, wrong list segment), it had already gone out to way more people than one bad decision should reach. the thing that would've actually helped is exactly what you're describing catching it mid-run, not in a post-mortem. quick question: does Prefactor's scoring work for actions that leave your own system (sends, API calls to third parties), or is it mainly for reasoning/tool-use inside the agent's own loop?

0
回复

How do you define "quality" here? That word does a lot of work and I imagine it means something different for a support agent than a coding agent.

6
回复

@sheikh_umair1 Absolutely right. SO we see risk and quality as two sides of the same coin. We break them out because if an agent is making a write action on a tool call, then its both a quality problem if it goes wrong and a risk problem if it doesn't land properly and affects teams reliant on the agents.

Quality to us is technical metrics, agent patterns - literally is it calling too many tools or looping too frequently but the real end game, and why we exist, is to answer the question, how do you know what your agent is doing? And more importantly, how can you tell if it doesn't do what you expect it to.

1
回复

@sheikh_umair1 That's very true, and it's why we've designed the system to be flexible -- so you can decide how you want to approach it. The diversity in evaluation frameworks just for coding agents is a sign of how complex (and fast evolving) it is.

3
回复

@sheikh_umair1 Very good question, quality is doing a lot of work in that sentence.

In practice you define what good looks like for your specific agent rather than us handing you a universal definition, because you are right that a support agent and a coding agent share almost nothing.

The common layer I'd say is the scoring and enforcement, not the criteria.

What would you want it measuring for the agent you are running?

1
回复

How does Prefactor decide what parts of the code need attention without creating unnecessary changes? Congrats @ethan_lee8 & team!

6
回复

@ethan_lee8  @hamza_afzal_butt We dont claim to know every agent and how they work. We give you the tools to track things as they go wrong, whether thats through meta data, llm as judge evals etc.

We help the unnecessary change bit by version controlling each agent, enabling different environments so you can track the problems as they occur in dev and when you're ready push them to prod.

2
回复

@hamza_afzal_butt Hey Hamza thanks so much for the support!

Slight clarification, we don't touch your code. Prefactor sits at runtime and records what the agent actually does on every run, then optional guardrails can hold a high-risk action before it executes.

So it's less "which code needs changing" and more "which agent stopped doing its job, and stop it now."

Are you running any agents in production at the moment btw?

1
回复

Hey Product Hunt, I'm Mukil — an engineer at Prefactor. 🔥

People always ask me:

"Mukil, you work with AI all day, isn't it crazy what it can do now?"

Sure, it's pretty cool. But I spend most of my time on the mostly ambiguous half of that sentence. Yes, it did something but was it the right thing? Did it actually do it or did it just say it did? Did it just hand someone the account details for the completely wrong account?

When an agent screws up, it often doesn't like to stop. It keeps going, fully confident and you find out hours or days later. And at that point its often when a customer's already upset or a refund went out that absolutely should not have. As an engineer this made me a little annoyed. The problem was never that I couldn’t see what the agent did. Dashboards will happily show me — right after it’s already done something I didn't want it to do. It's already leaked that PII, emailed the wrong guy or somehow spun up 15 subagents to run a database query.

With Prefactor you can see runs happening live, stopping the agent in its tracks. Every run gets scored live and lets the human know to step in if they need to. Instead of waiting 30 minutes to see this failed run on your custom dashboard (it's very pretty, I'm sure), you can step in and kill it while its executing.

In short, almost anyone can build an AI agent these days. And as the barrier to building them gets lower, the standard for running them well needs to get much higher.

We obsess over that second part, so you can spend less time investigating what your agent did at 3:14 a.m. and more time letting it do actual work.

Would love for you to try it and tell us your experience with it!

6
回复

Ethan asked how people keep tabs on live agents, so here is my honest answer. I run scheduled agent jobs daily for SEO reporting and roughly one in three submits used to report success while the page never actually changed, so every job now ends by rereading the rendered result and comparing it against what the agent claimed. Mukil's question of did it actually do it or did it just say it did is the exact failure I kept measuring. Can an eval in Prefactor check an external side effect like that, the page actually updated or the row actually written, instead of scoring the run transcript?

5
回复

@abdullah_javaid3 Thanks for your comment. Short answer: yes, and it's basically the loop you already built.

The re-read-and-compare step becomes the eval signal. Your job verifies the real outcome (re-reads the rendered page, or queries the row) and emits that as a span, and the eval scores against that ground truth instead of the agent's self-reported "success." So Prefactor never has to trust "I did it", it scores what actually landed.

That's exactly Mukil's line: the transcript is the claim, the emitted outcome is the verification, and the eval only counts the second one. Your 1-in-3 false-success rate is precisely what surfaces once you're scoring outcomes instead of runs.

4
回复

@abdullah_javaid3 Interesting question. Alongside the LLM activity you can store whatever you like; also you can store quality assessments against an individual run. Usually the goal is to understand what the success rate is, then use that information to improve the prompting and harness. For your situation a hybrid of prompting and deterministic assessment makes sense.

4
回复

@abdullah_javaid3 The 1-in-3 false success number is pretty crazy. how often does the re read catch something the transcript said was fine? and when it does catch one, is the fix usually a prompt change or was the tool call itself wrong?

asking becuase i keep finding teams who built the same verification loop by hand and nobody talks about it.

0
回复

Fantastic product and team, very excited to see them launch here, congratulations on the successful launch of the product!

5
回复

@matthew_browne1 Thanks Matt. It means a lot coming from you! Thanks for all of the support. To the moon!!!!!!

2
回复

IMO the 'act inside the same run' is what actually separates this from the dashboards that just score runs after the fact. Charting a bad run three days later never stopped anything. QQ - once an eval can hold or block at runtime, that eval is sitting in the critical path - how much latency does the inline check add before it lets a step through? congrats on the launch!

4
回复

@artstavenka1 generally the approach is to keep everything async in the runtime, this is to ensure minimal latency is introduced into production agents. Would you favor a different approach for the agents you manage?

2
回复

@artstavenka1 Great question. There are a couple of approaches, depending on your needs. You can choose to make your agent wait for permission to continue (synchronous) or just log the spans, and then that termination/feedback signal can come at a later stage when it's ready (asynchronous). Obviously that isn't the right approach in all scenarios but it can remove any possibility of latency. The other part of it is choosing how you're judging quality at different stages.

3
回复

@artstavenka1 Hey Art thanks for the question! I'll let @simonru or @matt_doughty answer this :)

1
回复

kind of worried this just becomes the new green checkmark people stop questioning. same problem, one layer up?

3
回复

@kellyops Thanks for the question.

I see a world where we are able to leverage multiple layers of evaluation framework to reduce chance of it going wrong. It's still software though, and even the best softtware breaks. But the key thing is knowing why it broke, when it broke, and who is the person responsible. It's remarkable how those simple questions can't be answered by so many companies we speak to.

Would love to understand what else you think is needed to fix that? Matt

0
回复

@kellyops Very valid concern!

Any scoring layer can turn into a checkmark people stop questioning, so what matters imo is whether you can see why a run scored the way it did and disagree with it, not just that a number exists.

We would rather be wrong loudly than right silently. What would make you actually trust a score rather than just accept it?

1
回复

I really like the idea of tracking how the agent is working, as it is performing actions. Being more proactive is better than finding out something went wrong several days ago. It's not to be overlooked that the interruption isn't just killing the agent run, but handing it to a human too. I am going to have to play around with this and my own agents.

3
回复

@philnash Thanks for the comment Phil. Would be happy to share some of the examples Ive built too, I'm building an eval triaging tool. 3 phases.

First phase: Tokenless Evals/ technical reviews
Second phase: pattern management against prior activity

Third phase: LLM as judge

All live. Let me know where you get to.

1
回复

@philnash Thanks Phil, exactly! Killing a bad run is only half of it, the handoff to a human is where the actual value sits, otherwise you have just built a more expensive way to fail.

Would love to hear what breaks when you try it on your own agents, that is the feedback we learn the most from.

What are you running them on at the moment?

1
回复

The runtime controls are the strongest part here. Being able to hold or block an agent when an eval fails turns evaluation into an actual safety layer, not another dashboard.

How do teams prevent a noisy or misconfigured eval from repeatedly stopping healthy production runs?

3
回复

@adityaharish2002 Absolutely. Thanks for the comment.

The beauty of the architecture, @simon_russell1 is being able to hold, insert a hitl/ agent, kill an agent mid flow. Really we would be encouraging anyone who inserts a killswitch to be absolutely sure that its the right approach.

It becomes a tiered model, where you only employ the strongest action once it's already gone through the other steps to avoid healthy runs being stopped.

Eg, How many loops have happened, how many tool runs have occurred against the norm, how long has the run gone on.

There is an argument you could have a healthy run and it just cost a lot, so you need to be clear on what those boundaries are.

1
回复

'no honest way to say which agents still do their job' nails it. when the llm-judge itself drifts, what keeps the eval honest?

3
回复

@andrewzakonov exactly right, who judges the judge, ad infinitum. How are you managing this at scale currently?

1
回复

@andrewzakonov Our core approach is tokenless. We/ you use the agent run to spot problems, risks, abnormal behaviour. When an LLM is needed, then you bring it in on demand rather than sampling. The idea is that when an LLM is needed, its being brought in for a specific task to confirm what the tokenless approach has already found.

You could apply the same logic to bringing in a HITL when a customer support conversation is going off the rails which is discovered through sentiment analysis (tokenless)

1
回复

@andrewzakonov That's actually a great point. Making a good judge LLM agent is actually a big challenge too - it needs just as much engineering as the thing it's judging. Luckily you can instrument the LLM-as-judge agent with Prefactor as well. (I guess you could instrument that recursively but you might run into diminishing returns at some point!)

2
回复

Can I use Prefactor for my fleet of Hermes and OpenClaw agents ?

3
回复

@nathan_ngz we officially support integrating with OpenClaw, and I have it on good authority that Hermes is very easy to instrument also. When in doubt, just point your agent at our docs, and SDKs on Github and let it cook!

1
回复

@nathan_ngz You can indeed -- there's an OpenClaw plugin. There's no Hermes plugin at this point but we have seen someone get Hermes to integrate itself with Prefactor to great effect!

2
回复

@nathan_ngz OpenClaw yes, Hermes not yet

1
回复

Working in high-trust environments, I don’t think the future is humans approving every AI action. It’s agents operating independently 95% of the time, with enough visibility and confidence that when they drift outside the guardrails, a human can intervene before it matters. Looks like you’re tackling that problem head on. Congrats @matt_doughty @simon_russell1 and team!

Send me my 1M spans to give it a solid rev up! ;)

3
回复

@simon_russell1  @emotf Done. You've got to set up an agent first and then you get them in your account.

You were one of our first conversations wayyyy back. LFG!

2
回复

@matt_doughty @emotf Thanks! Those 1M are yours if you sign up now :D

2
回复

Congrats on the launch.
When teams try Prefactor for the first time, how do teams catch the first bad run before users notice?

2
回复

@dmitrii_volosatov hey mate, thanks for the question. Every run is risk assessed, so if you see something in the run representing a high risk, you can insert a hold into the agent via the spans. Which stops the agent continuing. Or you can either manually or automatically insert a killlswitch, so the short answer is the customer should never see a bad run.

How are you handling this today?

1
回复

@dmitrii_volosatov Thanks Dmitrii! Short version is we score runs as they happen rather than sampling after the fact, so a bad run gets flagged while it is still executing instead of surfacing in a dashboard the next morning.

Most teams start with loose thresholds, watch what actually trips in week one, then tighten.

Are you in production already or still at POC?

1
回复

The "tokenless until you actually need an LLM" design is the part that stands out, most eval tools default to LLM-as-judge for everything and eat the cost whether it's warranted or not. Given you're scoring 100% of production traffic, curious how Prefactor handles agents that call other agents (sub-agent chains). Does a breach in a nested sub-agent bubble up and pause the parent run too, or does each agent in the chain get evaluated as its own isolated instance?

2
回复

@mittalpatel in the current design every agent whether it's a top-level agent or a sub-agent are treated as bespoke agents within Prefactor ie. they have their own identity, with spans attached to each of their own instances. So a breach in a sub-agent would only effect that agents run. You can of course decide whether that's also a stop condition for the parent, but that's very much up to the end user to decide. Admittedly sub-agents is something we've trialed a lot of design ideas for, and are admittedly not satisfied yet. That said we're always open to suggestions! What would your preference be for a case like this?

1
回复

@mittalpatel Thanks for a great question. We treat everything as a chain. Every reaction, step, call. All in one view.

Sub-agent break outs might occur in future versions if customers want it, but my gut tells me that people want to be able to see the whole workflow, subagents included. How they get represented is another question, but it doesnt make sense to cut them up.

0
回复

That's exactly the gap that matters. An agent can produce a perfect transcript and still fail the task in the real world. We've seen jobs report "success" while nothing actually changed. That's why the real evaluation isn't what the agent said—it's what actually happened. Can Prefactor validate external outcomes, like confirming a page was updated, a record was written, or an API state changed, instead of only grading the execution log?

2
回复

@md_khayruzzaman Absolutely -- we've designed it so that those outcomes can be recorded alongside the agent activity. They are a vital part of understanding the full outcome of an agent run.

1
回复

@md_khayruzzaman The way you'd handle that would be either connecting Prefactor directly into the MCP of your product and then associating the actions with spans inside the run, meaning you would be able to connect outcome to evaluation/ run quality.

Other ways might be to send completed outcomes directly into Prefactor so you top and tail the process. Would love to talk to you about how you would want that to work. The idea of self improving agents only works if you have the outcome connected to drive next steps.

0
回复

scoring the run live instead of charting it three days later is the part that actually matters. catching a bad agent after it already shipped the damage is just reporting.

2
回复

@alex_watson2110 Totally agree -- evaluation needs to be part of the control loop for an agent to really be trustworthy.

1
回复

@alex_watson2110 100%, are you building many agents yourself?

1
回复

I like that this is built around production instead of just benchmarks. That's where things usually get interesting. Congrats on the launch!

2
回复

@henry_habib Thanks Henry, and thanks for the comment. Benchmarks are great in controlled envs. Agents don't operate in hermetically sealed containers. They exist in the wild. We needed to meet that challenge.

1
回复

@henry_habib Thank you! Yes, production is where things tend to get real pretty quickly. Understanding how your benchmarks/evals translate to the real world is super important.

2
回复

@henry_habib thanks for the comment, Henry! We agree, an agent demo on a developers laptop is one thing, but once it hits a production environment that where the fun really begins. Are you deploying anything into production currently?

1
回复

I like that Perfector focuses on real production behavior instead of only passing evaluations. Many teams discover issues after release. How do you help teams trace the exact reason behind a quality drop across different agent workflow?

2
回复
@john_michael31 thanka for the comment! you can see every part of the agent run in Prefactor. Were API driven so you can see every span, every call, every agent step and act on them. Run live evals across any run or in line. it's being inside the SDK that allows us the the ability to catch any drift in real time.
1
回复

@john_michael31 There are various approaches, but the core of it is being able to store quality metrics against runs, and then extract the full information about those runs. Depending on the sort of agent you're building, your coding agent can use that directly to improve things.

2
回复

What kinds of failure modes or edge cases have you seen most often when an agent passes evaluation but fails in production, and how would real‑time scoring and drift alerts need to look for your team to trust them enough to act automatically?

2
回复

@swati_paliwal Thanks for your comment Swati.

We are not just focused on edge cases. The first thing is are the agents being actually assessed against whether they are doing their job.

That contract you've built for the agent is key. Eg, what tools are they allowed to call (or not in the case of OpenAIs agent), is it looping? Are parts of the flow failing causing excessive token usage.

Evals are point in time reference points for the agents themselves. When models change, when databases get different answers, drift happens. When sampling involved 1% of all agent runs or less (and given the cost of tokens, it is increasingly less), its not the 1% that will kill you. Its the 1% of the 99% you don't eval.

1
回复

Congrats on the launch! Does your tool make agents more smart over time by building right context around them, or is it still a responsibility of an agent?

2
回复

@nikitaeverywhere Hey Nik. Great question. The self-improving loop is something within its capabilities but not something we currently have as native.

I ran a POC for a customer where they wanted to have their rag database self improve. I did that by setting up a HITL when a conversation went badly via sentiment analysis running over every span. The rag database would then update with the answers drawn from the agent. So the rag db was self improving.

Everything is configurable so you can literally go in and do all sorts. I am passionate about working out how to use non token evals to allow 100% coverage and only use the token coverage for the triaged cases.

Would love to talk more when you've given it a go.

1
回复
#2
Cekura
The self-improvement loop for voice agents
369
一句话介绍:Cekura是一个为语音与聊天AI代理打造的生产环境测试、可观测性与自我修复平台,通过模拟数千场景、自动诊断故障并闭环修复Agent,解决开发团队手动调试低效、修复易引发新Bug的痛点。
SaaS Developer Tools Audio
AI代理测试 语音Agent 可观测性 自修复循环 回归测试 根因诊断 CI/CD集成 生产级评估 Prompt优化 过拟合防护
用户评论摘要:用户高度认可其自动复现Bug再修复、回归扫描防止过拟合的核心机制。核心关切包括:模拟场景的逼真度(如噪音、打断)、能否处理外部系统确认与Agent行为的分离、CI/CD自动门禁的集成程度、以及修复泛化能力的验证方式。
AI 锐评

Cekura解决了AI Agent开发中“测试-修复-打补丁”的无效内耗循环,这确实是当前工具链中一个被忽视但极其痛苦的空白。其价值核心并非简单的自动化,而是**将“诊断”和“修复”两个环节用闭环逻辑串联起来**,并从“优雅回放”的魔法转变为可验证、可回滚的工程规范。

值得一提的是,它严格设定了“复现错误后才能修复”的铁律,并在克隆环境验证,杜绝了“自愈”工具常引入的“黑箱”隐患。这解决了工程师最后1%的信任顾虑。然而,真正的挑战在于**复杂语境的保真度**——模拟的“噪声、打断”与真实用户的混乱、情绪化状态存在天壤之别。用户评论也尖锐指出,若场景集成为优化目标,Agent可能在测试中“高分低能”,尤其在处理那些“Agent做了正确操作但外部系统未接受”的复杂业务场景时,Cekura的“逐句转录+结果信号”能否持续精准区分“对话成败”与“业务成败”,尚存隐忧。

当前产品更像一件锋利的**外科手术刀**,精准但入门门槛不低。它要求团队已有成熟、可复现的测试场景定义。对于多数尚在“手工调试”阶段的小团队,从0到1构建高质量评估集本身就是巨大挑战。Cekura更像是一个“测试基础设施”,而非“即插即用的AI调优师”。其长期价值取决于能否持续降低用户构建和验证场景的成本,以及能否真正输出一个“通用化”的场景库,而非沦为特定团队的成本中心。

查看原始信息
Cekura
Cekura is the testing, observability, and self-improvement platform for production voice and chat AI agents. It simulates thousands of scenarios, catches failures, diagnoses the root cause, rewrites prompts and config, then re-validates with a full regression sweep. Unlike tools that hand failures back to your team, Cekura closes the loop by fixing the agent itself and proving the fix holds without overfitting.

Hey Product Hunt!

Sidhant here, co-founder of Cekura

Today we are launching self-improving loops for voice agents.

Fixing a voice agent has always been fragmented. Your testing tool tells you what failed, you diagnose it from transcripts, patch the prompt, re-run, and something else breaks. The tools find problems, but the fixing has always been a human walking between them.

Cekura collapses that loop. It runs thousands of simulated calls and groups every failure, explained in plain English. Optimise agent hands them to the Cekura Agent, or any coding agent you use, Claude Code, Codex, anything. It reproduces each failure, makes the change, and reruns until everything passes, then verifies nothing else broke. You review the diff.

Two rules: it must reproduce a bug before fixing it, and every fix is proven on simulated calls on a clone, never your live agent.

Free for everyone to try, starting today. We are in the comments all day. If you want to chat more, please feel free to book time here

12
回复

@kabra_sidhant must reproduce a bug before fixing it is the rule that actually matters here most auto-fix tools skip straight to patching and you end up trusting a fix you can't verify actually addressed the real failure. curious what happens when it can't reproduce something flagged in production does it flag that as a separate category, or just drop it?

0
回复

@kabra_sidhant Do most teams let Cekura suggest fixes, or do they mainly use it to identify issues?

0
回复

Root cause diagnosis for voice agents is genuinely hard because failures compound across turns. If it traces that accurately, that alone is worth the pitch.

9
回复

@peter_victor Completely agree. We trace the full conversation to find where behavior first diverged, rather than patching only the final symptom.

3
回复

Reading this made me think about how much time my team spends manually reproducing a bad call just to figure out what went wrong. Automating that alone would be huge.

8
回复

@ramish_saje Couldn't agree more. Hunting down logs to replay a bad call takes way too much time. Automating that reproduction and root-cause diagnosis was step one for us. Really appreciate the support!

2
回复

What stands out to me is the regression sweep after every fix. I've seen too many "self-healing" tools introduce new bugs while patching old ones. Curious how it measures whether a fix actually generalizes across edge cases.

7
回复

@kimberly_west Great question. Every fix passes through an Overfitting Gate first to catch issue-specific or overly narrow fixes, then gets re-validated with a full regression sweep across the evaluator set

1
回复

How does Vocera handle accuracy and context when conversations get more complex or users have different speaking styles? Congrats @kabra_sidhant & team!

7
回复

@hamza_afzal_butt Thank you! Cekura tests agents across varied scenarios, speaking styles, interruptions, and multi-turn context—not just isolated turns. It also preserves conversation history and evaluates whether the agent completed the intended outcome, so accuracy holds up as interactions get more complex.

3
回复

@hamza_afzal_butt Our evaluators are written against intent, not exact wording — conditions like "the agent asks for your name" match semantically regardless of how a specific agent phrases it, and grading runs over the full conversation using LLM/rubric-based judging rather than pattern matching. So the same test generalizes across agents with different tones, styles, or conversational complexity, without needing to be rewritten per agent.

1
回复

I've been burned by agents that pass QA and fail in production because the test scenarios were too clean. Genuinely curious how realistic it's simulated conversations sound.

7
回复

@almuddin_ansari Totally agree. We test interruptions, background noise, long pauses and unexpected responses, not just clean scripted conversations.

4
回复

@almuddin_ansari Great question , our simulated conversations use Conditional Actions, so the testing agent reacts dynamically to what the agent actually says (branching, interruptions, background noise, live data), not a fixed script. Please do check out our docs for more: https://docs.cekura.ai/documentation/key-concepts/evaluators/conditional-actions

2
回复

Voice agents are unforgiving compared to chat, a bad pause or mistimed interruption sticks out immediately. I appreciate that it treats voice and chat as needing the same rigor rather than bolting voice support onto a chat-first tool.

5
回复

@james_will1 Thanks! That's basically what we built Cekura for.

Cekura doesn't answer the calls, it stress tests the agent that does. We generate personas that talk like real users: different accents, interruptions, mumbled order numbers, background noise, code switching mid sentence.

And we score full conversations, not single turns. Did the agent hold state after a topic switch? Did it recover from mishearing something on turn 3?

When it fails, you get the root cause and the fix passed to your coding agent, not just a red flag on a dashboard.

2
回复

@james_will1  This is the exact thing we refused to compromise on. The moment you evaluate a voice agent by reading its transcript, you've deleted the dimension where it actually fails — timing, barge-in, overlapping speech. So voice is tested at the audio layer, not as chat with extra steps. Appreciate you calling it out.

2
回复
Love what you and the Cekura crew are building! This is truly solving the problem end to end. I got struck on this loop even for a simple voice agent I built with ElevenLabs for my personal website. So getting this solved for mission critical voice agents is a game changer! Rooting for you guys :)
5
回复

@satiwick1 Thanks for your encouraging comment. Means the world to us. If you get a chance do try cekura for your agent at cekura.ai

3
回复

The rule that a bug has to be reproduced before it gets fixed, and validated on a clone instead of the live agent, is the part that stands out to me. I build a voice companion that calls aging parents daily, and my worst fear is a silent regression in an emotionally sensitive call that nobody catches until it actually matters. How do you generate realistic edge cases for messy real speech (long pauses, hearing loss, someone talking over the agent) versus clean synthetic voices? That gap is where most of our failures hide.

4
回复

@igorgurovich That's exactly the gap we designed for. Our simulations use Conditional Actions to reproduce real pauses, mid-sentence interruptions, and real-voice recordings, not just clean TTS , along with personalities for speaking pace and background noise. That's where messy-speech edge cases get caught before production.

3
回复

the loop-closing part is the interesting claim. the failure mode i would worry about is the simulator quietly becoming the thing you optimise against, so the agent gets very good at passing your scenario set and no better in production.

we run agents against third-party systems we do not control, and the split that ended up mattering most was "the agent did the right thing" versus "the other side actually accepted it". those two look identical in a transcript and come apart constantly in practice.

curious how the regression sweep decides a fix generalised rather than fit the scenarios. is the held-out set drawn from real production calls, or generated the same way the training scenarios were?

3
回复

@whateverneveranywhere That’s exactly the failure mode we guard against. We separate “the agent attempted the right action” from “the external system accepted it” using provider state, tool results, and outcome signals—not transcripts alone.

The regression sweep includes the full validation set plus happy/edge cases, with overfitting checks for scenario-specific fixes. Production-derived held-out cases are the strongest validation.

3
回复

It's like living in the future

3
回复

The rule that a fix has to be reproduced and validated on a clone before touching the live agent is the right engineering call for anything running in production customer support. The part I want to understand is the CI/CD integration: when a new prompt version or agent config gets committed, does Cekura run the regression suite automatically as a pipeline step that can block a deploy, or is the test run still something the team triggers manually? That boundary between an automated gate and a manual check is usually where teams cut corners under deadline pressure.

3
回复

@hazy0 Cekura evals can run as a step in your CI/CD pipeline via our GitHub Actions integration. We also provide a curated Infrastructure Suite that we recommend adding to that pipeline for infra regression checks.

2
回复

Love seeing the focus on proving that a fix doesn't introduce regressions elsewhere. Reliable AI systems need repeatable validation, not just faster debugging.

2
回复

Simulating messy real-world conversations with interruptions, pauses, and background noise feels much closer to production than traditional scripted evaluations.

2
回复

@pratham_patel8 Thanks, appreciate it

0
回复

Lets Go Team!!!!

2
回复

Congrats on the launch. The hard part of a self-improvement loop is the blast radius. When an agent rewrites its own behavior from production calls, a fix for one flow can quietly bend a compliance flow sitting next to it, and in healthcare that is the thing every security review hunts for. Closing that loop safely, so the agent improves without drifting out of its guardrails, is the whole game, and it is a genuinely hard problem to have taken on.

2
回复

@roguetink Thank you, really appreciate it

0
回复

the without overfitting part is doing a lot of heavy lifting in that description, and its the right thing to be worried about. an auto-fixer that cant prove the fix generalized is just prompt roulette.

2
回复

@alex_watson2110 Agreed , that's exactly the risk we designed against. Every applied edit passes through an Overfitting Gate before validation, then gets re-run against the full evaluator set, so a fix only counts once it holds broadly, not just on the case it was written for.

0
回复

I'm a huge fan of the work the Cekura team is doing. Closing the development loop with testing, evaluation, and paths to verifiable improvement is the single most valuable thing you can do in an agentic coding context. This is hard for voice agent development, for a bunch of reasons, including that there are many moving parts, gathering metrics robustly can be tricky and subtle, and success criteria are complex. Cekura has created great building blocks and high-leverage agent skills/workflows for solving these problems. Congratulations to the team on this launch, and all the effort and good thinking that went into it!

2
回复

@kwindla Thank you so much! We really appreciate your help with this

0
回复

Love seeing more attention on production reliability for voice AI. The regression validation especially stood out. Curious whether Cekura can prioritize issues by business impact (for example, failed payments vs. minor conversation hiccups) or if everything is treated equally.

1
回复

@yash_jain49 Great question. You can easily prioritize issues on which to self-improve your agent on Cekura. You can create custom metrics depending on your use case and can attach them to scenarios on which you are improving your agent and the self-improvement loop will consider the failure points on the attached metrics to improve your agent upon.

0
回复

Congrats on the launch! Really like that you’re focusing on proving fixes instead of just identifying failures. One thing I’m curious about: how do you decide when a suggested fix is reliable enough to recommend versus flagging it for manual review?

1
回复

@ujjval_goury Hi Ujjval the loop run a suggested fix multiple times to see if the fix consistently resolves the failure mode. If the loop isn't sure about anything then it hands to the user for manual review.

0
回复

Voice agents fail in ways text evals never catch — interruptions, latency, someone talking over the bot. The "loop" framing is the right one: testing voice once at build time is basically useless. Would love to know how many simulated calls it takes before the improvements show up.

1
回复

@lucasjpols Fair question , it's less about a fixed number of calls and more about iterations: the loop diagnoses, fixes, and re-validates until the full evaluator set passes. We've seen quite a lot of improvement starting from just 1-2 rounds itself

0
回复

Thanks — the GitHub Actions path is the answer to that question. The follow-up I'd have is whether the infra regression suite runs against a live agent clone or replays recorded sessions, because customer-support agents with integration state tend to behave differently on replay versus a live environment.

1
回复

@hazy0 Hi Hazy, every Infrastructure Suite case is a real simulated call against a real running agent.

0
回复

Does Cekura provide regression testing capabilities to catch quality drops when agents are updated or retrained?

1
回复

@james_khuzuma Yes , every prompt/config fix is re-run against the full evaluator set as a regression sweep before it's considered done, and the same evaluators can be run in CI/CD on every agent update so quality drops get caught before they reach production.

0
回复

How customizable are the evaluation metrics? Can teams define their own quality benchmarks based on their specific use case?

1
回复

@william_leon1 Very customizable , beyond our pre-defined metrics, teams can define their own custom metrics, including LLM-judge and Python-based ones for fully custom logic tailored to their specific use case.

0
回复

For teams building voice based agents specifically, does Cekura account for latency and audio quality issues, or is the focus primarily on conversational logic?

1
回复

@sophie_louis Both, not just conversational logic. We have dedicated Conversation Quality metrics for latency (with percentile breakdowns) and interruption/turn-taking timing, plus Speech Quality metrics for audio itself , pitch, clarity, jitter, unnatural speaking-rate changes , all computed straight from the call audio

0
回复

How does Cekura evaluate more nuanced qualities like tone, empathy, or conversational flow, beyond just accuracy or task completion?

1
回复

@louis_hallie Yes , beyond accuracy/task-completion, we have built-in metrics for exactly that: CSAT and Sentiment score the caller's tone and satisfaction, Verbosity and Unnecessary Repetition flag conversational flow issues like over-explaining or re-confirming the same thing, and Voice Tone + Clarity looks at delivery quality from the audio itself. For anything more specific to your agent's context, you can also define a custom LLM-judge metric

0
回复

Does Cekura support testing across multiple languages and accents, given how critical that is for conversational AI reliability?

1
回复

@eden_halls Yes , language and accent are both first-class testing surfaces. Personalities let you run the same evaluator across different languages (including code-switching, like Spanglish) and different regional/non-native accents, with a Transcription Accuracy metric that flags exactly where STT struggles per variant, so you can compare pass rates side by side instead of relying on one aggregate number.

0
回复

When monitoring production calls, how does Cekura detect quality issues in real time, and what kind of alerting or reporting does it provide to teams?

1
回复

@charlotte_henry For production monitoring, Cekura runs metrics continuously on live calls and clusters failures into recurring failure-mode themes automatically, so you see patterns instead of re-reading every call. On top of that, configurable alerts (failure, trend drift, threshold breach, new failure mode, failure-mode spike) route straight to Slack with the call and context attached.

0
回复

Nice launch! What CI/CD tool integrations do you support?

1
回复

@avz Hi Julien, GitHub Actions is the one we support out of the box. For some other CI system there's no separate integration needed: our CLI does the same job. You can check more on CI/CD here: https://docs.cekura.ai/documentation/guides/github-actions-ci-cd

0
回复

For teams already using CI/CD pipelines, how much setup or configuration is typically required to integrate Cekura, and is it compatible with most existing tooling?

1
回复

@cameron_collins1 Setup is minimal , once your evals are set up on Cekura, integrating into your CI/CD pipeline is just one small step, and it works with most existing tooling, not just GitHub Actions.

0
回复
#3
Lottie Creator 2.0
After Effects for the web, built on Lottie
329
一句话介绍:Lottie Creator 2.0 是一个在浏览器中直接完成矢量动画设计、AI辅助创作、交互逻辑编排与生产导出的在线工具,解决专业动效设计需安装桌面软件及团队协作低效的痛点。
Design Tools Graphics & Design Animation
在线动画制作 矢量动画 Lottie Motion Copilot AI 状态机 运动令牌 无桌面软件 浏览器工具 运动设计系统 交互设计
用户评论摘要:用户普遍盛赞Motion Copilot降低创作门槛,尤其缓解“空白画布焦虑”。多人反馈State Machines使交互设计更亲民。建议集中在:期望AI进一步细化能力,并希望看到更多社区复用的最佳实践和模板。
AI 锐评

Lottie Creator 2.0 扛着“网页版After Effects”的大旗,在2023年这个节点上,确实精准切中了一个长期被忽视的需求:高级动效制作的“桌面软件依赖”和“团队交付成本”。其真正的价值并非完全替代AE,而是通过三个“去复杂化”动作制造了差异化优势:一是用AI Copilot降低从零开始的创意门槛,解决“开场难”;二是用State Machines和Motion Tokens将交互逻辑与品牌资产组件化,解决了动效在团队内部规模化、标准化应用的难题;三是全程云端与dotLottie导出方案,直接消除了文件版本混乱与跨平台兼容的隐性成本。

然而,必须泼一盆冷水。宣称“After Effects for the web”是危险的比喻,因为AE的生态壁垒在于其庞大成熟的用户群、脚本、插件以及针对复杂影视级合成(如粒子、抠像)的底层渲染管线。Lottie Creator 2.0目前更接近“面向UI/UX动效的Figma”——擅长品牌动画、微交互、Lottie素材生产,但绝非全能。其AI Copilot目前主要作为起点助推器,而在精细关键帧控制和复杂动画逻辑编排上,评论区反馈仍需手动大量调试,这意味着“入门简单,精通仍难”。

真正要解决的问题是:它能否说服那些已经熟练使用AE的设计师切换工作流?如果不能解决“从AE迁移的全套资产损失”,它更可能扮演“低代码动画的协作层”角色——让不懂动效的PM或前端用AI快速生成原型,让初级设计师交付标准化Lottie,资深设计师则继续在AE里磨精品。团队若想在工具红海里存活,不该沉迷于对标AE,而应死磕Motion System和AI Copilot的智能化深度,让自己成为“动效驱动的交互系统”底座,而非一个更轻巧的绘图软件。

查看原始信息
Lottie Creator 2.0
Yes, we're calling it: After Effects for the web. Design vector animations from scratch in your browser, move faster with Motion Copilot AI, wire up interactivity with State Machines, and export production-ready Lottie + dotLottie. No desktop app. No dev handoff.

Hi PH!

On behalf of the Creator team, we are beyond excited to finally share our latest milestone with you: Lottie Creator 2.0!

Over the last year of building Lottie Creator, the same problem kept surfacing: on the modern web, there is a clear gap between simple tools and tools that enable you to unleash your full creative potential. That's where Lottie Creator 2.0 comes in to bridge that gap!

We’ve done our best to put YOU in the driver’s seat, all while automating the boring stuff away.

Boss on your back and you need to quickly align an entire animated brand kit? Use the Motion Systems and standardize look and feel across multiple animations. Working on a one-of-a-kind masterpiece? Use the new Graph Editor to tweak every important frame to tell your story (and lets be real: every frame is important). Need to blaze through adding interactivity for your animations? Use a combo of State Machines and Copilot to perform the repetitive clicks, while your brain focuses on your next creative idea.

We hope you give Lottie Creator 2.0 a shot, and let us know about your experience. The team is actively sourcing feedback (the good, the bad, the ugly) for us to take the product to the next level, so please don’t hesitate to reach out to us.

We’re so excited to see your animations out there in the wild!

32
回复

@karinfam Really excited to see everyone's experiments with Creator 2.0! It's a huge milestone and a major improvement. Can't wait to see what the community builds with it, using the new tools and expanded capabilities available on it.

0
回复

Motion Copilot is top of my list to try. Getting past the empty timeline is the real bottleneck. Congrats on Shipping!

5
回复

@karinfam Amazing!!

0
回复

Hi Product Hunter!  

Lottie Creator creator here (no pun intended). This is what's behind the Lottie Creator 2.0.

  • We started Creator with a challenging goal: bring advanced, production-grade animation tools to the browser while lowering the barrier to creating motion.

  • Historically, producing sophisticated animations for digital products has required complex desktop software, specialized expertise, and continuous installation, maintenance and updates.

  • With Creator 2.0, we’re building a different kind of motion tool—one that is approachable when you’re juts getting started, yet powerful enough for experienced animators to create detailed, high-quality work.

  • Creator gives you the precision you’d expect from a professional animation environment, including keyframe animation, a graph editor, and advanced easing controls to build complex motion directly in your browser.

  • We’ve also designed Creator to be AI-first. Motion Copilot helps you move from an idea to an editable animation, within seconds. AI-powered raster-to-vector conversion remove repetitive steps and give you a stronger starting point. AI doesn’t replace the creative process—it augments your creativity where your judgment and craft matter faster.



Some of the biggest additions in Creator 2.0 include:

  • Advanced keyframing and graph-editor controls for precise motion

  • Motion Copilot and AI-powered creation workflows

  • State Machines for building interactive animations

  • Motion Tokens for creating dynamic, data-driven motion

  • Motion Systems for standardizing animation across products and brands

  • Multiple animations and states packaged into a single dotLottie file

  • An extensible plugin ecosystem for expanding what Creator can do


Creator is for motion designers, product designers, developers, and teams that want to treat animation as a core part of the product experience—not as a decorative asset added at the end. Our larger vision is to make sophisticated motion easier to create, systemize, and ship at scale, without taking control away from the creator. We’d love your feedback on the advanced animation workflow, our AI features, and how Motion Tokens and Motion Systems could fit into the way you build motion today.

Thanks for checking out Lottie Creator 2.0! 

16
回复

Really proud to be part of the team behind Lottie Creator 2.0!
A lot of work went into this release, and it’s great to finally share it with everyone.

I’m especially excited about the improved Motion Copilot. You can just describe what you want, and it can help you create animations using powerful features like Motion Tokens and State Machines.
It honestly feels a bit like magic. ✨

If you try it, please share your feedback with us. good or bad. We’re very open to hearing what could be better and what you’d like to see next!

13
回复

@inerds This is damn impressive. Being able to describe an idea and watch it turn into an animation feels a bit unreal. Can't wait to try Motion Copilot!!! congrats to the whole team! ✨

1
回复

Massive congrats to the Creator team on this huge 2.0 milestone! 🎉 While I’m not on the core Creator team myself, but my team had an absolute blast collaborating with them to improving and integration dotLottie Motion Tokens and State Machines into this release.

Seeing these features come to life directly in the browser is honestly surreal. State Machines completely change the game by making interactive animations incredibly accessible, and Motion Tokens opens up so many dynamic styling possibilities for production. Really proud of the cross-team effort here. I genuinely can't wait to see all the interactive experiments the community builds with this! 🚀

11
回复

@theashraf ❤️

1
回复

Congrats to the Lottie Creator team! I'm impressed by how easy it's gotten to make animations. Something like this would have required hours of tutorials in the past, but Creator is really intuitive.

I also really appreciate that it works on the web. I've been playing around with Linux a lot lately (on top of my Mac/Windows computers), and browser-based software makes compatibility a breeze. I don't need to worry about downloading and installing anything.

And the Motion Copilot AI is probably the first time I've seen a motion design tool implement AI properly, so major kudos to the people who put that together!

10
回复

@georgeatlottiefiles Thank you, George! Don't be shy to share some of your made-on-Linux animations, we would love to see those!

1
回复

Best part of any launch: seeing what the community builds next. Show us what you make with Creator 2.0 especially motion copilot 2.0!

8
回复

Motion Copilot has changed the way I experiment with animations. Instead of starting from a blank timeline every time, I can quickly generate a direction, iterate on it, and spend more time polishing the final result. It's become a regular part of my workflow.

8
回复

@kshitij_minglani1 Glad to see you're enjoying it

3
回复

One thing I really like about Creator 2.0 is that everything happens directly in the browser. There is nothing to install, and I do not have to keep switching between different tools just to try something out. I can open it, start working, and see how an idea feels almost immediately.

That small difference makes the whole process feel much more approachable. I am more willing to experiment, make quick changes, and test different motion ideas because there is less setup getting in the way. It feels easier to stay focused on the work itself instead of thinking about the tools around it.

7
回复

Love hearing this @wenjieshen !
Creator also saves your work automatically, so you can always come back and pick up where you left off. Keep exploring, and please share any feedback or ideas with us!

0
回复

I have spent a lot of time with LottieFiles Creator over the years, so seeing it out in the world feels really special. It has been a rewarding journey helping shape a tool built for motion designers. I hope it helps you create faster, experiment more, and enjoy the process as much as i have.

6
回复

Directly gonna share this with our team of designers :)

6
回复

@busmark_w_nika Your team might especially love the Motion System inside Creator 2.0. It lets teams standardize color, easing, and motion presets in one place. Would love to hear what they think!

1
回复

@busmark_w_nika Thank you for the support, Nika! We're excited to see your team shipping those animations!

1
回复

First off, congratulations to the team on Lottie Creator 2.0! It’s been a milestone we’ve all been counting down to.

For me personally, my favorite features are Motion Tokens and State Machines. Interactive and responsive animation is just so fascinating to me, and being able to build that with the power of dotLottie still blows my mind.

As a non-designer, getting started was definitely a learning experience. But once I got the hang of a few tricks, everything started to click and the logic became more clear. I can't wait to see what everyone creates! 😊

6
回复

@lubleep Thank you for all your support in this launch Mariyam, we really appreciate it!

1
回复

@lubleep I'm with you! Apart from its web-based nature (so I don't have to jump between tools or install heavy software) and the super helpful Motion Copilot, Motion Tokens stole my attention the most. They make editing and scaling the use of animations a lot smoother and structured. Congratulations Creator 2.0 team!

1
回复

What I love about Motion Copilot is that it gets me past the hardest part, which is starting. I feed it a rough idea, it gives me something to react to, and from there I just keep refining until it feels right.

6
回复

@serhiii Thank you for this, Serhii! Let us know if you've got any more feedback, we'd love to make it easier to get started in Creator.

1
回复

I honestly didn't expect Motion Copilot to become such a big part of my workflow, but here we are. 😄

As someone who works with motion design and animation regularly, I constantly have ideas I'd love to try—but many of them never make it past the "what if" stage because prototyping takes time. Motion Copilot has completely changed that.

Instead of spending hours building a rough proof of concept, I can quickly explore different animation ideas, iterate on them, and see what works. That freedom to experiment has made me much more creative because the cost of trying something new is so much lower.

It's become one of those tools I open without even thinking about it. Whether I'm exploring interactions, testing animation concepts, or simply looking for inspiration, it helps me move from idea to prototype in minutes instead of hours.

The biggest value for me isn't just the time it saves—it's that it encourages experimentation. I've ended up creating animations I probably would have abandoned before because they seemed too time-consuming to prototype.

If you work with motion design, animation, or interactive experiences, Motion Copilot is one of those tools that quietly becomes indispensable once it's part of your workflow.

6
回复

@pecsundar honestly motion copilot has saved me hours of chasing animation directions. the new version goes wayyyy further. it builds animations from scratch, so i start with something and I can keep refining it instead of staring at an empty canvas.

4
回复

I've been using it for the past few months and I'm so proud of what the team has built here. I love the Physics Editor: makes confetti animations so fast and easy. Also Motion Presets and Easing Curves really help companies build their definitive Motion System, which enables a whole new level of consistency and speed on teams.

But without a doubt, my favorite part is still Motion Copilot. Watching an idea turn into an animation in seconds with a simple prompt never gets old. Feels like absolute magic.

5
回复

It's been great working on Creator 2.0. I'm not an animator myself, and honestly Motion Copilot is what finally let me make animations I couldn't before. A lot of this release came from feedback and comments we received, from small fixes to State Machines and the new Graph Editor. Still plenty on our list. Can't wait to see what the community builds. 🚀

5
回复

The State Machines feature is what I wanted to understand more — specifically how complex the exported dotLottie gets when you layer multiple interactive states. I use Lottie for small animated microinteractions in web projects and the dev handoff is usually where the friction shows up. Does the dotLottie bundle the state machine logic so the player handles transitions automatically, or does the developer still need to write state-triggering code around it? That answer changes whether this replaces the whole handoff or just the design side.

5
回复

Great question, @leo404 !

State machine logic is bundled inside the .lottie file, including the states, transitions, conditions, actions, interactions, and inputs.
dotLottie player reads that definition and handles the transitions at runtime.

For interactions like clicks, hovers, pointer events, or animation completion, the triggers can live inside the state machine, so developers do not need to recreate that logic in code.

If a transition depends on something outside the animation, such as app state or an API response, the developer only needs to pass that value or event to the state machine. The transition logic still stays inside the .lottie file.

Adding more states will add more animation and state data, but it remains a single compressed bundle. So for many microinteractions, this should make the handoff much simpler.

For more technical details, you can read our Interactivity & State Machines guide and the full dotLottie 2.0 specification.

0
回复

@leo404 Hey Leopold, I worked on state machines. Inad's reply nailed it on the head but feel free to ask any more questions if you have any. I'm super excited that we implemented the gesture controls across platforms so that devs don't have to themselves. It can be a complete frictionless handoff, cheers!

2
回复

@leo404 Thanks for this comment, Leopold!

You can wire up transitions directly in the Creator's state machine mode and the dotlottie player will handle it, once the right condition or trigger is met. We have a tutorial here, if you need more details!

Don't hesitate to reach out in our community if you need more help with this.

1
回复

Hey product hunters, so stoked to see what you’ll be creating with Creator! 

Been experimenting a lot with the Motion Copilot myself. As someone without motion design experience, being able to quickly animate and iterate ideas has been really fun. Can’t wait to see how it fares in the hands of everyone else, whether you’re a seasoned motion designer, a complete beginner or somewhere in between!

5
回复

@leemjenli This is amazing work. I will definitely play more!

0
回复

Congrats to everyone who shipped this 🎉

Motion Copilot is getting most of the attention today, and fairly so, but the update I'd point people to is the Graph Editor. Easing is where an animation either feels right or feels slightly off, and that's usually the part that sends you back to desktop software. Having real curve control in a browser tab is what makes Creator 2.0 feel like a tool you can finish work in, not just start it.

5
回复

From design assets to full-fledged motion creation, watching the evolution of LottieFiles' ecosystem continues to inspire. Lottie Creator 2.0 makes adding interactive motion so accessible without compromising on quality or workflow speed.


Kudos to the team on another epic launch! Wishing you all the success on PH today! 🙌✨

4
回复

@rankarpan Thank you for the kind words and all the support for this launch, Arpan!

0
回复

I'm really happy with the result I get from Creator. Not only I can create complex animations with Creator, but my workflow also is much faster and more efficient. Simply copy and paste .svg, animate, save it to my workspace and ship it!, all in one platform. I'm also thrilled to try really cool plugins inside Creator as well. Great work, team!!!!

4
回复

@ngoc_nguyen_lana Thank you for the kind words. Can't wait to see what you'll create with Creator!

0
回复

Congrats to the whole team on shipping 2.0! 🎉

I've been somewhat involved in Creator's development on and off since the earlier stages, most recently contributing to the plugin system. I've also been using it internally as a testing ground for some of the work I'm doing, and the improvements in this release are really noticeable. Smoother, faster, more responsive, and just a lot closer to what you'd expect from a polished production tool.

Proud of how far it's come. Excited for people to try it out.

4
回复

Really love the new Creator 2.0, everything feels fluid. Can’t stop playing witht the copilot to quickly generate the animation with just a prompt.

4
回复

@amirul_bin_abdullah Thank you for the kind words. Looking forward to see what is possible.

0
回复

Really proud to see what the team has been able to build and achieve so far! Been trying it out internally and am amazed to see the potential that the community will be able to achieve with this product.

4
回复

Such a powerful motion design tool, especially now with Motion System! But my personal favourite hidden feature is being able to change themes 💙

4
回复

I run an app service packed with micro-interactions, and animation is a big part of making it feel fun and alive. I’ve always wanted to use Lottie more, but as a non-designer, the complexity of existing tools made it hard to get started. LottieFiles Creator and Motion Copilot 2.0 feel like a whole new world ~_~ they make the creative process so much more accessible and inspiring ✨

4
回复

@jinui We are glad to be a small part of your process. Keep creating and come back to us if we could help.

0
回复

Not gonna lie, motion design used to feel intimidating to me, but Creator 2.0 makes it feel approachable.

Motion Copilot is my favorite, feels like having a motion designer assistant right inside the tool. Congrats team! 🎉

4
回复

@jenelim Yes, Creator 2.0; your personal motion agent! Can't wait to see what you make.

0
回复

After trying quite a few browser-based animation tools, Creator 2.0 stands out as one of the few that feels powerful without being overwhelming. Motion Copilot is a nice touch for getting ideas moving quickly, the Graph Editor gives you the control when you want to fine-tune things, and I was genuinely impressed by how smooth and responsive the editor feels, even with more complex animations. Looking forward to seeing where this goes. 🚀

4
回复

Having been on Creator since the early builds, seeing 2.0 go live is pretty surreal. I spent most of this cycle on the performance side, and that's meant obsessing over a ton of under-the-hood details. That's making sure everything from State Machines and the new Graph Editor to Motion Copilot feels insanely smooth, stable, and ready for real production work. It honestly feels like a completely different product now.

The part I'm actually waiting on is what everyone does with it, as that's always the best bit of a launch. Thanks to everyone who reported the rough edges and kept telling us what was missing. A lot of 2.0 is your feedback.

4
回复

@kadeer Thank you for this comment, Abdul, fully agreed! A lot of 2.0 and onwards will be feedback from this very same community.

1
回复

My favorite part is how easy it is to iterate. Prompt, tweak, refine, repeat. Watching an idea transform into an animation in seconds still feels kind of wild!

3
回复

Honored to be a part of Creator engineering team.
Creator has been an ongoing project that's been incredibly fun to work on, from the earlier versions up until today. It's so much more stable, intuitive and capable than ever before. I hope you enjoy using it as much as we've enjoyed building it.

3
回复
#4
Hardbook
The freelancer booking link that signs the contract for you
206
一句话介绍:Hardbook是一个将日历预约、合同签署和押金支付整合在一个链接中的工具,专门解决自由职业者“口头确认后客户反悔”的痛点和签约流程断裂问题。
Productivity Freelance Calendar
自由职业 预约链接 合同签署 押金支付 日历同步 无客户端注册 履约效率 工作流自动化 生产力工具
用户评论摘要:用户普遍认可“无需客户注册”和“签字与付款一体化”的设计。核心建议包括:支持合同条款的事先协商修改、提供可调整日期的轻量级改签流程(而非作废重签)、增加基于时间的分级取消赔偿机制。
AI 锐评

Hardbook本质上是将“轻量CRM”与“电子合同签署”做了场景化缝合,它的聪明之处不在于技术创新,而在于精准砍掉了自由职业者交易链路中两个最大的摩擦点:客户因拖延导致的“反悔窗口”,以及日结算模式下模糊的履约确认。把日历选定、合同签署和押金支付按流程硬性捆绑,形成“签字即付款”的经济锁定,实质上是将交易从“口头意向”前置到了“资金托管”状态。

但产品的护城河未必深。核心交互(日历+签名)均有成熟的API可调用,分拆复制并不困难。目前最严重的功能缺失是对“合同可变性”的处理缺位:自由职业者与客户的合作往往需要前后多轮条款博弈(如范围变更、交付物细节),但Hardbook目前的路线似乎是“签约即锁定,改签即作废”,这在简单代拍场景可行,却无法覆盖高客单价、长周期项目的动态需求。此外,单纯的美国律所默认模板和自由文本条款生成器,在面对不同法域(如GDPR或国内劳动法)时有极大概率出现法律效力瑕疵,这可能是个隐蔽但致命的雷。

这款产品的真正价值不在于“签名”,而在于用流程效率把“口头承诺”转化为“落袋资金”。它若能开放更灵活的条款协商入口、以及基于时间线的分期付款能力,才可能从“锁定轻单的工具”升级为“自由职业者的核心经营系统”。目前,它更像是一个极佳的单点破局者,而非终局产品。

查看原始信息
Hardbook
The gap between agreed dates and a signed contract is where freelance jobs die. Hardbook fixes it. Your client picks dates and signs the contract in one flow. No app, no account, no chasing. Built by a freelancer to stop deal decay and lock in work.

Hey Product Hunt 👋

Hardbook is a booking portal + contract signing in one link for day-rate freelancers.

I lost a job to a client who agreed on dates, never signed, and by the time I followed up the window had closed. The gap between "sounds good" and "officially booked" is where most freelance work dies — not on quality, not on price, just on friction.

The flow: send your Hardbook link, client picks dates from your Google Calendar, signs the contract, and — if you’ve switched deposits on — pays one, all in the same session. No app, no account on their end. Nothing moves forward without a signature, and the deposit goes to your own Stripe account, not mine.

motion designers, video editors, photographers, anyone billing by the day — to try it on a real client and tell me what's broken.

What's your biggest point of friction between "agreed on dates" and actually being booked?

Btw
- PHUNT65 + monthly → ✓ "65% off for 3 months"

- PHUNT20 + annual → ✓ valid (20% off the first year)

4
回复

@suchback In my experience hiring freelancers, the friction is never the calendar, it's locking scope and terms before anyone starts working. If the contract signs at booking, can the client edit terms or scope before it goes through, or is it a fixed template? Curious how you handle the back-and-forth that usually happens before a real engagement.

1
回复

As a freelancer, the hardest part is often not finding work it's getting the commitment locked in. Nice solution! 🚀

3
回复

@gordon_bennett not getting the commitment used to be my middle name, but now it’s harry :)

0
回复

Congrats on shipping! Every freelancer has a story about starting work on a verbal yes — putting the contract right into the booking flow is a clean fix. Wishing you a great launch day.

1
回复

@ben_kahan Thanks.

“Every freelancer has a story” is so true — I built this assuming it was mostly my problem, and so far nobody’s told me it isn’t also theirs.

0
回复

this is a real problem, not a nice-to-have. i've watched exactly this happen to freelancer friends - verbal yes, calendar hold, then silence right up until the date, and there's no clean way to force the moment where it becomes real. the deposit-tied-to-signature bit is the smart part imo, way more than the signature itself. one thing i'm curious about: once a booking is signed and the client needs to shift dates (not cancel, just move), does that need a brand new link/contract or is there a lighter reschedule path that keeps the same deposit and terms intact?

1
回复

@omri_ben_shoham1 No lighter path today. Moving dates on a signed booking means cancelling and rebooking — new contract, and the deposit doesn’t travel with it.

Though there are two shapes of this and only one is a gap. If the client knows they need you around some dates but isn’t sure yet, that’s a pencil — soft hold, no contract, no deposit, meant to move. That case is covered. And the rigidity of the other one is deliberate: a hard booking is called that because it sits in a contract and you can build your month around it. If dates slid freely it’d be a pencil with extra steps.

But moving isn’t breaking. A signed booking shifting by a week is a real thing that shouldn’t cost you the contract and the deposit. That’s the gap you found, and it’s a fair one.

1
回复

@omri_ben_shoham1 couldn't agree more! money talks. signature is secondary .

0
回复

This is actually useful. Freelancers waste so much time going back and forth before a project even starts.

0
回复

The no app, no account part for clients is what makes this actually usable, nobody creates yet another login just to book a shoot. I can see photographers and video folks adopting this fast. Are deposits or partial payments at booking on the roadmap, since chasing the deposit is usually the other half of the problem?

0
回复

@doganakbulut Already shipped, not roadmap. The client signs and goes straight to card entry in the same flow — an authorization hold, captured when the booking confirms, paid into the freelancer’s own Stripe account.

You’re right that it’s the other half. That’s exactly why it’s tied to the signature rather than sitting in a separate invoice nobody opens.

0
回复

Every freelancer I know has been burned by starting work on a "verbal yes." Folding the contract into the booking link means the paperwork happens when the client is most motivated. Smart sequencing. Are the templates jurisdiction-aware, or one standard agreement?

0
回复

@lucasjpols Neither, exactly. There’s no library of per-jurisdiction templates — I’m not going to claim coverage I haven’t had reviewed.

What’s there is a generator. You give it your governing law, your practice type, and your positions on IP ownership, liability cap, confidentiality term, termination notice, late payment, dispute resolution and payment terms, plus any custom clauses in plain English. It drafts an MSA against those, in your language. Then it’s yours to edit, and it’s version-locked at signing so you always know exactly what any given client agreed to.

Some of the presets lean US — AAA arbitration, state courts. Governing law is free text so you’re not stuck with that, but I won’t pretend the defaults are neutral.

0
回复

Congrats on the launch, Adam. Skipping a client account and doing everything through one link is exactly the friction Hardbook is trying to kill. Since that link ends up carrying all the access control, is it a long random token that can't be guessed or iterated, or something closer to a sequential booking ID? That's usually the first thing I'd poke at on a no-login flow like this.

0
回复

@vollos Signed tokens, not sequential IDs — each one works for exactly one action on exactly one booking, and it expires.

/booking/2 would be a fun afternoon for someone and a very bad one for me.

0
回复

Love the no-account approach, that alone removes so much friction. One thing I'd add: a simple payment trigger tied to the signature, like letting the client leave a deposit or set up the first milestone right in that same flow. It would close the loop from "agreed" to "funded" without needing a separate invoice step.

0
回复

@redditlurker That part already exists. Client signs, and goes straight to card entry in the same flow — no separate invoice step. It’s an authorization hold that gets drawn when the booking actually confirms, released if it doesn’t. On the paid plans, pass-through to the freelancer’s own Stripe account.

Milestones aren’t there though. Only the deposit at signing. Fair gap.

0
回复

love the framing that the job dies on friction, not price or quality. one thing I'd want as a freelancer: a signed contract still doesn't stop a client from just not showing up or cancelling last minute once the date arrives. is there any deposit or cancellation-fee mechanism built into the flow, or is that still something you'd have to chase separately even after they've signed

0
回复

@galdayan You’re right that a signature doesn’t stop anyone. Nothing does. What changes the conversation is money having already moved.

Deposits are built in, taken in the same flow as the signature — client signs, card gets authorized, captured once the booking confirms. Goes to the freelancer’s own Stripe account; I don’t hold it or take a cut of it. So if they walk two weeks later, that money is already yours rather than something you’re chasing.

What isn’t built in is a sliding cancellation scale — the 50%-inside-7-days kind of thing. The deposit is a flat amount today. And I deliberately don’t adjudicate what happens after a cancellation; who’s owed what is between the two of you, not something I should be deciding.

You listed kill fee in your other comment too. Twice in one day from the sharpest person in my thread is a fairly strong hint.

0
回复
#5
Leaping AI
AI agents that call and text in multi-day campaigns
198
一句话介绍:Leaping AI 是一款为家装、屋顶维修等实体行业设计的 AI 语音与短信代理平台,通过多日、多线程自动化外呼与接听流程,解决企业漏接电话、跟进缓慢及语言障碍等痛点,实现预约安排与客户服务的无人化运行。
Artificial Intelligence Home improvement
AI语音代理 短信自动化 多日营销活动 家居改造 销售跟进 CRM集成 合规外呼 多语言 预约调度 B2B SaaS
用户评论摘要:用户普遍关注合规与体验细节:最受好评的是其跨渠道统一的拒绝处理机制(语音拒接同步到短信黑名单)。核心争议点在于:1)多日活动中,客户情况变化(如已找别家)如何避免AI显得“刻板”;2)手动转接人类客服时的摘要准确性;3)大规模外呼面临的运营商垃圾标记(SPAM)问题。创始人回应了号码健康度管理与转接流程。
AI 锐评

Leaping AI 的亮点不在“AI打电话”本身,而在其主动将自己嵌套进实体行业复杂销售周期的勇气。从横向平台到垂直“家装”的收敛,是明智的,因为这个行业有高频、低客单价转化、强季节性且极度依赖电话沟通的特性,AI的“无休串联”天然比销售软件(HubSpot)更直接。

但你必须警惕“自动化”产生的幻觉。评论中两个尖锐问题直击其软肋:一是客户状态动态变化(如已修好房顶),AI若仍按脚本推进,会从“帮手”变成“骚扰”;二是“老线索复活”,CRM里的“无明确拒绝”与客户记忆中的“我已拒绝过”之间存在灰色地带,这比漏接电话更伤品牌信任。本质上,你在做的是将人类销售中“通过直觉与情感判断时机”的模糊艺术,硬编码成了一套严格的时序规则。

产品的护城河,并非对话流畅度(业内已趋同),而是那个“跨渠道统一DNC记录”和“运营商号码健康度管理”的底层合规基建。对于小企业主,数据隐私是麻烦,而对巨头,这是生死线。下一步真正有价值的方向,不应是更多渠道,而是引入“意图识别与生命周期预测”——当AI检测到客户当前语境(如“我问问老婆”)与历史状态(3天沉默)矛盾时,自动降频或切换策略。否则,一旦客户抱怨率飙升,运营商封号或法律风险,将比竞品丢单更致命。成也“动作”,败也“时机”。

查看原始信息
Leaping AI
We allow companies that operate in the physical world (home remodeling, roofing, trades) to automate inbound and outbound calling & texting and run multi-threaded campaigns spanning several weeks. Our AI agents can hold 100+ parallel conversations, speak in multiple languages and integrate into any CRM. We help put appointment scheduling, customer service & confirmation calls on autopilot and eliminate missed calls, slow speed to lead and insufficient lead follow ups for our customers.

Hey ProductHunt!

I am Kevin, maker at Leaping AI.

Today we are launching our AI voice and texting platform, specifically for home remodeling companies. We started originally as a horizontal voice AI calling platform and have gradually niched down to the home remodeling industry.

Why? We signed up our first home remodeling companies because they approached us inbound via our website and we discovered that they have a list of specific pain points that makes our voice AI solution almost a no-brainer.

Pain points that we address

Home remodeling companies miss calls on the weekend or after-hours because no one is in the office, struggle to reach out to web leads fast enough over the phone and therefore lose leads to competitors, cannot serve Spanish speaking customers because usually no on on the team speaks Spanish and cannot (at scale) reactivate old leads that didn't buy in the past but might now be interested again.

Our solution

Our solution is AI agents that can take inbound phone calls and make outbound dials, specifically to schedule appointments and confirm appointments. They can hold 100+ conversations in parallel, are always active and stick to the script (a big problem in the industry). We can deploy both English and Spanish speaking agents. Inbound leads are being reached out to in under 10 seconds.

If leads do not pick up, our AI texting agents will send them a text message and follow the same conversation flow.

Our differentiator

Something that differentiates us is that the AI calling and texting agents can now be weaved together in multi-threaded campaigns. It's similar to sales software, like HubSpot, where you can create multi-day sequencing with email sequences mixed together with LinkedIn outreach and cold calling. In our platform, our customers can create similar campaigns where new leads, that have opted in to be reached out to (strict requirement), will be contacted via calling and texting until they pick up and schedule an appointment. Leads can at all times opt out of these campaigns by texting STOP via SMS. We maintain our own DNC database and strictly respect that.

It can be configured how long the campaigns last, what the exit conditions are and also what the cadence should be (e.g., how many calls and texts on day 7 vs. day 14 of the campaign). Home remodeling companies already have a similar process in place with humans and now our AI agents can complement that team and provide more firepower.

We have been live with this already with several large brands in the space, e.g. Bath Experts, Aspen Contracting, etc.

Our ask

Please give us feedback on our solution and what features we should build next. We are constantly iterating and are always open for new inspiration.

4
回复

@kevin_wu25 the DNC database + strict opt-in requirement stood out most voice/text automation launches lead with the tech and bury compliance as a footnote, you led with it. we run multi-day sequencing on the email/LinkedIn side and cadence tuning is the hardest part to get right too aggressive and reply rate craters, too spaced out and leads go cold before you've made contact. curious how you're deciding cadence per campaign right now is it something home remodeling companies configure themselves based on their own experience, or are you seeing patterns across customers that inform a default?

0
回复

@kevin_wu25 i'd say you're building an interesting stuff. curious how you're handling consent and identity when the same lead gets both a call and a text from different agent threads, is that tracked at the campaign level or per agent? btw, congrats on the launch team Leaping Ai.

0
回复

The "similar to HubSpot sequencing but for calls/texts" framing makes total sense, and the channel-agnostic DNC record (spoken stop = texted STOP) is the kind of detail most teams skip until it bites them. We deal with a version of this on the WhatsApp/email side, once someone's mid-conversation with our AI and a human needs to step in, we keep the same thread going so there's no "starting over" moment.

Curious how the handoff feels on your end when a homeowner insists on a real person mid-call, does the human rep see the full call history instantly, or is there a summary step first?

1
回复

@mittalpatel There's a summary step. We condense the conversation down to what matters: enough that the human never has to re-ask something the homeowner already said, but not so much that they're hunting for the important pieces while someone's waiting on the line. Full transcript is there if they want it, but the summary is what they land on.

1
回复

Kevin, the multi-day campaign is the part I would stress test, and it is not the dialing. You covered opt in and STOP over SMS, which is more than most voice products say out loud. What I would check is the spoken revocation.

If someone tells the agent on day 7 to stop calling, that counts even though it never touched your SMS keyword path, and it has to end the remaining calls and the texts. I work in a regulated space where consent is tracked per channel, and the failure always has the same shape: the person revokes on the channel in front of them and the sequence keeps running on the other one.

Does a spoken stop write to the same DNC record as a texted STOP, and does it close the whole campaign or only the call leg?

1
回复

@clemente_lopez1 Yes: one DNC pool, channel-agnostic. A spoken stop writes to the same record as a texted STOP, and once the number is on the list it suppresses every further attempt across all channels for that customer, so it closes the whole campaign, not just the call leg. It also works from any direction: if they call into an unrelated inbound agent and ask to stop, that agent can add them and the outbound sequence ends too. That's how we make sure nothing falls through the cracks.

0
回复
This is really impressive. The multithreaded campaigns are really interesting. Curious how you handle conversations that need a human to step in mid flow? 
0
回复

the reactivation angle is the one that would worry me most honestly. an old lead who didn't buy could mean "not ready yet, try again later" or it could mean someone who told a rep months ago they weren't interested and that just never got logged properly. both look identical in a CRM as "no purchase, no explicit opt-out on file." if the campaign resurfaces someone who already said no once, even politely, that's a much worse experience than a missed call. how do you handle that gap between what's actually in the DNC list versus what a homeowner remembers telling someone on the phone six months ago?

0
回复

@kevin_wu25 Congrats on the launch! You’ve deployed voice agents into messy, real-world workflows where conversations can span multiple calls and texts over several weeks. What surprised you most about where AI works well and where customers still clearly want a human?

0
回复

congrats on the launch!

0
回复

the multi-day part is what stands out to me too, but from a different angle than the consent/DNC thread below - a lead's situation can just change between calls. someone's roofing quote request from day 1 might be moot by day 5 because they already hired a competitor or got it patched themselves. does the agent pick up on that kind of signal mid-call and update the campaign, or is there a real risk of the agent sounding tone-deaf by following up on a problem the homeowner already solved elsewhere, which seems like it'd burn trust faster than a missed call ever would

0
回复

Missed calls and slow speed-to-lead are massive pain for local service businesses, banks especially lol. Putting appointment scheduling and confirmations on autopilot with multi-day campaigns is a game-changer for trades. Congrats on the launch, goated!

0
回复
Congrats on the launch! Multi-day AI campaigns are an interesting approach. How do you measure success across longer customer journeys compared to traditional outreach?
0
回复

multi-day is the part that sounds small and is not. a campaign running for weeks means the agent's picture of a contact has to survive between sessions, and every stale field is a chance to say something that was true last tuesday.

the thing i would want measured is the gap between "we placed the call" and "a human heard it". we work on an adjacent problem, confirming a form submission actually registered on someone else's system, and the honest version needed independent confirmation rather than our own send log. voicemail, a carrier drop, and someone who hung up in two seconds all look like a completed attempt from the sender's side.

how are you counting those?

0
回复

Good luck with the launch guys!

0
回复

Leaping has come so far from its first launch! What challenges did you have to solve to maintain context and memory across these channels?

0
回复

That's interesting. When a homeowner insists on a real person, how does the handoff actually work?

0
回复

@dhiraj_patel5 The agent picks up on the request and transfers the call to an available rep, along with a short summary of what's been discussed – so the rep comes in knowing who they're talking to and what the homeowner already said, instead of starting from zero. If nobody's available, it doesn't dead-end: the request gets logged and a callback scheduled rather than leaving the person stuck with a bot.

0
回复

The horizontal to home remodeling niche-down is a smart call, the physical-world trades are so underserved by voice AI. I work on the consumer side (daily check-in calls to aging parents), and the thing that quietly kills outbound for us is carrier spam labeling: even a wanted, friendly call gets flagged "Scam Likely" and never picked up. With 100+ parallel outbound calls, how are you handling number reputation and STIR/SHAKEN attestation so calls actually connect? Would love to hear what held up at scale.

0
回复

@igorgurovich We run our numbers through a third-party reputation service that tracks number health and handles remediation, so we see a number degrading before connect rates drop and can rest it instead of burning it. You'd want to keep the number of dials per day per number low (usually below 100 dials/day is optimal). 100+ parallel calls is fine as long as no single number carries it.

0
回复
Awesome product!
0
回复
#6
EasyCircuit
Hardware prototyping, as simple as vibe-coding
176
一句话介绍:EasyCircuit是一个AI电路设计助手,用户用自然语言描述需求,即可自动生成电路原理图、匹配真实库存零件、输出面包板到洞洞板的组装方案,让零电子工程经验的人也能快速完成硬件原型制作。 ### 关键词 AI电路设计, 硬件原型, 电子设计自动化, 元器件采购, 自然语言生成, 面包板, 创客, 嵌入式开发, 硬件入门, 产品众筹工具
Design Tools Prototyping Hardware
用户评论摘要:用户称赞“自动匹配真实库存零件”和“面包板到洞洞板”流程是痛点解决方案;核心问题集中在:电气约束检查(电流/安全)是否足够、能否导出PDF/SVG原理图、是否支持向工厂BOM过渡。开发者回应已实现结构级检查(如GND引脚、I2C连接),对市电(AC)和锂电池(LiPo)有硬性安全拦截。
AI 锐评

EasyCircuit精准切中了“会写代码但不会电路”的创客群体断层,其价值不在于AI的智能程度,而在于将硬件原型从“知识密集型”降维为“资源密集型”。它把最折磨人的两个步骤——电气图设计与零件选型采购——一键打包,让用户只需关注“想要什么功能”,而非“哪个电阻不会烧”。

产品目前没有试图做“全功能EDA”的野心,而是聪明地选择了“面包板→洞洞板”这个有限但真实的场景。核心亮点有三:一是零件库与实时库存挂钩,二是结构化检查(而非LLM碰运气)拦截空焊和危险配置,三是引导用户从可实验的原型出发,而非直接跳向PCB。创始人对“安全边界”的清醒认知值得肯定:不给新手操作市电的机会,用硬停止而非软警告,这对硬件产品至关重要。

但需警惕“剩余问题空间”的挑战。当前检查仅基于拓扑和典型功耗估算,对模拟信号、高频、噪声等场景无感知。一旦用户突破“模块组合”的边界(如需要自己设计放大电路、滤波器),工具会迅速失效。此外,290个零件库目前覆盖有限,对“冷门传感器+特定电源组合”可能无法给出可靠方案——这恰恰是许多“有趣项目”的核心痛点。

真正的长期价值,在于它能否积累“设计-验证-烧录”闭环中的涌现数据:用户描述什么、AI怎么拆解、搭建后哪些失败、如何修正。这条数据飞轮才是构建下一代“硬件设计经验库”的关键。目前产品停留在“入门加速器”,离“设计助手”还有一段距离。如果只卖套装和简化流程,护城河不深;但如果能借用户数据反哺设计建议,未来可能成为面向消费者的“硬件设计知识引擎”。

查看原始信息
EasyCircuit
An AI circuit copilot that designs your project and sources the parts automatically — no electrical engineering experience needed. Describe it in plain language, get a verified parts kit, and build it staged from breadboard to soldered perfboard.
Hey Product Hunt 👋 I built EasyCircuit because of a much weirder project: an instrumented orchid growth chamber. I wanted to run real causal-inference experiments on plant growth — controlled interventions on temperature, humidity, light — which meant I needed a custom sensor/actuator circuit with an ESP32, humidity sensors, a pump, heaters, LEDs, all wired together correctly. I know how to write code. I did not know how to design a circuit. Every time I tried, I'd hit the same wall: which pins are safe to use, what resistor values won't fry a sensor, whether a part I liked online was actually in stock anywhere. Days would disappear into datasheets before I'd wired a single connection. So I built the tool I wished existed: describe what you want in plain English, and a copilot designs the full schematic, lays it out for a breadboard so you can prototype before anything is soldered, and matches every part to a real, in-stock supplier — then bundles it into one made-to-order kit if you want it shipped. It's still early — 290+ parts in the catalog today, more added weekly — and it's meant for exactly the person I was a few months ago: comfortable with code, new to electronics, tired of losing a weekend to a part that turned out to be discontinued. Would love your feedback, especially if you try describing something weird and see what it designs. That's the best way to find where it breaks.
1
回复

Sourcing real in-stock parts is the underrated feature here, half the pain of hobby electronics is finding out the part in the tutorial went obsolete years ago. The breadboard to perfboard progression is a nice touch too. How does verification work under the hood, does it actually check electrical constraints like current limits or is it matching known patterns from the parts library?

1
回复

@adamkamaneh Good question, and the honest answer is it's pattern/topology matching today, not full electrical-rules simulation. It checks things like whether there's a GND pin, whether an I2C device has both SDA and SCL, whether the fuse is before the load. Real checks, but structural. I looked hard at whether Ohm's-law current/voltage-limit math is actually the missing piece, and mostly it isn't: almost everything in the catalog is a pre-made breakout module, so the resistor sizing and regulator selection are already solved on the module itself, the same reason nobody hand-calculates resistor values for an off-the-shelf sensor breakout. The real gap was narrower: total power budget, whether your pump plus fan plus heater plus servo actually fits the supply you picked. That's live now too: it sums conservative, well-known typical current draw for actuator-class parts and checks it against the supply's stated amp rating, and only surfaces when both sides are actually known, never a fabricated number. So: real, scoped safety checks stacking up, not a general verification engine. Glad the sourcing and breadboard-to-perfboard progression are landing, that's the part I care most about getting right. If you're into what we're building, we'd love a follow at @try_easycircuit on Instagram.

0
回复

The orchid chamber origin story is great, that is exactly the kind of project people abandon once the wiring gets confusing. This would get me to finally attempt a hardware side project instead of stopping at the idea stage. With 290 parts in the library so far, how do you decide what gets added next, sensors, motors, or whatever users request most?

1
回复

@doganakbulut Glad you like it! On the parts question, it's a mix. We use AI to simulate the most common project types beginners actually build. Then in production, whenever someone's actual request needs a part that isn't stocked yet, that gets captured and queued for sourcing rather than just failing quietly. So the library grows from anticipating the popular scopes and from the real long tail people ask for, nothing gets left off, it just might take a bit to source. Follow us on Instagram at @try_easycircuit if you like what we're building.

0
回复

Do you see EasyCircuit evolving into a full KiCad workflow? It would be amazing if beginners could go from a plain-English idea all the way to a PCB while learning why design decisions were made.

1
回复

@tarqiya_forgah Yes, that's exactly the aspiration, not a hypothetical. I met a 13 year old at a hackathon who built a robotic dog from scratch: 3D printed the body himself, soldered the electronics, had a real PCB manufactured, and trained the kinetics with reinforcement learning. That's proof this is achievable, the ceiling is genuinely that high. Our mission is making that pathway available to everyone, not just the kid who happens to already have an engineer in the family or years of a head start. Going from a working prototype to real PCB fabrication isn't built yet, that part is honest, but what you pointed at, learning why design decisions were made, is already the core of how this works today, not something we'd bolt on later. The copilot already explains its reasoning at the schematic and breadboard stage, which pin is safe and why, what a fuse is protecting against. Extending that same explain as you go approach all the way to a manufacturable board is the natural continuation of the actual product philosophy, not a pivot. Appreciate you painting the picture, and if you're into what we're building, we'd love a follow at @try_easycircuit on Instagram.

0
回复

Hardware is the last place where iteration still costs weeks and real money, so lowering the barrier here matters more than another web-app builder. The question I'd have as a user: how far does it get me before I need an EE to check the work? Congrats on shipping.

1
回复

@lucasjpols Appreciate that, and it's the right question for a beginner to ask before committing time. Honest answer: further than you'd think for a real working prototype. My own instrumented orchid growth chamber, sensors, pump, heater, the whole thing, was built this way and actually works, running real experiments, not a demo. That's the bar I care about, one person going from an idea to a working unit on their desk. Where it stops today is mass production, real industrial electronics, PCB fabrication at scale, that's genuinely more work and not solved yet. We're focused on getting people started first, and figuring out how far that path can extend toward manufacturing is the next question for us, not something we're pretending is already done. Thanks for the sharp question and the congrats. Follow us on Instagram at @try_easycircuit if you like what we're building.

0
回复

the fried-sensor case is one thing, but I'd want to know about the failure modes that are actually dangerous rather than just annoying - anything touching mains voltage or LiPo battery charging. does the tool flag current/voltage limits that could cause a real safety issue (overheating, fire risk) as a hard stop, or is that still something a beginner could accidentally wire past since it technically "works" on the breadboard?

1
回复

@galdayan Fair callout, and you're right that fried-sensor vs. fire-risk are very different bars. As of today it's a hard stop, not a soft warning. If a prompt shows clear intent to wire directly into mains/wall-AC, the Copilot refuses to generate anything and tells you to use a pre-built, enclosed mains-rated module instead, checked before it even reaches the model, not relying on the LLM to "remember" to refuse. That's a deliberate positioning choice, not just a limitation. EasyCircuit is for people who don't know what they're doing yet, and mains is exactly where not-knowing gets genuinely dangerous rather than just expensive. It doesn't shut the door on mains-adjacent projects though (smart plugs, appliance control): the catalog already has 5V-coil relay modules and DC-AC solid-state relays, so EasyCircuit will design the low-voltage control side and require a sealed, mains-rated module for the actual switching, never bare wires. Same idea for LiPo/Li-ion: a bare cell in a design now requires a protection circuit (TP4056/BMS) present, flagged as a blocking check. It's scoped hazard-detection, not a general electrical-rules engine, but the two failure modes that can actually hurt someone are real, tested guardrails now, not just hoped-for LLM behavior. Thanks for pushing on this. If you spot other failure modes worth a hard stop, I'd genuinely like to hear them. And if you're into what we're building, we'd love a follow at @try_easycircuit on Instagram.

0
回复

The prototype is the easy half, honestly. I do sourcing and QC in Yiwu, and what usually stalls people is the handoff: the factory asks for a BOM with acceptable substitutes, and the answer is "whatever the prototype used." Then someone swaps a capacitor to hit a price target and nobody catches it until the first batch lands.

Curious whether you export anything a factory could actually quote from, or is that outside what you're solving?

1
回复

@supplymo Good question, and it points at a real seam. Today EasyCircuit is focused on breadboard to perfboard: prove the design on a breadboard, then a packed perfboard layout you hand-solder from a made-to-order kit. It's built for one person building one unit, not a production run. What you're describing, a factory quoting from a BOM with tracked acceptable substitutes before a batch lands, is the next stage after that, going from a working prototype to real manufacturing, and it's a natural next step for this rather than something bolted onto what exists now. The right version of it needs substitute/equivalent mappings per component, which is genuinely the next thing to build once the perfboard side is solid. Given you do exactly this handoff for a living, I'd rather shape that next step with someone who's seen the capacitor-swap failure happen than guess at it myself. Worth a direct conversation if you're open to it, feel free to DM.

0
回复

A schematic view export would be amazing, especially as a downloadable PDF or SVG. Right now you get the parts kit but if I want to tweak the design myself or share it with a maker friend, I have nothing to work from visually. That would make this way more useful for anyone learning along the way.

0
回复

Looks like I'm about to lose myself in this for quite a few fun evenings! :)

0
回复

@mike_kosenkov Glad you like it ;) Enjoy the evenings, and if you build something fun, would love to hear about it.

0
回复

the orchid growth chamber backstory sells this way better than a generic pitch would. one thing that gives me pause with AI-designed hardware specifically vs AI-designed software: a bad software deploy just rolls back, but a bad resistor value or wiring choice doesn't show up until you've already soldered a real board and maybe cooked a sensor. does the breadboard stage actually catch those mistakes before anything's committed, or is it mostly a physical-layout check rather than a real electrical sanity check?

0
回复
#7
Ycode AI Agents
Build websites with AI
164
一句话介绍:Ycode AI Agents 允许用户在可视化网站构建器中接入Claude、OpenAI等AI模型,让AI直接理解并修改现有项目,解决传统AI建站工具“一次性生成、无法持续迭代”的痛点。
Open Source Website Builder Web Design
AI网站构建 可视化编辑器 多模型接入 自带API密钥 CMS内容管理 设计迭代 开源 AI代理 网站开发工具 无供应商锁定
用户评论摘要:多数用户认可BYOK(自带密钥)模式透明且无加价,避免传统AI积分制的成本不透明。核心问题聚焦于:AI修改是否有版本回滚(回复称有检查点和撤销功能)、如何确保AI对CMS写入的真实性(回复强调由服务端数据库报告,非模型自述,且每次操作后读取权威快照)、多模型切换是否保留上下文(回复确认可跨模型无缝续聊)。
AI 锐评

Ycode AI Agents 的核心突破不在于“用AI建站”,而在于解决了AI建站行业的一个结构性矛盾——绝大多数工具(如Framer AI、Wix ADI)通过黑盒生成一次成型,用户若需迭代只能推倒重来。Ycode反其道而行:AI不再掌控全流程,而是作为现有项目的协作者,基于用户已有的设计、内容、组件进行增量修改。

这种定位的聪明之处在于:它切中了从“原型”到“产品”的真实工作流痛点。数据显示,网站创建后80%以上的时间花在维护和迭代上,而非初始搭建。Ycode的“检查点-撤销”机制、服务端双校验(数据库确认+状态快照)和开源BYOK模式,共同构建了一个对专业用户友好的信任架构——模型可以出错,但工具层不会撒谎。

然而,风险同样存在:第一,支持4个模型意味着维护成本骤升且无法统一优化prompt,不同模型在视觉编辑场景下的精度差异可能导致用户体验分裂;第二,“切换模型保留上下文”虽然技术上可行,但上下文包含大量结构化指令(DOM结构、CSS状态),模型间的注意力迁移可能引发语义漂移;第三,BYOK虽然透明,却可能劝退那些不懂API配置的轻度用户,让产品天然偏向开发者群体。Ycode需要在“灵活”与“易用”之间找到一个更精妙的平衡点,否则它最终可能只是一款面向技客的“AI编辑器”,而非大众化的AI建站平台。

查看原始信息
Ycode AI Agents
Connect Claude, OpenAI, Gemini or Grok to Ycode and use your preferred AI directly inside the website builder. Describe what you want to change, and the agent can update designs, manage CMS content, and build or improve components for you

Hey Product Hunt! 👋

We built Ycode to give designers and developers more control over how they create websites, and this launch takes that idea further.

You can now connect Claude, OpenAI, Gemini or Grok and use your preferred AI directly inside Ycode. The agent can help you refine designs, manage CMS content, and create or update components without leaving the editor.


Instead of generating a website once and starting over, the AI works with your existing project and helps you keep improving it.

We’d love to hear what you think, how you would use it, and what you would like us to build next. Thanks for checking out Ycode! 🙌

8
回复

@lunenas Congratulations on the launch!

I really like this direction. Most AI website builders help you create something from scratch, but real projects spend far more time evolving than being created. Having AI understand and work with an existing project feels much more practical.

I'm curious, after seeing people use this, what's the most common workflow? Do they rely on AI more for design iteration, CMS management, or component building? It would be interesting to see where users naturally trust AI the most. Wishing you an amazing launch!

0
回复

Love that I can use my own Claude/OpenAI key instead of buying AI credits with a markup. You pay your provider directly and Ycode charges nothing for AI usage.

4
回复

Huge congrats🙌 on the launch @Ycode team.. giving creators choice over their AI model directly inside a visual builder is a game-changer qq Is there a version history rollback for AI edits, so if a prompt yields a messy layout we can instantly restore the canvas to the previous state?

3
回复

@vikramp7470 Yes! When the AI makes a change, it creates a checkpoint of your canvas first. So if a prompt gives you a layout you do not like, you can hit Undo right in the AI chat and instantly snap back to how things were before that prompt, no manual cleanup needed. 🙌

4
回复

the bring-your-own-model part is the interesting risk here. we run a browser-driving agent and swapped the loop model across four candidates on one fixed task, and the failure that cost us most was not a model that errored. it was one that reported every step as successful while committing nothing, so the run read clean and the output was empty.

if your agents write to the CMS, that shape is worth guarding against, because a tool result saying it updated something is the model's own claim rather than the CMS's. do you read the record back after a write, or take the tool return at its word?

2
回复

@whateverneveranywhere We guard against exactly that. Tools execute server-side against the project DB, so a tool return is the database's report, not the model's claim — errors propagate into the loop instead of being narrated away. Then every turn ends with an authoritative snapshot read back from the DB (state + server-compiled CSS), and that's what the canvas renders — so a silently failed write would be visibly missing, not invisibly "successful". Final backstop: everything is draft-first with a per-turn Changes card built from actual writes, and only a human can hit Publish.

3
回复

Since you support Claude, OpenAI, Gemini and Grok, are users bringing their own keys per provider, or would you consider a single unified endpoint to abstract the model routing and billing? Curious how you're handling failover when one provider's API is degraded.

2
回复

@tian_yi1 Bring-your-own-key per provider. Each project can hold a key for Claude, OpenAI, Gemini, and Grok side by side — keys can be shared with the whole project or scoped "only me" so contributors bill their own accounts. We deliberately didn't put a unified endpoint in the middle: Ycode is open source and we don't want to be a billing middleman or add markup — you pay your provider directly and see exactly what each session cost in the usage badge.

No silent cross-provider failover (we'd rather not swap models behind your back mid-build). Instead, provider errors surface as readable messages in the chat, and if a provider is degraded you can switch models from the picker mid-conversation — the session context carries over, so the new model picks up where the previous one left off.

1
回复

Love the approach here. Bringing your own API key and paying providers directly is a much more transparent model than the usual markup. Pair that with being open source, and it gives users real flexibility without vendor lock-in. Wishing the team a fantastic launch!

2
回复

Giving users the freedom to work with Claude, OpenAI, Gemini, or Grok directly inside the builder is a strong level of flexibility. Can users easily switch models within the same project while keeping the previous context?

2
回复

@nico_mandera Yes, absolutely! You can switch models anytime, right from the chat (even mid-conversation). Pick from Claude (Opus 5, Fable 5, Sonnet 5), GPT (5.5, 5 Mini), Gemini (3.1 Pro, 3.5 Flash), or Grok (4.5, 4.3), and swap between them whenever you like. The full conversation context carries over, so you can start on one model and switch to another without losing your thread or repeating yourself, and it's all working against your current project state either way.

3
回复

@nico_mandera Yes, you can switch models right in the chat composer, mid-conversation. The session context carries over, so the new model picks up where the last one left off.

3
回复

Most AI websites tools optimize for convenience @lunenas . This seems to balance convenience with flexibility

2
回复
#8
Jotform Website Widgets
Build and embed no-code website widgets in minutes
156
一句话介绍:Jotform Website Widgets 让用户在无需编码的情况下,快速为网站添加评论、预约、聊天、弹窗等互动组件,解决网站功能碎片化和多工具管理繁琐的痛点。
Marketing Website Builder No-Code
无代码网站组件 网站互动工具 表单扩展 拖拽式构建 嵌入组件 客户互动 预约系统 聊天插件 弹窗公告 Janform生态
用户评论摘要:用户普遍认可其无代码拖拽体验和与Jotform账戶的集成价值。主要疑问集中在:与其他工具相比的独特优势(评论回答强调150+组件与统一账户);多组件同时嵌入时的性能轻量性问题(官方回复称可独立加载并持续优化);建议增加条件逻辑的可视化预览模式。
AI 锐评

Jotform这一步棋走得既聪明又保守。聪明之处在于,它精准抓住了现有数百万表单用户的“剩余需求”:既然你信任我用表单收集数据,那么网站上的预约、聊天、FAQ等互动环节,自然也该由我来承包。这本质上是在做“客户生命周期价值”的深度挖掘,将一个工具型产品向轻量级网站运营平台延伸,意图降低用户的多工具切换成本。

但冷静来看,产品本身并无颠覆性创新。所谓的“150+组件”更多是对市面上成熟模块(如Calendly的预约、Intercom的聊天)的整合与无代码化,而非底层技术突破。评论中已有用户敏锐地指出“individual widgets may not be entirely new”,恰恰点出了其软肋:缺乏差异化的核心能力。当用户需要高度定制或复杂业务逻辑时,这套工具链的灵活性可能很快见顶,而Jotform此前在表单领域的“大而全”有时反而意味着“精而不深”。

更值得关注的风险在于性能与生态依赖。虽然官方强调了独立加载,但多个Widget在同一页面叠加后的实际性能表现,才是决定用户是否长期使用的关键。此外,将网站交互层深度绑定在单一平台账户上,一旦Jotform调整定价或功能策略,用户将面临较高的迁移成本。

总而言之,Jotform Website Widgets对现有用户是高效实用的“锦上添花”,对追求极致性能或复杂交互的网站开发者而言,则更像一个“备选方案”而非“必选答案”。其真正的价值,在于验证了“表单工具向网站交互平台演进”这条路径的商业可行性,而非产品本身的技术壁垒。

查看原始信息
Jotform Website Widgets
Upgrade your website with customizable no-code widgets. Create interactive experiences with reviews, booking, chat, popups, countdowns, FAQs, media galleries, announcements, and much more, then personalize every widget and embed it on any website in minutes.
👋 Hi Product Hunt! We're excited to introduce Jotform Website Widgets. Over the years, we've helped millions of people build forms without code. One thing we kept hearing was that users wanted an equally simple way to improve the rest of their websites, not just their forms. That's what inspired Website Widgets. Whether you want to showcase reviews, add booking, chat with visitors, answer FAQs, display media, or launch a countdown, you can now create and customize website widgets in minutes and embed them on virtually any website, no coding required. We designed this launch to be as accessible as possible, so you can start building right away and explore different widgets for your website with ease. This is just the beginning, and we're already working on expanding the widget library with more ideas. We'd love to hear what you think: 👉 Which widget are you most excited to use? Or if there's a widget you've always wished existed, let us know, we'd love to build it. Thanks so much for checking us out, and we can't wait to hear your feedback!
6
回复

@aytekintank Which widget would make the biggest immediate difference on your site; and how would you use it? For example, would you add a booking widget to reduce back-and-forth messages, a reviews widget to boost trust, or a chat widget to capture visitors in real time? Tell us what problem you’d solve first and why.

0
回复

@aytekintank the FAQ/chat/booking bundle makes sense as a natural extension you already have the trust surface from forms, so widgets feel like the same toolkit for the rest of the site instead of a new product to learn. which widget are people actually reaching for first at launch booking, or chat?

0
回复

Congratulations! There are quite a few website widget tools out there. What do you think is the one feature that makes Jotform Website Widgets stand out from the rest?

1
回复

@saksham_shukla3 Thanks so much! I'd say it's the combination of variety and simplicity. With 150+ customizable website widgets, you can add everything from reviews and booking to chat, popups, FAQs, countdowns, and more all from a single Jotform account with a familiar no-code experience.

0
回复

jotform is been my all time used product and this new feature makes it even better

1
回复

@bibhash_dutta Thank you so much! That truly means a lot to us. We're thrilled to hear Jotform has been a part of your workflow, and we hope Website Widgets becomes another tool you'll enjoy using. Thanks for your continued support!

0
回复

Jot has evolved a lot! Been watching Jot for so many years now. Congrats on the launch!

0
回复

Honestly, the drag and drop builder is way more intuitive than I expected. Threw together a feedback survey in like two minutes without touching any code.

0
回复

honestly the drag-and-drop builder feels really polished, like every element just snaps into place without that finicky alignment frustration you get with other form tools.

0
回复

Used Jotform to spin up a feedback form for a small project and was honestly surprised how fast it came together. The drag and drop felt snappy on my laptop, and the conditional logic just worked without me needing to mess with anything.

0
回复

Jotform has grown far beyond a form builder, and this feels like a logical extension of the ecosystem rather than a random add-on. The individual widgets may not be entirely new, but having reviews, popups, FAQs, booking, chat, and forms under one account could remove a lot of small-tool fragmentation for website owners. I’m especially curious about performance — how lightweight are the embeds when several widgets are used on the same page? Congrats on the launch!

0
回复

@andrey_ivanchenko Thanks so much! We really appreciate that perspective, that's exactly how we think about it.

Rather than asking users to juggle multiple tools, we wanted to bring the most common website enhancements together in one place, all within the Jotform ecosystem.

As for performance, each widget is designed to be embedded independently, so you only load the widgets you actually use.

We know page performance is important, especially for websites using multiple widgets, and it's something we're actively optimizing as we continue to improve the product.

We'd love to hear which widgets you'd use together on the same page!

0
回复

One thing I'd love to see is a built-in conditional logic preview mode that lets me test branching paths visually before publishing. Would save a lot of back-and-forth with test submissions.

0
回复
#9
Firstpass
Actionable input so you nail every launch
143
一句话介绍:Firstpass是一款为产品发布者设计的预览检查工具,能在发布前模拟产品在Product Hunt、Glaze Store等多个平台页面上的真实展示效果,并标注出文案被截断、歧义等致命问题,解决因不同平台字符限制导致营销信息失效的痛点。
Design Tools UX Design
发布检查 文案预览 产品上线 字符限制 智能提醒 内容优化 创业工具 营销校验 平台适配 用户反馈
用户评论摘要:用户普遍认可“先预览后发布”的实用价值,并关注:能否支持X、iOS应用商店等更多平台;能否解释问题背后的原则以便学习;工具是否考虑SEO搜索关键词的隐形损失;以及如何评分只看截断点而不看检索价值。此外,有用户提出GDPR隐私合规问题。
AI 锐评

Firstpass切中了一个极细微却极痛的“最后一公里”问题:无数产品死在发布瞬间的信息错位。创始人用亲身经历和300多条上线的校准数据,证明了“60字横幅≠40字商店≠33字网格”这个被99%的发布者忽视的客观事实。其核心价值不在于“改写”而在于“呈现”——它像一面冷酷的镜子,让创作者直面一个陌生人在不同场景下如何真正“看见”自己的产品。这种反直觉的视角转换,往往比任何AI建议都更具杀伤力。

然而,产品的护城河令人担忧。核心功能本质上是“字符预检+规则化提醒”,技术上不存在不可逾越的门槛,很容易被其他审查工具或AI插件作为附加功能实现。评论中已有用户指出,工具目前“只关心读者何时停止阅读”,却完全忽略了“用户如何搜索、排名词如何存活”等更复杂的营销维度,这从逻辑上暴露了Firstpass的偏科——它解决的是形式问题(适应性),而非效率问题(转化率)。更关键的是,创始人的回复中承认“还没有完美答案”,这暗示着产品目前仍停留在“发现痛点”阶段,而非“彻底解决痛点”。

从评论热度看,用户最关心的并非工具本身,而是它揭示的深层问题——“为什么我总能发现更好的办法,却总在发布后才后悔”。Firstpass如果仅仅做一个提醒工具,将很快沦为“发布者安慰剂”。真正的进化方向应是:基于历史数据预测哪个字符改动能带来更多曝光或留存,甚至联动平台API进行实时优化建议。否则,它终究只是发布前的一针强心剂,而非产品增长的战略引擎。

查看原始信息
Firstpass
If you've ever shipped a launch and watched your copy get truncated, ignored, or misread — Firstpass shows you every surface a stranger meets your product on (the Product Hunt page, the leaderboard row, the Glaze Store grid) and pins handwritten crit notes to what breaks in each one, before you publish. The same tagline lives or dies differently depending on where it lands: PH gives you 60 characters, the Glaze Store caps at 40, the browse grid cuts near 33. Nobody checks — until it's live.

Hey hey fam — time for me to jump in the maker seat!

Since January I've done nearly 200 coaching calls focused on launch positioning and product feedback for Product Hunt.

I've watched more great products die on bad copy than bad code.

And now thanks to agentic coding, non-developers like myself can finally bring their ideas to life.

The problem is — all those non-devs need to connect with their audience, and AI slop isn't going to cut it.

To stand out in a leaderboard that regularly exceeds 1000 products launched every 24 hour period, you need powerful, muscular prose that grabs the audience by the lapels and demands they pay attention.

Firstpass gives you feedback before you launch⁉️

Not "make it pop" — a real adversarial content crit.

Here's what few people realize: the same tagline lives or dies differently depending on where it appears.

Product Hunt gives you 60 characters. The Glaze Store caps at 40. The browse grid where people actually scan cuts near 33. A tagline that's perfectly legal in one place is unpublishable in another — and you find out after you've shipped.

How it works: you paste your launch copy once. Firstpass renders it on every surface a stranger will actually meet it — PH page, the leaderboard row, the store grid — and pins handwritten crit notes to whatever breaks in each. It scores against a real rubric: tagline, description, gallery, maker comment, message consistency. Calibrated against 305 live listings.

🔎 Why it's different: other tools rewrite your copy. Firstpass shows you where a stranger stops reading — per surface, before you publish. It's a design crit wall, not a chatbot.

🕵🏻 The proof: I ran Firstpass on this launch. It caught that my tagline gets cut mid-word on the Glaze grid — the exact failure the app exists to find — and flagged an earlier draft of this comment for ending on two questions. It critiques itself.

It's early — pre-beta and free while under development. If you find the concept compelling, you can get it from the Glaze store and if you like it, I'd love your support in the Glaze Awards: 🫶🏻

6
回复
@chrismessina congrats on the launch! I like the idea a lot. is there any gdpr concerns?
0
回复

@chrismessina Quick question: what’s the single most common copy mistake you see that actually costs launches upvotes or attention, and can you share one tiny, practical rule we can apply right away to avoid it?

2
回复

@chrismessina Congratulations on the launch! 🎉 This is a genuinely useful idea. I like that Firstpass doesn't try to rewrite everything—instead, it shows exactly where your messaging breaks across different launch surfaces before you publish. That's practical feedback makers can actually act on. Wishing you a fantastic launch! I'm curious: after analyzing 305+ listings, what was the most common messaging mistake you found among otherwise great products?

0
回复
I found Firstpass on Glaze awards but I didn’t see how to support it!
2
回复

@ishita_jindal2 oh! If you download Glaze, search in the store for Firstpass, and open its page, you should see an upvote button, like here:

0
回复

The useful part for me is that it shows where a stranger stops reading instead of rewriting the copy for you. One question: when a crit note flags a break, does it explain the principle behind it so it carries to the next launch, or is it scoped to that one listing?

2
回复

@alieksia great question.

The intention is to identify a principle when content misses the mark (so you can address the underlying concern), but if it's just a matter of truncation, that's less a principle and more a pragmatic consideration meant to help you see how listings might harm your presentation.

0
回复
Chris, is this optimized for PH or also for X and HN launches?
1
回复

@ishita_jindal2 only Glaze and Product Hunt, but I will be adding more!

I'm a little worried that X is harder to optimize for, since it now supports such long posts and is an art unto itself!

Firstpass is better for launches where you have fixed metadata or content fields, like the Chrome Web Store or the App Store, etc.

0
回复
This is dope! I'm working on a project right now so this timing is perfect. I've been a lurker on your other work, so excited to see how well this works.
1
回复

@zackdn woo, thanks! Give it a shot and lemme know how I can improve it!

What project are you working on?

0
回复

Launching today myself, sitting at 3 points, so your timing is either perfect or cruel.

The feedback I actually needed came earlier than launch day. Our tagline went through four versions and the one we shipped came from cutting it to fit the 60 character limit, not from anyone telling me the earlier ones were weak. That forced edit did feedback's job by accident.

Does it work on a pre-launch draft, or only once the page is live?

1
回复

@supplymo this is definitely meant to be used before launch — to make sure that your copy looks good and reads well in the various contexts where it'll appear.

You could still use it on launch day to see if you can squeeze any more clarity in.

I see that you went with "12 free tools to check a China supplier before you pay". Here's the Firstpass analysis:

0
回复

Reading copy where it actually renders instead of where you wrote it is the whole trick, and it generalizes past launches. Last week I ran a copy and alignment pass on my agency site and every fix that mattered came from checking each string on the rendered page, the mega menu and the footer both read fine in the file and wrong in the browser. Out of the 305 live listings you calibrated the rubric against, which surface kills the most taglines? My money is on the 33 character browse grid, since nobody writes for a surface they cannot see in the editor.

1
回复

@abdullah_javaid3 I re-ran the rubric against the total 500+ launches...! I imagine that the most important surfaces are defined by the funnel — so wherever the most number of people are likely to see your copy.

On Product Hunt, that's the leaderboard; for Glaze, that's probably the store homepage, where you really only get 40 (!!) characters:

0
回复
Love the name ahaha!
0
回复
0
回复

Chris, the per-surface render is the right primitive, and I want to push on one axis the crit wall can't see by construction.

I do this dance on the App Store, where the same string holds two jobs at once. The title is a headline a human skims in a results row, and it's also an indexed retrieval field. So when a character limit forces a cut, the visible fix and the invisible cost land in different places: the tagline reads tighter and cleaner, and you've quietly dropped the word you were ranking on. Nothing on the rendered surface tells you that happened. You find out weeks later in the impressions, if you go looking.

Product Hunt is friendlier here because the leaderboard row is mostly a reading surface. It isn't purely one though. There's search, there are topic pages, and there's whatever the models are ingesting now when somebody asks them for "tools like X".

So: does the rubric have any notion of a term that has to survive, or is it scored entirely on where a stranger stops reading? A crit that says "this cut reads better and costs you the only word anyone searches for" is a note I've never gotten from any tool, and I've had to run that check by hand every time.

0
回复

@narek_keshishyan great call out, and one that I don't have a perfect answer to!

Another way to phrase your point is that there are two audiences of product metadata: human readers and robot indexers.

Firstpass currently (given its initial focus on Product Hunt launches) prioritizes human readers scrolling through the leaderboard, where you might be up against 20-30 competitors and need to stand out.

Longer term, or in other contexts like the App Store, or in Google results, it might be more important to over-index on SEO/AEO terms or phrases that can actually turn off or annoy humans.

Finding a balance between both is an unsettled science, and may never be.

But it's something I'll consider for Firstpass — especially since Product Hunt itself is prioritizing AEO performance more these days!

0
回复
#10
FlowTask 2.0
Company brain for AI Agents
136
一句话介绍:FlowTask 2.0 是一款企业级AI知识中枢,通过审批层连接Slack、WhatsApp、邮件等分散渠道,实时更新结构化上下文,让AI代理在不泄露隐私、不重复输入的前提下获得分钟级精准的公司记忆。
Productivity Developer Tools Artificial Intelligence
AI知识库 企业记忆 AI代理上下文 MCP协议 数据审批层 实时同步 隐私隔离 Slack集成 WhatsApp集成 工作流自动化
用户评论摘要:用户普遍认可审批层设计,但核心质疑集中在:手动审批是否会成为日常堵塞?记忆更新后,过时或矛盾事实如何自动淘汰?多个代理同时读取同一脑库时,如何防止因时间差导致的行动冲突?另外有用户反馈Google Workspace和Slack连接器故障,影响体验。
AI 锐评

FlowTask 2.0 切中了一个真实但狡猾的痛点:AI代理不是不够聪明,而是“失忆”太快。企业数据散落在Slack、邮件、WhatsApp等孤岛,调用时要么喂进一堆垃圾,要么每次都要手动灌上下文——成本高、时效差、隐私难保。FlowTask的解决方案本质上是在“数据入口”与“AI上下文”之间加了一层带闸门的实时管道:通过审批层过滤私人信息,通过MCP协议向多代理统一供给不断刷新的结构化记忆。这个思路是对的,甚至可以说比单纯做RAG或向量数据库更接近企业真实需求——后者只解决“搜索”,不解决“切片和管控”。

但评论区的追问暴露了产品尚未真正解决的三个裂缝:第一,审批层的可扩展性存疑。日常业务中Slack和WhatsApp每小时涌进数百条消息,如果审批是人工队列,那么“审批层”就会从安全盾牌退化为效率瓶颈。第二,事实时效与管理机制缺失。记忆每分钟更新,但旧事实是否自动过期?矛盾事实如何处理?当前产品只回答了“如何让信息进来”,却没有回答“如何让错误信息出去”。第三,多代理并发动作的一致性未解决。两个AI在相差一分钟内读到不同版本的事实,分别采取行动并导致业务冲突——这不是Bug,而是在线脑库架构下必然出现的时序问题,目前产品缺乏版本戳或事务一致性保障。

从社区反应看,FlowTask 2.0的真正价值不在于“连上所有数据源”,而在于它让企业首次有了一个可审计、可授权、可追踪的AI上下文治理框架。但坦白说,它现在的版本更像是“一个聪明的数据管道”,距离“公司大脑”还很远。方向正确,工程落地仍需补上一致性、时效性和自动化审批三块关键拼图。

查看原始信息
FlowTask 2.0
Company Brain for AI agents, Fragmented communication and data of a company which is left on channels Emails, Slacks, WhatsApp and LinkedIn connect it all in one place with approval layer attach so personal and works chats don't gets mixed and connect it with any ai agents simultaneously, where records are keep getting updated minutes by minutes and so AI agents are getting updated as well so use AI agents with less cost of context and don't have to repeat.

A company brain for AI agents (the question is why)

So I posted one simple question on Reddit about a frustration I have while using AI.

got 8.5k views in 5hrs with 45+ comments

Almost nobody said the problem didn't exist, instead everyone shared their own workaround.

Some use Chatgpt Projects. Some maintain PROJECT.md or CLAUDE.md files.

Some build knowledge graphs. Some create automation pipelines with n8n or Zapier.

Some use RAG connectors and custom memory systems.

The pattern was surprisingly consistent.

The real problem is Claude doesn't know what's happening minute to minute people said use an API or connect it manually. But then you're still feeding it everything including the useless stuff so we added an approval layer you decide what goes to AI combined with MCP, no need for Make, Zapier, or n8n it just works.

And what's the output

AI agents 10x better, knows what's happening minutes by minutes + no daily updates

5
回复

the everyone independently duct-taped the same fix pattern in your Reddit post says it all that's always a sign the fix belongs in the platform, not everyone's personal setup. genuinely curious: does the approval layer stay a one-time setup per source, or does it become an ongoing queue someone has to clear daily?

0
回复

@bibhash_dutta the approval layer for whatsapp and slack is smart to block personal chats but if a team sends 100s of messages a day. does someone actually have to sit there and manually approve every single one before the agent sees it?

1
回复

Very Happy to announce FlowTask on Today's Product Launch. After working for months on out SaaS we have finally launched FlowTask 2.0 which has much better Operations Management, now comes with a larger AI context for companies with large number of employees. Hope ya'll enjoy it!! Happy Coding.

2
回复

I like that you're tackling the context problem instead of trying to cram more data into the model. Congrats on the launch!

2
回复

Congrats on the launch. The approval layer is the part I like most — most tools that pull from Slack and WhatsApp just take everything, so personal and work chats get mixed. Letting you decide what actually enters the brain is a smart call.

One thing I keep thinking about: when the memory updates every few minutes from so many channels, how do you handle facts that go stale or contradict each other? Like a decision made in Slack last week that gets reversed today — does the brain catch that on its own, or does someone have to approve the change?

2
回复

The "approval layer so personal stays personal" line is the part I'd want to understand more. Once Slack, WhatsApp and Gmail are all feeding the same brain and any agent can read from it via MCP, is the approval/redaction happening per-source before it ever enters the brain, or per-agent at query time? Asking because those give very different guarantees — one keeps sensitive stuff out entirely, the other trusts every agent to respect scope once it's already in.

2
回复

This is something I say to every C-level I talk to: if a person's knowledge doesn't become the company's knowledge, the hire was meaningless. Most teams lose that expertise the moment someone's out sick or leaves. Turning scattered Slack, email, and WhatsApp threads into one living memory solves something I run into constantly. Well done.

1
回复
Congrats on the launch; a really interesting project. I’m interested in the approval layer specifically as this is something I’ve been wrestling with in my own business. Controlling what the model gets is definitely a critical factor in success on company implementations. So, I’m curious; how does that work in practice? If I’m approving what gets through, that’s a queue where I’m the gatekeeper and quite possibly the blocker. No one wants that admin role if they can help it. Is it rules based? Does it learn what you keep rejecting or is it a manual job? What happens if someone stops approving things for a fortnight? Does the agent know it’s working off a stale picture? Or does it carry on regardless, confidently working from its outdated view?
1
回复

The approval layer on top of the ingestion pipeline is what makes this safe to use with real company data — without it, connecting WhatsApp and Slack is effectively giving every AI session access to every conversation that got pulled in. The thing I want to understand is the MCP server lifecycle: when an agent session connects to the FlowTask MCP, does the knowledge snapshot stay consistent for the duration of that session, or can it update mid-session as the brain ingests new messages? That matters for multi-step workflows where context drift mid-task would break the run.

1
回复

the staleness questions above are the obvious concern, but there's a related one nobody's asked yet: if Claude and ChatGPT are both reading from the same brain at the same moment via MCP, and the brain is updating minute by minute, could two agents working in parallel end up citing two different 'current' answers to the same question just because one queried a few minutes before an update landed and the other after? not a bug exactly, just a consistency question that matters more once you've got multiple agents acting on the same context instead of one person reading it.

1
回复

@omri_ben_shoham1 good catch, and I'd push it one step further - it's not just two agents citing different facts, it's two agents acting on different facts. if Claude reads the brain before an update and goes and sends an email based on the old state while ChatGPT reads it a minute later and takes a different action based on the new state, you don't get a wrong answer sitting in a chat window, you get two real-world actions that contradict each other and nobody notices until someone has to clean it up. curious if there's any kind of version stamp on a read, so at least you could trace which snapshot an action was taken against after the fact

0
回复

"No more stale CLAUDE.md files" is the right problem to name — we run agents across eight repos and those files rot faster than anyone gets around to updating them.

The thing I'd worry about with an always-updated memory is a quieter kind of staleness: a fact that was true the day it was written and silently stopped being true when the code changed underneath it. Nothing in the original Slack thread ever says it expired. Does the approval layer carry any notion of a fact aging out, or is it mainly a gate on what gets in?

1
回复

Congrats on the launch! 👏

Curious—what's the one workflow your users keep coming back to?

1
回复

both the Google Workspace and Slack connectors are broken which means there's no way to actually proceed into the app and try it out. bit disappointing and not a great experience for new potential customers

1
回复
#11
Superunit
AI agents that verify employment by phone, email & fax
125
一句话介绍:Superunit通过AI智能体(Ava)自动拨打电话、发送邮件和传真,替代人工完成就业核实流程,解决背景筛查、房贷申请、租赁审核等场景中HR部门“接电话-等回电-转交”的繁琐痛点。
Artificial Intelligence Human Resources Banking
AI代理 就业背景核实 自动化呼叫 人力资源流程 抵押贷款审核 租赁审查 数据验证 电话机器人 传真自动化 合规核查
用户评论摘要:用户关注AI通话是否明示身份(是)及HR对AI的信任问题(提供签署授权函/转接人工)。质疑30%未完成案例的卡点(雇主失联/需人工介入),并担忧语音克隆风险(需验证回调机制)。赞赏解决HR核实“灵魂折磨”的实际价值。
AI 锐评

Superunit的定位精准而务实——它没有妄想用AI颠覆整个招聘或金融流程,而是选择了其中最枯燥、最反人性、但行业共识度最高的“螺丝钉环节”:就业核实。125票的社区热度不算亮眼,但200k次完成量和70%的次日完成率足够说明产品已有弹药。

产品最大的价值锚点在于“吃掉脏活”:HR部门手动接电话、等传真、留语音邮件,这种低效消耗在大企业合规流程中长期存在。而Ava的自动化并非简单外呼,而是嵌入了“研究-拨打-导航电话树-传真-输出审计报告”的完整闭环。对背景调查公司和抵押贷款方来说,这直接转化为人均产能提升和客户体验跃迁。

值得警惕的是,医疗医保行业的验证具有严格监管(录音/转录需授权许可),客户评论已明确指出“通话中提及患者姓名即触发法规记录”。Superunit明确暂不涉足该领域,说明其对合规复杂度有认知,但这意味着TAM上限受限——目前依赖背景调查、金融租赁等相对宽松领域,一旦行业普遍推行AI监管法案,类似产品可能首先被要求增加“透明播报+实时验证码”交互层。

另一个风险点是安全逆向:若克隆Ava的脚本对HR进行社工攻击,平台缺乏即时防伪机制(如回调验证或一键查询入口),就会从一个效率工具演变成信息泄露的敞口。创始人回复“建议HR要求提供签署授权函”这一防御姿态,反而透露出产品目前更多依赖人工警觉,而非技术层面的可信通信协议。

整体来看,Superunit是在一个“人人都知道疼,但没人愿做大”的缝隙市场里,用AI做了标准化的工程实现。其护城河不在壁垒,而在体量:做得越久,积累的电话树模板、雇主联系人库、传真线路质量越高。如果后续能将已完成验证的雇主信息脱敏后形成动态“可信雇主图谱”,产品将不仅是一个工具,更可能成为行业验证基础设施。但在此之前,合规透明度和反欺诈信任构建,是决定它能飞多高的两翼。

查看原始信息
Superunit
Every time someone gets hired, applies for a mortgage, or leases an apartment, someone has to confirm they actually worked where they said they did. That means millions of calls, emails, faxes (yes, in 2026), to HR and payroll departments. Superunit's AI agents do all of that from start to finish. They research, call, email, fax and handle documents until the verification is done.

Peter, the fax detail is the part nobody outside these workflows believes. Healthcare runs the same rails, and the queue shaped exactly like yours is insurance benefit and eligibility verification, which sits directly on whether a practice gets paid.

What changes there is the content of the call. The moment your agent says a patient name to a payer, the recording, the transcript and the auditable result are all regulated records, and whoever holds them needs a signed agreement before a compliance team approves the pilot. That requirement stops this kind of automation more often than the phone tree does.

Have you pointed Ava at payer calls yet, and could a customer have the recording and transcript dropped so only the result survives?

1
回复

@clemente_lopez1 I know, AI running faxes feels very dystopian right? But it's about 2% of companies require it that way. I've definitely heard of this more in healthcare (have you heard of Tennr?). I didn't know that about the regulatory hurdle with the patient name - I can imagine it being a compliance nightmare. I do believe with HIPAA compliance that recordings and transcripts don't need to survive for the result to, but not 100% sure. We're not planning to enter that space in the near future.

1
回复

Hey Product Hunt 👋 I'm Peter, co-founder of Superunit.

Every time someone gets hired, applies for a mortgage, or signs a lease, somebody has to confirm they actually worked where they said they did. In practice that's millions of phone calls a year to HR and payroll — leaving voicemails, sending emails, sending faxes (yes, in 2026), waiting days for a callback, getting transferred, leaving another voicemail.

Background screeners, lenders, and property managers hire whole teams to do this. It's tedious, expensive, and turnover is brutal because the work is mind-numbing.

Tbh we didn't set out to fix this. We went through YC's S24 batch working on something else entirely (in the accounting space), and at Demo Day it landed with a thud, which sucked. By the end of the year we made the call to rip it all out and start over.

The restart came from a chance conversation: a product leader at Checkr told us employment verifications were "soul-crushing." We got curious, started building in January 2025, and haven't looked back.

So we built Ava (not the most original agent name, I know…), an AI agent that runs employment verifications end to end. She does the contact research, calls HR and payroll, navigates phone trees, sends the emails and faxes, handles compliance, and publishes a clean, auditable result back to your system.

In the past year we've made 1.5M calls and emails to complete 200k verifications across 17 countries for 50+ customers, including some of the largest background screeners and lenders in the US. 70% completed, most back in under a day.

Building this also gave us a look at the verification layer almost nobody gets to see, so we're publishing some of it today in our State of Employment Verification 2026 report. One finding that surprised even us: when a verification routes through a named third-party platform, ~73% of the time it's a single company. More in the report → superunit.com/launch

Happy to answer anything, how Ava handles a live call, the compliance side, the international data, whatever you're curious about. Ask away, excited to be here. 🙏

0
回复

@petersuperunit the 73% stat is the kind of thing that only shows up when you've actually run the volume nobody guesses that from outside. 70% completion in under a day for something that used to mean voicemail-tag for a week is a real number, not a marketing one. curious about the 30% that don't complete fast is that mostly employers who genuinely can't be reached, or does Ava hit a wall where a human has to take over (weird phone tree, someone insists on a callback, etc.)?

0
回复

@petersuperunit glad to see AI being applied to this workflow!

0
回复

this idea is great! one thing I'm wondering about is how the calls work. does it clarify its an AI?

0
回复

@ethan_cheng It does!

0
回复

with voice-phishing sounding this convincing now, I'd actually worry less about HR being cautious of you and more about the reverse: what stops a bad actor from cloning Ava's script and running the same kind of call to social-engineer real employee data out of an unsuspecting HR rep, using your legitimacy as cover? is there a way for the person on the other end to verify a call is genuinely from you, like a callback number or portal link, rather than just trusting whoever calls and says they're an automated verification agent?

0
回复

Great idea and product for solving an everyday problem. As a business owner, I would get verification phone calls (and even faxes), all the time. Businesses want to respond, to help employees with their loans, rentals, etc, but the existing system was very cumbersome and annoying. Anything that streamlines and simplifies it is valuable, great use of technology.

0
回复

@howard_lind Thanks for the comment, appreciate that you've felt this before!

0
回复

Congrats on the launch, sounds like you've made some good progress. Question - what happens to a verification you can't complete? Does that happen often?

0
回复

@mike_stachowiak Great question! We have about a 70% completion rate directly from the employers we contact. If we can't get it from them most commonly we'll collect proof docs from the applicants (e.g. W2s, paystubs). Another common path is to get it from a third party that holds the payroll info (e.g. Equifax). Sometimes we'll return it to our clients for them to give it a shake. And sometimes it's just not gettable (e.g. business no longer exists).

1
回复

Congrats on your launch! Do receivers of the call like HR or employers get cautious cause its an AI Agent on the phone and don't want to confirm sensitive information?

0
回复

@jacklyn_i Yes they do! In those cases we always let them know that we have a signed release from the applicant that we're happy to send over by email. Or that they have the option to speak to a live person on our team if they'd like.

0
回复
#12
Cercle
The bat signal for your closest friends.
121
一句话介绍:Cercle 是一款让用户通过轻量级心情签到更新桌面小组件,直接向密友传递状态与需求、即时破解孤独与回避式沉默的信任型小圈子社交工具。
Health & Fitness Messaging Social Media
心情日志 密友圈 桌面小组件 情绪感知 孤独症 社交轻量连接 情感互助 签到机制 隐私边界 小圈子社交
用户评论摘要:用户关注互惠门槛导致沉寂者被动脱离;高频低落的信号噪音削弱警觉;隔离偏好(想独处)与小组件通知的矛盾;长期沉默是否应作为信号;Android平台缺失;习惯养成对持续使用的挑战。
AI 锐评

Cercle 的立意让人无法忽视——它痛击了当代社交里最虚伪的温柔:用“怕打扰”包装的冷落,用“大家都很忙”掩盖的无视。把“你还好吗?”这个高沟通成本问题压缩成一次点击,放到桌面首页,让友谊从猜测变成视觉信号,这个设计直觉是对的。

但真正致命的不是它解决了什么场景,而是这条评论里反复被撕开的裂痕:互惠门槛。产品用“你必须先签到才能看见朋友的状态”来维持活跃度,这听起来很公平,但它绑架了那批最需要被看见的人——那些已经沉默多日、点开App却没有力气完成一次签到的用户。他们依然被锁在门外,悄无声息。而另一边,高频低落者的信号会被朋友可视化为“常态”,起初的关切变成滑过图标时的麻木。你猜对了,“想独处”这个最诚实的选项,最终还是原封不动地推到了对方桌面上。

如果把Cercle看作“情感的暗号系统”,它确实降低了求助的好奇心和开场的压力。但暗号系统的悖论在于:它依赖双方同时启用信号仪才能交换意义。一旦某一边断电,整个网络就变成空心的。目前产品对沉默和噪音的单向处理,说明它还没有想清楚“谁来支付门槛”,也没有想清楚“传感器失灵时,系统要不要主动报警”。这不是一个功能缺陷,而是对孤独问题本质理解的深度问题。

Cercle是体面的、有温度的、功能完成度不低的。但它目前最像一座精致的灯塔——照亮了那些已经站在一起、只是没说话的人。而你真正要盖的,是一座在暴风雨里依然能搜索生命信号的雷达。

查看原始信息
Cercle
Cercle is conquering the loneliness epidemic by bringing back connections with the people that matter most. Check in with yourself and share your mood with your close friends or your entire Cercle. Take the guesswork out of your friendships, know how they're feeling before that phone call or text. Your friends' mood updates live on your home-screen widget, so you know who needs you before they have to ask. Built after watching a series of close friends live and struggle in silence

Sheeta, the founding story has a detail the rest of the thread hasn't picked up. Your friend went quiet for ten months because she didn't have it in her to have a conversation. That same person also doesn't have it in her to complete a daily check-in.

Which is why the reciprocal gate is the thing I'd pressure-test. You offered it to Irene as the retention answer and it clearly works for the engaged middle. But it's a lock, and the people it locks out are the tail the app was built for. Three weeks into a bad stretch someone opens Cercle, gets asked to log first, and closes it. Her friends see nothing, and she sees nothing either, on the day it matters most.

I build in a different category, in-the-moment support for parents, where the user is depleted by definition, and the rule I keep coming back to is that any gate in front of the thing that helps gets paid by whoever is worst off. Mutuality is a fair principle when everyone is fine.

Two things I'd want to know. Does lighting the bat signal bypass the gate in both directions, so a low mood logged after a long silence still reaches friends who also haven't logged lately? And do you treat prolonged silence as a signal in itself? In your own story the ten months was the message, but a mood log renders silence as no data.

2
回复

@narek_keshishyan Hi Narek! Users actually showed they were more likely to log a negative mood to their friend on Cercle because they didn't have it in them to have a conversation. The check in takes seconds while a conversation can go from minutes to hours.

Lighting the bat signal notifies all your friends to "check in and see how your friends are doing". This nudges people to see how their friends are doing without telling the explicitly that someone is struggling. If someone doesn't log for more than 3 days, the app checks in and asks them to log.

0
回复
My friend went through extreme depression and didn't speak to us for 10 months. Lack of communication made it hard to know what they need. Countless texts and calls went out, but we never got a response. Months later she finally reached out saying "I wanted to tell you guys what's going on but I didn't have it in me to have a conversation. I just wanted to light my bat signal that I am struggling and have you guys know without needing to explain until I am ready". That's when it hit me. We watch our friends on social media, send a like and pretend it's connection. We avoid conversations and say "maybe they're busy", "what if they aren't in a place to chat", "maybe now isn't a good time to reach out" but in reality we are dressing up avoidance as consideration. WHAT IF WE TOOK THIS CONFUSION AND PRESSURE OFF FRIENDSHIPS AND STILL HAD A WAY TO BUILD CONNECTIONS, AUTHENTICALLY? Cercle allows you to check in on your friends, see their active mood updates in app or through the widget on your phone, and it tells you what your friends need. How users are currently using Cercle: - Long distance couples are able to log and know how the other is doing. They know how the other is feeling so there isn't any miscommunication when they FaceTime. For example, boyfriend is stressed from job hunting and girlfriend had a great day at work. Instead of the girlfriend getting upset over the tone shift, she knows how he's feeling before they even call (and vice versa). - Siblings know how each other is doing, especially when keeping touch is hard and they live apart. They can be there for each other when they light the bat signal. - Friend groups that aren't able to keep up every day but want to know how friends are doing so they can reach out to chat or to help them through something.
1
回复

@sheetaverma I think a lot of people want to check in but don't know how. Has there been a feature that's helped start those conversations naturally?

0
回复

the reciprocal-gate thread above is the real design tension here, but there's a second one on the other side: what happens when someone logs low mood often, not once after ten months of silence, but as their baseline. does the friend group eventually stop reacting the same way to a widget that's been on "struggling" for three weeks straight? the whole point is catching the person who's usually fine and suddenly isn't, but I'd worry the signal gets noisier the more someone actually needs it long-term, not less.

0
回复

@omri_ben_shoham1 There's definitely a difference between each low mood. If the mood tells people they are struggling, the difference is "feeling chatty", "enjoying peace". "want to isolate". This allows people to know what they need and what is happening. The note also allows people to add details so people know what's going on and if it's something they can handle. Currently we are working on this: if the app sees continuous low mood logging, it will offer the user options to feel better along with mental health resources.

0
回复

People almost never say they're struggling, they just go quiet. Having your closest circle's state show up on your home screen before the phone call, without anyone having to perform okay-ness, matches exactly what I've seen after years of running group retreats. Loneliness spreads in silence long before anyone names it. Rooting for this one.

0
回复

@alex_gidirim Thank you so much Alex. Hope you and your Cercle enjoy this!

0
回复

The reciprocal check-in mechanic — you have to log before you can see — is smart because it makes the value exchange explicit and keeps it feeling mutual rather than one-sided. What I'm curious about is the Cercle size itself: is there a cap on how many people can be in one circle, or does it scale to the size of your friend group? That changes whether this works as a tight inner-circle tool or something bigger with looser ties.

0
回复

@leo404 There is no cap :) We are currently working on adding more groups. For now you can share your mood with your "favorited friends" which can be your people of choice, all friends, or just with yourself.

0
回复

Great idea! Wish it was on Android too.

0
回复

@rich_sun working on it soon :)

0
回复

Love that the mood lives on a home screen widget, so support becomes proactive instead of waiting for someone to find the words to ask for it.

0
回复

@ilko_kacharov exactly! thank you for your support!

0
回复

Would you consider making available on Android as well? :)

0
回复
0
回复

I really like the idea of focusing on close friends instead of trying to build another social network.

I'm curious: after people complete the first few check-ins, do they tend to keep using it consistently, or is building the habit the biggest challenge?

0
回复

@ir3ne People use it consistently because it allows them to see how their friends are doing. In order to see your friend's moods you must log in your own otherwise you can't view it. This brings consistent momentum!

0
回复

This is a considered take on a hard problem. A question about the self side rather than the friend side: does a daily mood check-in change how well people come to understand their own patterns over time, or is the log mainly a signal for others?

0
回复

@alieksia Both! You get a private mood view so you can see how you've felt in a month and what your most logged mood is. Your friends cannot see this. You can only see how they're doing once you log in your own mood.

1
回复

Sheeta, the screen that stopped me was the third check-in step, because Want to Isolate is the hardest state to build for and you already named it.

The preference lives in the log, but the delivery surface may not honor it. A low mood still lands on a home screen widget, and the person seeing it is the same friend from your story who sent countless texts. I build AI for clinicians and the rule there is identical: a stated preference about being contacted only counts if the system downstream enforces it, otherwise it is a note nobody acts on.

Does the widget hold or soften an update when someone picks Want to Isolate, or does the mood go out either way?

0
回复

@clemente_lopez1 great question! The mood goes out anyway so that people are aware that friend needs their space. If someone logs a low mood and says they "need a friend" then a notification will go out to their friends saying "check in on your friends" prompting people to go on the app so they can see that a friend is struggling.

0
回复
#13
MCP-Billing
OAuth 2.1 + usage-based Stripe billing for MCP servers
118
一句话介绍:MCP-Billing 是一个面向MCP服务器的自托管后端样板,通过内置OAuth 2.1认证、API密钥轮换和基于用量的Stripe计费,解决开发者快速搭建计费与认证基础设施的痛点,尤其规避了AI生成代码中因边缘逻辑错误导致计费失效的隐性风险。
API SaaS Developer Tools GitHub
MCP服务器 OAuth 2.1 API密钥管理 Stripe计费 自托管样板 用量计费 Redis限流 Next.js TypeScript 开源认证
用户评论摘要:用户高度认可开放核心计费模块为可信度加分;重点质疑:1)不同API密钥是否支持不同定价层级;2)非浏览器客户端(如桌面原生应用)的OAuth流程适配情况;3)Stripe重试时是否自动去重防止重复计费;4)对“零宕机密钥轮换”的实际效果和端到端测试环境有强烈需求。
AI 锐评

MCP-Billing的价值不在代码,而在“场景认知”与“信任定价”的精准计算。

核心逻辑很聪明:将用户从“AI代码陷阱”中唤醒。Maker用自身踩过的坑——Stripe webhook因错误处理不当导致计费事件静默丢失——直接命中AI开发者的痛点:现在的AI能快速生成“看起来对”的代码,却无法构建“在线路上对”的架构。这恰恰是样板项目最值得溢价的地方。

但这款产品并非无懈可击。首先,从评论反馈看,它面对的是“有使用经验的开发者”而非新手——你需要清楚自己的API密钥是否需要多层级定价、是否服务于桌面原生客户端(loopback/OAuth 2.0 Device Grant),这些缺失的场景直接从“通用样板”降级为“特定场景样板”,狭窄了目标市场。其次,79欧的一口价对标的是“开箱即用的信任”,但核心计费模块开源后,尾部用户完全可能只取开源部分自建计费——产品护城河取决于OAuth 2.1和密钥管理部分的独有难度是否有足够壁垒,目前看OAuth部分仍有适配鸿沟(如对RFC 8252的支持不完整)。

最大的销售杠杆在于“先验后买”——释出开源计费引擎作为信任锚点,这是良心之举,也是克制策略。但这恰恰反向要求:79欧必须买到的是一套“既扫清AI生成痕迹、又吃透Stripe计费哲学”的硬核工程决策,而非另一个Next.js文件夹。一句话评价:商业模式是聪明的,产品边界是务实但有短板的——垂直领域的精益工具,不是万能脚手架。市场反馈需要看是否有一波“吃过计费亏”的MCP开发者愿意为“少踩一个坑”买单。

查看原始信息
MCP-Billing
Self-hosted Next.js/TypeScript boilerplate: full OAuth 2.1 + PKCE, API key management with zero-downtime rotation, usage-based Stripe billing, and Redis rate limiting. 7 modules, 300+ tests. One-time €79, no revenue share, no platform lock-in. The metering core is free and open source on npm.
a few months ago i started building custom MCP servers on the side. the tool logic took me an afternoon; the auth and billing infrastructure took weeks. after spending four hours reading RFC 9728 at 3 AM, i ended up parking it with a static API key. i eventually decided to build it properly using Claude Code as a pair. and yes, AI writes code insanely fast now, that’s literally how this project started. but what an LLM draft doesn't do by default is make the right architectural decisions when money is on the line. early on, the AI generated a webhook handler that returned the exact same 400 error for an invalid signature as it did for a temporary DB outage. that meant Stripe never retried the event, and billing updates vanished in silence. finding, testing, and fixing those silent failures is the actual work. generating code is cheap now; the decisions, the tests, and the edge cases someone already paid to discover are not. you don’t have to take my word for it, though. I extracted and published the core metering logic as a standalone, free MIT package on npm (`mcp-metering`). you can install it, inspect the code, and judge the engineering quality yourself before deciding if the rest of the stack is worth €79. it’s a full proof of quality, not a stripped-down teaser. the full boilerplate includes full OAuth 2.1 + PKCE, per-user API key management with zero-downtime rotation, usage-based Stripe billing, and Redis rate-limiting. it’s organized into 7 decoupled modules with 300+ tests. one-time payment of €79, full source code, no revenue share, no platform lock-in, and a 7-day refund policy. would love to hear your feedback on the architecture or any of the documented trade-offs in the README!
0
回复

@marc_gil1 The webhook 400-error story is the part that'll stick with people an LLM producing code that looks correct but silently drops billing events on a DB outage is exactly the kind of bug that doesn't show up until it's already cost someone money. publishing the metering core as inspectable MIT before asking for €79 for the rest is a smart trust move too, most boilerplate sellers just show a demo video. how many of these edge cases (like the webhook one) did you catch through testing vs. only found after they'd already happened once?

0
回复

The decision to keep the metering core open source while charging for the rest is genuinely smart, basically lets people trust the math before they trust the product. Love that there's no revenue share either, feels refreshing.

0
回复

The decision to ship the metering core as free open source while keeping the full stack paid is a really thoughtful move, it lets developers build trust with the hardest part before committing to the boilerplate.

0
回复

The webhook story is the right thing to lead with. I shipped a Stripe integration this month where everything was green in test mode and the first real card 400'd on a currency param test mode never asked for - and separately found receipts failing silently because the send helper never throws and a try/catch was eating the reason. Money bugs don't announce themselves; you find them by walking the path with a real card and reading the logs like a skeptic.

Which is why I'd second Grace's checklist ask, and add one architectural vote: treat metering as a ledger, not a counter. Reserve before the metered call, settle exactly once keyed on the attempt, refund on failure - Gal's retry-dedupe question mostly dissolves when the settle is idempotent by construction. The counter tells you what happened; the ledger makes sure it only happened once.

Publishing mcp-metering free and inspectable is a genuinely good trust move. €79 for the edge cases someone already paid to discover is fair.

0
回复

The zero-downtime API key rotation is the detail that would save the most pain in practice — keys usually get replaced when something breaks, not during planned maintenance, and any rotation downtime cascades into support tickets. The question I would ask before building on this: does the usage-based metering support different pricing tiers per API key, or is billing flat across all keys on the account? That matters when you want to give different API keys to different customer tiers without running separate instances.

0
回复

The OAuth 2.1 + PKCE half is the part I'd have paid for. We ship an MCP server next to a local CLI binary, and the awkward case was authenticating a client that has no redirect target of its own — we ended up binding a loopback listener on 127.0.0.1 and handing back a http://localhost:<port>/ca... redirect URI. That took longer to get right than the billing did.

Does the boilerplate cover that non-browser client path (loopback or device code), or is it aimed at hosted MCP servers where a normal redirect URI already exists? Also curious what led you to one-time €79 with no revenue share when the metering core is already open source — that's a deliberate-looking choice.

0
回复

The Stripe webhook example is a strong trust signal. For MCP servers, billing is one of those areas where AI-generated code can look complete while the edge cases quietly decide whether the product is usable.

One onboarding thing I’d want as a small builder: a checklist or test mode that proves the money path end to end — auth, metering, retry behavior, failed payments, and key rotation — before I connect a real server.

0
回复

the webhook silent-failure story in your maker comment is the one that would actually keep me up at night, not the OAuth stuff everyone's asking about. on the usage-based billing side specifically: if Stripe retries a webhook (which it does on any non-2xx, including your own transient DB hiccups), does mcp-metering dedupe on the event id before incrementing usage, or is double-counting on retry something the integrator has to guard against themselves? that's the kind of bug that doesn't throw an error, it just quietly overcharges someone until they notice their invoice looks wrong

0
回复

The OAuth half of this is the part I would pay for before the billing half. I connect MCP servers to my Claude setup as a user most weeks, and the pattern from that side is blunt, a server whose auth works on the first try gets used the same day, and one that fails sits unauthenticated for weeks. I have one in that exact state right now. Does your OAuth 2.1 flow cover the interactive consent dance the desktop AI clients run, or is it aimed at headless API consumers with keys?

0
回复

@abdullah_javaid3 

it covers the real interactive consent flow, not an auto-approve shortcut. when a client kicks off the OAuth 2.1 PKCE authorization code grant, the user lands on an actual consent page in the dashboard to review requested scopes and click "Allow" or "Deny" (backed by TOCTOU protection in the server action and explicit tests for both decisions). so it's definitely built for interactive consent rather than just headless API keys.

that said, to be 100% transparent about a current gap that might affect your exact setup: right now the redirect_uri validation relies on a web-app whitelist checking for https:// or standard http://localhost:port/. if the desktop client you're using relies on RFC 8252 native app patterns like loopbacks with dynamic ports or custom schemes (like claude://), my current validation regex won't cover it — it's a narrow whitelist, easy to extend but not built in yet.

i also haven't tested it end-to-end against a live Claude Desktop client instance yet—the test suite exercises the server actions and handlers directly with mock clients.

curious though, what's the auth setup or custom scheme on the server sitting unauthenticated in your setup right now? would love to know if it's hitting that exact loopback/scheme restriction so i can prioritize it.

2
回复
#14
Conduit
AI agents purpose-built for Hospitality
116
一句话介绍:Conduit是针对酒店行业的AI智能体,能自动处理客人从咨询到操作的全流程(如入住申请、保洁调度等),解决传统AI助手只回复不办事的痛点,将沟通与后台系统执行无缝衔接。
Customer Communication Travel Artificial Intelligence
酒店AI智能体 自动化客服 PMS集成 酒店运营效率 语音代理 多语言支持 AI+CRM 后台自动化 OTA对接 应急转人工
用户评论摘要:用户盛赞其解决“半夜无人接电话”的痛点,并关注语音代理在紧急情况(如客人锁门、医疗问题)下的转人工机制。创始人回复已设硬编码规则,可定义“必须转人工”的话题(如火灾、客人愤怒),且支持带背景信息的温暖转接。
AI 锐评

Conduit的真正价值不在于“写回信”,而在于“关工单”。它切中了酒店业一个被忽略的深层痛点:AI客服助手往往止步于“我来查一下”,然后让运营人员继续开四个页面、手动操作PMS、对接保洁系统——这根本没有降低人的劳动强度,只是把打字环节外包给了AI。Conduit的聪明之处在于,它把AI从“对话层”下沉到“执行层”,让智能体能直接调用酒店已有的技术栈(门锁、支付、排班、PMS),形成从请求到完成的有效闭环。

116票的产品猎手投票和Marriott、Hilton等品牌背书说明行业认可度不低,但真正的考验在于“长尾异常”:当客人因系统权限不足、日历冲突、支付失败等无法自动完成的操作出现时,Conduit的纠错机制和人工交接效率如何?目前只看到对紧急事件的硬编码转移,日常逻辑偏差的处理尚未被审视。另外,与300多个品牌、拥有50M+对话和30亿美金预订值的战绩固然亮眼,但酒店系统的碎片化程度极高,PMS供应商千差万别,能否实现“开箱即用”的深度整合,将决定Conduit是成为酒店业的Zendesk,还是又一个需要大量定制化SOW的项目。技术上SOC 2 Type II合规是基本门槛,但真正能把这款产品推上主流的是它对酒店运营流程的颗粒度理解,而非Agent架构的炫技。

查看原始信息
Conduit
Conduit automates guest conversations and the work behind it. Conduit agents reply across email, WhatsApp, OTAs and social, voice calls. Agents dispatch subagents to act inside your PMS and tech stack: dispatching cleaners, updating calendars, closing the loop. Agents learn from every interaction, and you can inspect every step. 140+ languages, SOC 2 Type II. 50M+ conversations and $3B+ in reservation value for 300+ brands including Marriott, Hilton, Nobu and Fairmont.
Hey Product Hunt, Punn here, co-founder and CTO at Conduit. Most "AI for hospitality" tools stop at the reply. They draft a nice message to the guest and hand the actual work back to a human. The guest asks for a 2pm check-in, the agent says "let me check", and then someone on the ops team opens four tabs. We spent the last year on the other half. Conduit's agents read your PMS, your ticketing, your door locks, your payment system, and take action in them. Early check-in request: the agent checks availability, dispatches housekeeping, charges the fee, confirms with the guest, logs it. No tabs. Three surfaces: - Chat agents on email, WhatsApp, OTAs, socials - Voice agents for the calls nobody picks up at 11pm - Internal agents your team talks to like a coworker Our first launch was an inbox with agents bolted on. We built it in that order and it was backwards. This one is the fix. Would love your feedback, especially from anyone running ops at scale. I'm in the comments all day.
0
回复

This is incredible. Congrats @punn_kam and the Conduit team!

0
回复

@punn_kam I like that the goal isn't just faster replies but actually resolving requests. Have you seen that translate into better guest satisfaction scores?

0
回复

the "calls nobody picks up at 11pm" line is the real value prop here, but that's also exactly when the riskier calls happen, a guest locked out with a medical issue, a smoke smell, someone who sounds actually panicked versus just annoyed about checkout time. does the voice agent have a hard trigger for escalating straight to a human on-call instead of trying to resolve it itself, or is that judgment call left to the model reading tone in the moment? for something touching door locks and payment systems already, I'd want the emergency path to be the most boring, hardcoded part of the whole system.

0
回复

@omri_ben_shoham1 Yes, we have configurable escalation rules that allow customers to define topics that should ALWAYS be transferred to a human instead of being handled by AI. Common examples include lockouts, smoke or fire-related issues, and situations where a guest expresses frustration.

When one of these conditions is met, the AI immediately escalates the call. It also performs a warm transfer. During the transfer, it briefs the receiver on who the guest is and any relevant information the guest has already shared.

This creates a much better guest experience, especially in stressful situations because guests don't have to explain the issue again!

0
回复
#15
Liminal
A workspace & 2nd brain for you, your agent, and your team
115
一句话介绍:Liminal是一个让人类、AI Agent和团队在同一工作区通过本地Markdown/HTML文件协作的“第二大脑”,解决了Agent生成内容后无法实时、美观共享与协同编辑的痛点。
Productivity Artificial Intelligence
AI协作工作区 第二大脑 团队知识库 实时同步 Markdown编辑器 本地优先 Agent原生 文件协同 云端同步 开源替代
用户评论摘要:用户普遍共鸣:本地Markdown+Agent协作是真实痛点。核心建议:1)需解决笔记过时导致Agent误引的“陈旧性”问题;2)随文件增多需加入搜索或索引机制;3)大赞“按写作者追加日志”避免冲突,以及合并优于“后写覆盖”的语义。离线可用与细粒度权限获肯定。
AI 锐评

Liminal的走红并非因为其技术有多颠覆——Syncthing、Git + Markdown渲染器组合也能拼凑出类似效果。它真正的价值在于精准命中了AI编程时代一个被忽视的“基础设施真空”:当Claude Code、Codex等Agent成为新的生产力中心,传统协作工具(Notion、Google Docs)和文件系统(VS Code预览)都在这个新工作流中暴露出巨大的摩擦。

产品最聪明的地方在于“不做中间商”。它放弃了通过MCP(Model Context Protocol)去“适配”AI,转而拥抱Agent最擅长的本地文件读写。这让Agent的工作成本降至最低(无Token浪费),同时保持了最终用户(人类)的编辑体验(WYSIWYG UI)。这种“文件为王”的极简哲学直接切中了AI重度用户的要害。

然而,Liminal面临的挑战比它想象的更大。用户的“陈旧性”和“规模索引”两大拷问直指产品天花板:一个只做“展示与同步”而缺乏知识管理(可重验证、自动标记、语义搜索)的“第二大脑”,本质上是一个进化版的文件浏览器。当文件从几百个变成几千个,当多个Agent写入冲突,仅靠“UI合并”和“LLM搜索”将很快捉襟见肘。

它更像是一个优秀的“MVP”——解决了一个真实但狭小的问题:Agent工作流的渲染与共享。要想成为真正的“第二大脑”,它必须从“文件浏览器”进化为“知识仲裁者”,学会自动标注数据可信度、追踪知识变更、并主动处理冲突。否则,它不过是又一个精致的“临时方案”,等Notion或Cursor们反应过来,原生集成Agent文件支持后,其护城河将瞬间消失。

查看原始信息
Liminal
Your agent writes files locally in markdown and HTML, the formats it already speaks. Liminal renders them as clean UI instantly and syncs to your team in real time. Share a link and everyone sees the same live workspace, editable in the browser. No sending markdown over Slack, no hosting your own HTML, no waiting on Google Drive. No MCP calls, no proprietary formats, no wasted tokens. Over time, it becomes a shared second brain for your team and every agent working with you.

Hey Product Hunt 👋

I'm Justin, the maker of Liminal, and I'm stoked to finally share what I've been building!

Heads up: this is the first cut. I'm launching now to find out if the friction I've been feeling is real for other people too. Please let me know if you run into any bugs or problems, I'll get them fix asap!

Quick backstory

Last summer, the way I work changed. Agents got good enough to become the place I did the work. Claude Code became my main interface, the intelligence layer everything flowed through. My job shifted from writing and producing to directing an agent to do it.

Notion and Google Docs weren't built for this. They were bolting agents into their apps, but that's not where my agent lived, and their MCPs were slow and expensive to talk to. Claude Code worked much better and faster with local markdown files, a format the agent already speaks. So I stopped writing in Notion and Google Docs and went all in on local markdown.

I was doing all of this in VS Code because that's where I coded. But viewing and editing markdown there is clunky. You use the preview to render it nicely, then jump to the raw side to actually edit. So I built a WYSIWYG editor on top of the local files.

Then I tried to share something. Nobody wants to be sent a raw markdown file, and Google Drive doesn't render it nicely. Passing files back and forth killed any hope of collaboration. So I built real-time live collaboration on top of local files: a lightweight CLI watches your workspace folder and syncs bidirectionally with the cloud. Share a link and your teammate opens the same live workspace in a browser. Their edits land back on your disk instantly.

Once that worked, the thing I hadn't designed for showed up. My agent and my teammates' agents were all reading and writing into the same workspace. Liminal had quietly become a shared second brain. Company knowledge that keeps updating and compounding as the team works, sitting on every agent's local disk as plain files they can directly access.

What Liminal is

A collaborative workspace for humans, agents, and teams. Your AI-native second brain. Bring your own agents. Own your files.

How it works:

🟢 Work with your agent, not just through it. Your agent writes files to disk in markdown or HTML. Liminal renders them as clean UI the moment they save. Review and edit naturally in the browser, without wrestling with raw markdown.

🟢 Share instantly, without the friction. A lightweight CLI syncs your workspace to the cloud in real time. Send a link and your teammate is in the same live workspace. No sending markdown over Slack, no hosting your own HTML, no waiting on Google Drive.

🟢 A shared second brain for every agent on the team. Files live locally, so every agent has direct access to everything the team has written, decided, and learned. No external calls, no lag, no losing context between sessions.

🟢 No middleman, no MCP tax. Direct file access means no MCP round-trips, no proprietary formats, no wasted tokens. Context is cheap, fast, and always fresh.

The ask

Would love your feedback, especially if you're living in Claude Code, Codex, or another agent all day: does this map to your experience, or am I solving a problem only I have?

Justin

4
回复

@justinonthelam Abdullah's staleness question above is the real one a note pointing at something that no longer exists is worse than no note at all, since it sounds confident either way. curious if this is on your roadmap or an accepted tradeoff for now.

0
回复

the "everyone with the link sees the same live workspace" part is the piece I'd want to understand before using this for real team work - once an agent is writing files into a shared space, it's going to write things that were meant for one person's eyes at some point (a draft perf note, a client's rough numbers, whatever). is there any per-file or per-folder permission layer, or is the model still "anyone with the link sees everything," just like the underlying markdown files would if you handed someone the folder

1
回复

@galdayan 100%! Liminal is built with per file access controls, just like Notion or Google Docs. Users can have their own personal workspaces, or limit access to files in a shared workspace!

0
回复

Everyone's building second brains for individuals. Almost nobody's built one for a founder plus their agents plus their people sharing the same context. That's the harder and more interesting problem. Good luck with the launch.

1
回复

@alex_gidirim Thank you! The funny thing is that I never meant to build a second brain. I stumbled into it by focusing on building a great workspace for humans and their agents -- it just turned out that the result is a shared collection of files/docs/knowledge that's easy to access and stays up to date because that's where all of the work is

0
回复

You're not solving a problem only you have. I've run exactly this architecture for months — a folder of markdown that multiple agents on multiple machines read and write, synced through git because nothing better existed. It quietly became the most load-bearing thing in my stack, so the friction is real and so is the payoff.

Some conventions from living it, offered as free product opinions since you said Liminal doesn't have any yet:

Per-writer append-only files — each machine and agent gets its own daily log — make conflicts structurally rare instead of resolved-after-the-fact. Shared files become the exception, and the exceptional ones are the ones worth merge UI.

One small always-loaded index with a hard rule that detail lives elsewhere. At 500 files the index is the search: the agent greps the index, not the corpus. Omri's question is the one that decides whether this scales.

An authoritative current-state file that wins when it disagrees with older notes. Abdullah's staleness case bites hardest — my standing rule is that notes record what was true when written, and anything actionable gets re-verified before an agent acts on it. A tool that aged or flagged entries when the thing they point to disappears would be the first genuinely new feature in this category.

And merging over last-write-wins is the right call — this week two of my agents wrote the same knowledge as different commits and taught me that multi-writer brains want merge semantics everywhere, not just the editor. Good launch. This one's aimed at a real hole.

1
回复

@ryan_davis23 that's great feedback, thank you so much!

0
回复

this is one of those launches where the backstory tracks exactly with my own workflow - i ended up doing the same VS-Code-preview-then-raw-edit dance before giving up and just living in Claude Code with a folder of notes. the part i'm most curious about is scale rather than sync: once a team's second brain grows to a few hundred files across months of work, how do you actually find the right one? plain markdown in folders works great at 20 files and turns into its own kind of clutter at 500 - is there real search/structure on top, or is it still mostly relying on the agent to know where to look?

1
回复

@omri_ben_shoham1 It turns out that LLMs are amazing at search already, and giving agents the ability to just grep your entire workspace is pretty effective even when you have hundreds or thousands of files! So today Liminal has no separate search capabilities and I believe that AI model companies will build better search faster than I can

0
回复

Love the positioning.

What type of teams are seeing the most value so far?

1
回复

@getfoundertools This should most benefit people, teams, and companies who are all in on using a single agent as their interface for work and want to collaborate with each other, but TBD since this is day 1!

0
回复

@justinonthelam The concurrent write answer above covers editing at the same time, but I'm curious about the offline case: if you're working locally with the sync CLI not running, say on a flight, does everything just work as plain markdown with nothing lost?

1
回复
@clement_avq yup! All the files sit locally on your machine so you have always-on offline access. Once you sync back up, there’s merging and conflict resolution baked in to handle any changes that conflict
1
回复

'work with your agent, not just through it' is the real reframe. when two agents write the same file at once, does it merge or last-write-win?

1
回复
@andrewzakonov there’s merging and version control, though right now conflict resolutions is done in the UI. Next step is to make it so agents can handle that for you!
0
回复

You asked whether this maps to people living in Claude Code all day, and from my desk the answer is yes, the friction is real. My whole setup runs on a folder of markdown memory files with an index loaded each session, and it quietly became the most durable part of my workflow because notes survive every context reset. The failure mode I keep hitting is staleness, a note written three weeks ago names a file or a flag that no longer exists and the agent recommends it confidently anyway. Does Liminal do anything to age or flag entries when the thing a note points to changes, or is pruning still on the human?

1
回复

@abdullah_javaid3 Right now Liminal has no opinions about your data and files, it's just a workspace that makes collaboration between humans, agents, and teams easy and accessible, but I can see a future where it does more than that!

0
回复

finally something that doesn't make me wrangle markdown back and forth. plugged in an agent output and it just showed up as a clean page my coworker could edit, which honestly felt kind of magical

0
回复

@sedanurz5c4 I'm glad you found the experience useful!

0
回复
#16
Phantom
Voice-first AI agent that operates your Mac
112
一句话介绍:Phantom 是一款常驻 Mac 刘海区域的语音优先AI代理,无需切换应用或打开新窗口,用户可直接通过语音或文字在任意界面操控电脑完成任务,解决了传统AI助手需中断工作流、操作繁琐的痛点。
Productivity Education Artificial Intelligence
语音AI代理 MacOS 刘海交互 AI工作流 桌面自动化 效率工具 无干扰操作 屏幕上下文感知 AI控制电脑 未来工作方式
用户评论摘要:用户肯定刘海设计带来的低打扰感与屏幕上下文感知的实用性。核心质疑集中在操作的**可逆性**:语音误识别导致误删文件、发送消息等不可逆操作的后果不明。担忧持续监听隐私及会议误触发。建议明确撤销机制与唤醒方式。
AI 锐评

Phantom 的“刘海交互”在形式上确实是一次漂亮的减法——它终结了“打开新标签页-粘贴提示-等待返回-复制粘贴-切回工作”的碎片化流程,将AI从次级窗口提升为操作系统的原生交互层。这种设计哲学值得肯定:AI不应是工作流的打断者,而应是内嵌于动作的延伸。

然而,产品目前暴露出的最大缺陷并非功能不足,而是**信任模型的重建失败**。评论中反复出现的“可逆性”质疑,直指AI代理在PC领域落地的核心矛盾:传统鼠标键盘的每一次点击都是物理确认,而语音天然带有歧义性与不可逆性。当一个错误命令能直接导致文件误删或消息误发时,Phantom并未给出任何超越“Ctrl+Z”的叙事。对于“不可逆操作”,其默认的假设是“用户能即时发现并撤销”,这在快速对话场景中几乎不可能。

更深层的问题在于,它混淆了“效率”与“控制权”。用户在评论中担心的不是AI不够聪明,而是它太“主动”。Voice-first 的真正价值不在于代替用户点击,而在于将用户的意图精准地转化为系统动作;但目前的执行层面,产品更像是在用新交互包装旧命令,缺乏针对语音误识别、上下文冲突(如会议误触发)的容错与校验机制。

Phantom 的愿景是“AI不只是回答问题,而是采取行动”,这恰恰暴露了它当前的短板:它试图直接越过“行动意图确认”的保险栓,让AI从“参谋”变“执行者”,却没有给出用户敢于放下鼠标的心理安全垫。一句话:交互革命喊得漂亮,但安全底线需要重写。在未实现“可撤销的确定性”之前,它仍是一款危险的效率玩具。

查看原始信息
Phantom
Phantom is a voice-first AI agent for macOS that is always accessible via the Mac notch. Instead of opening a chatbot window and moving work into a separate conversation, users speak or type from anywhere on their Mac, without interrupting their work. Phantom understands the relevant screen and app context and completes tasks directly within the user's workflow. The possibilities are only limited by the imagination of the user.
Introducing a new generation of AI interfaces. For the past few years, we’ve interacted with AI in essentially the same way: Open a new tab. Type a prompt. Wait for an answer. Copy it. Switch back to what we were doing. Powerful intelligence, but trapped inside an outdated interface. I built Phantom to change that. Phantom is an AI assistant that lives entirely in your Mac’s notch. It’s always there when you need it and out of the way when you don’t. Instead of explaining every step, you simply tell Phantom what you want to accomplish. It can navigate your Mac, work across applications, click, type, organize, and carry out tasks just as you would; and even more. I believe the future of AI isn’t another chatbot you have to visit. It’s an intelligent interface built directly into the place where your work already happens. AI shouldn’t just answer questions. It should take action. Meet Phantom. 👻 www.heyphantom.app #AI #ArtificialIntelligence #MacOS #Productivity #FutureOfWork
0
回复

@benaja_heger voice-first for actually operating the Mac (clicking, organizing, navigating apps) is a different trust bar than voice-first for dictation one misheard command and it's clicking the wrong thing or closing unsaved work. what's the confirm/undo story for anything that can't be easily reversed?

0
回复

the reversibility question above is the big one, but there's a smaller version that would hit way more often: if it's voice-first and always accessible, is it listening for a wake word continuously, or do you have to trigger it? because the failure mode I'd actually run into daily isn't a misheard delete, it's being on a screen share or call and someone else's voice on the meeting triggering an action mid-conversation. does it know the difference between you talking to it and you talking to a person while it happens to be open?

0
回复

Finally tried Phantom and the notch placement feels genuinely clever, I can just dictate a quick question without leaving whatever I was working on. Screen context awareness works better than I expected for grabbing text from a doc.

0
回复

I like the idea of making AI available everywhere instead of living in yet another chat window.

I'm curious: after using Phantom internally, what are the tasks where people naturally switch to voice instead of typing? I'd love to know what became the biggest "aha!" moment for your users.

0
回复

the part I'm stuck on isn't privacy, it's reversibility. "click, type, organize, carry out tasks just as you would" means it can also delete the wrong file, send a message before you meant to, or submit a form, and voice commands are inherently less precise than a mouse click you can see land. what's the undo story for an action that's already irreversible by the time you notice it misheard you? until that's answered clearly I'd be nervous giving something this much control over my actual Mac, not just a sandboxed chat window

0
回复

How about privacy? How and where do you process it? Do you then store it anywhere?

0
回复

@kamil_infeld Stupid question

0
回复
#17
qsa.sh
External security scan of your own IP, in your terminal
109
一句话介绍:qsa.sh 是一款通过一行 `curl` 命令,在终端内对服务器公网 IP 进行外部安全扫描的工具,帮助用户快速了解自身主机暴露在互联网上的开放端口、服务版本及已知漏洞,无需注册账户,无数据存储,解决服务器管理者对自身安全配置“看不见摸不着”的焦虑。
Developer Tools Business Intelligence Security
外部安全扫描 终端工具 Curl一键扫描 端口映射 CVE漏洞检测 Nmap扫描 Nuclei扫描 服务器安全 免费扫描 微SaaS
用户评论摘要:用户赞赏其零注册、零存储、仅通过Curl返回文本的简洁设计,以及15秒中止窗口的隐私机制。主要担忧集中在闭源代码的信任问题(Curl请求本身是否安全),以及是否支持Serverless/Edge环境。作者回应称不执行管道下载,扫描软件开源,并将发布自动报警教程。
AI 锐评

qsa.sh 的“一行Curl”设计在理念上是一个漂亮的减法——它把传统安全扫描的注册、配置、数据留存等冗余环节全部砍掉,直接面向“我想知道互联网怎么看我”这一原始需求。这种极简主义带来的体验是其他安全SaaS难以比拟的:秒级响应、无信任负担、终端原生。然而,其真正的价值远不止于便利性,而在于它巧妙地将“安全扫描”从专业运维的工具,降维成一个类似于 `ping` 或 `curl ifconfig.me` 的日常诊断命令,这有可能触及一个高度碎片化的长尾市场——那些拥有VPS的独立开发者、小团队、甚至运维新手。但产品存在显著的硬伤:闭源包装、依赖外部IP情报库(且拒绝代理和云平台)、深扫依赖付费。核心扫描软件虽开源,但“闭源包装”在安全行业天然与信任感对立,尤其当执行命令本身就是攻击面时,用户的质疑绝非无病呻吟。Pro版(65535端口异步扫描)和Deep版(全Nuclei邮件报告)能否撑起付费逻辑,完全取决于用户对“自己手动搭一套同样工具的时间成本”的评估。如果产品能开放扫描脚本或提供自动化告警管道,其“微SaaS”模式才可能跑通。目前看,它是一个优秀的体验原型,但距离商业闭环还差一个稳固的信任基建和明确的差异化商业功能。

查看原始信息
qsa.sh
Run curl qsa.sh for a one-command external security scan of your server's own public IP — naabu, nmap + vulners, and nuclei map your open ports, service versions, and known CVEs, streamed straight to your terminal in ~30 seconds. See exactly what the internet sees of your host: no account, nothing stored. Free scans run live; paid Pro (all 65,535 ports, async) and one-time Deep (the full nuclei firehose, emailed) dig deeper uncovering vulnerabilities below the surface.

Seeing my own setup the way a stranger on the outside would is oddly reassuring, and I like that it just tells me plainly what looks open without any fuss.

1
回复

@amine_aziz_alaoui It is a very useful tool. I will be publishing a blog post soon with a how-to on setting up a crontab diff checker to auto-alert users when new vulns pop up. The entire micro-SaaS idea behind it is that everything is self-service. If you want auto-alerts with a diff checker, you create it, but our website will give you guidance. This keeps us from storing any data on users' server scans.

1
回复

Nice that curl here just streams text back, no piping into bash like typical installers. The 15-second abort window before scanning starts is a smart consent mechanism too.

1
回复

@aidan_codefox After doing the research that was the first thing I noticed and wanted to eliminate compared to many others out there, no piping/downloading (no trust) and wanted some type of terms acceptance, so I figured treat it like a checkout. It worked out nicely. Thanks for the feedback.

0
回复

Really like how frictionless this is. curl qsa.sh scanning the IP you're calling from means it's guaranteed your own host with no account or consent faff, which is a genuinely clever bit of design, and not retaining results is the right call for trust.

One question from my corner: for those of us mostly on serverless/edge now (Cloudflare Workers and the like) where there's no box with an open IP to point at, is there a version of this that makes sense, or is it firmly aimed at VPS/server setups? Either way, nice work.

1
回复

@dalemooney This version uses IP intelligence from WorldIP.io to classify the requesting IP, and it actively refuses detected proxies, VPNs, mobile devices, and Cloud platforms. However, because you're executing the command from your server and not the proxy egress, the scanner should pick up the actual IPv4 for proper scanning. All of my testing was completed on my web servers running Cloudflared.

2
回复

@dalemooney I'd like to ask, does it support both Linux and Windows?

0
回复

honestly the one-curl approach is kind of genius — no signup wall, just instant results streamed to your terminal. love that the free scan stays honest about what the internet actually sees instead of burying the real findings behind a paywall.

0
回复

This i s great! Would love to try it! ...but .....Is it opensource? I couldnt find a github link. Runing curl is in it self a security risk. Basically you allow anything to run on your computer. No gatekeeper nothing. So having it be closed source is a very red flag 😅. Its fine to do from big vendors like claud eor openai. But small vendors, opensource + curl is a must. IMO. Thanks 🙏

0
回复

@conduit_design I completely understand the concern! To clarify, you never pipe anything into your terminal here (e.g., curl qsa.sh | bash). We don't require that like other sites do, so nothing is downloading or executing code on your server. curl qsa.sh simply sends a standard web request. QSA reads the headers, and if the request comes from curl, it provides the results directly to your terminal (just like a website displays a page, but formatted for the CLI). It is a 100% external scan, not internal. To give you another example on this, another site of mine is curlhub.sh which "curl curlhub.sh" will provide you with an entire set of useful tools.

Is it open source? All of the scanning software we use (naabu, nmap + vulners, and nuclei) is open source. We've just put them together in an easy-to-use tool, though the website's wrapper code itself is closed source. The micro-SaaS idea is purely to save time for people who don't want to set up and maintain an entire security stack. One quick command gives you basic data, while the premium packages offer deeper, sub-surface scans.

0
回复
#18
G.I.A.ac
Build real, working apps from a single sentence
106
一句话介绍:G.I.A.ac 让用户只需输入一句话描述,就能实时生成可直接部署到自有 GitHub 和 Vercel 账户的完整应用,并内置业务专属 AI 助手,解决传统 AI 应用生成器代码不开放、易被厂商锁定的痛点。
Developer Tools
一句话生成应用 AI 编程助手 代码自主可控 实时构建展示 GitHub 部署 Vercel 部署 AI 客服嵌入 SaaS 无锁定 应用生成器 低代码开发
用户评论摘要:用户关注与 Lovable、Replit 等工具的差异,核心质疑两点:一是代码是否真正开放、能否导出后继续迭代;二是产品对新手的最快价值场景在哪。开发者回应强调自有账户部署、实时编码过程可见、内置行业 AI 客服,并采用先试用后付费模式。
AI 锐评

G.I.A.ac 切入了一个被不少玩家忽视的痛点:AI 生成的代码到底归谁?当 Lovable、Replit 们热衷于“平台内闭环”时,GIA 直接把代码推向用户的 GitHub 和 Vercel,看似是技术选择,实则是商业模式的叛逆——放弃托管抽成,转而对代码导出收费。这招聪明且危险:聪明在于,它精准击中了开发者和创业者的“锁仓恐惧”,尤其对曾因 Notion、Webflow 涨价而被迫迁移的群体而言,一句“your code, no lock-in”比任何华丽演示都更具说服力。危险在于,一旦导出后用户自行大改代码,GIA 的 AI 能否再次理解和迭代这些“被修改过的”代码?开发者在评论区的追问恰恰命中了这个命门——当前版本的 GIA 大概率依赖自身构建时的内部状态,对导出后再编辑的代码缺乏逆向理解能力。这会导致一个尴尬场景:用户要么永远留在 GIA 的构建环境中修改,要么第一次导出后就与 AI 辅助功能“断交”。这不叫真正的自主可控,这是“一次性的自主可控”。另外,标记“Built solo in Morocco”虽是情怀加分项,但单人维护的实时构建工具在并发、崩溃修复、功能迭代上的可持续性,投资人和早期用户都应保持审慎。一句话总结:GIA 在产品哲学上打对了牌,但后续“断联后如何续命”的技术难题尚未真正回答。如果它能实现跨语言、跨框架地理解和迭代用户导出的代码,那它才配得上“General”之名。

查看原始信息
G.I.A.ac
General Intelligence Architect — describe an app and watch it get built live.

How is it different from Lovable or Replit (or any other similar tools that are able to code your own solution)?

5
回复

@busmark_w_nika hey Nika, Great question — three real differences:

  1. Where your code ends up. GIA deploys to your own GitHub and your own Vercel accounts. Not hosted on my platform, no per-app hosting fees, export anytime. Most tools in this space keep your app inside their walls — that's the business model. Mine is: you own it, you can leave whenever.

  2. A real coding agent under the hood — not a loop over a pile of templates. You see the actual files being written, tested, and repaired in real time — not a loading bar and a reveal. If it makes a mistake, you watch it fix it. Nothing to take on faith — the build IS the demo.

  3. Every app ships with a built-in AI assistant that knows that specific business — services, prices, availability, pulled from the app's own data — answering that app's customers from day one. A nail salon gets a booking receptionist, automatically. I haven't seen another builder do this.

And the pricing philosophy: the full build + live preview is free, no card. You pay (€9.99/mo) only to download the code or deploy. Try before you pay, always.

Genuinely happy to be compared side by side — same prompt in both, see where the code lands. That's the test I built GIA to win.

0
回复
Type one sentence — "a booking site for a nail salon in Paris" — and watch G.I.A build it live: real code, working app, no black box. Every app ships with a built-in AI assistant that knows that business (services, prices, availability) and answers customers from day one. Deploy to YOUR own GitHub and Vercel — your code, no lock-in, export anytime. Try everything free in a full live preview; pay €9.99/mo only when you're ready to ship. Built solo in Morocco. Describe it. Watch it build. Ship it.
1
回复

@elhoucine_idousaid The deploy to your own GitHub and Vercel, export anytime part is the detail that actually matters more than the one-sentence generation most AI app builders keep you locked into their hosting, so you're one pricing change away from being stuck. curious what happens after the first export though: if I ask G.I.A to add a feature later, does it re-read my (possibly now-edited) exported code, or does it only work if I keep building inside the platform?

0
回复

Congrats, what early workflow usually gives users the quickest win?

0
回复
#19
RecipeBook by Shofo
Buy video training data by the hour featuring 25M+ clips
105
一句话介绍:RecipeBook是一个视频数据平台,让AI模型训练者通过语义搜索、投票排序,以3美元/小时的价格快速购买精准的视频训练数据,省去传统销售沟通和漫长等待的痛点。
Developer Tools Artificial Intelligence
视频数据平台 AI训练数据 视频数据集 语义搜索 数据标注 数据集购买 数据采集 机器学习数据 视频筛选 训练数据市场
用户评论摘要:用户称赞搜索后投票重排的效果和按小时购买的便捷流程。核心问题在于数据版权:用户担心“公开可用”不等于“商用合法”,可能面临诉讼风险。团队回应称数据收集不涉及登录、伪造账户等,并引用公开数据判例法。此外,元数据(如字幕、互动数据)目前较弱。
AI 锐评

RecipeBook的价值不在于“25M+视频”这个数字,而在于它精准切中了AI训练数据市场中一个被忽视的痛点:采购效率。传统模式下,无论自建还是购买,数据团队都要在工程消耗和销售沟通中耗费大量时间,而RecipeBook用“搜索-投票-按需购买”的闭环,将采购周期从数周压缩到一天内,这是真正的效率革命。3美元/小时的定价更是极具侵略性,直接撕开了高溢价数据市场的口子。

但产品最大的隐患写在了评论区:版权问题。创始人声称“收集方式不涉及违规”并援引公共数据判例法,但“公开爬取”与“商用训练”之间的法律灰色地带正在成为行业雷区。OpenAI、Stability AI都因类似做法吃官司,RecipeBook不能仅靠自说自话的条款免责。若无法提供明确的授权链或可追溯的合规层,对商业客户而言就是定时炸弹。

此外,元数据薄弱(无字幕、无互动数据)意味着搜索依赖纯视觉语义,这限制了高质量场景(如对话理解、情感分析)的可用性。产品是好产品,但想从“实验室好工具”升级为“企业级基础设施”,必须先解决法律合规和元数据深度问题——否则,低价和效率只是加速踩雷的油门。

查看原始信息
RecipeBook by Shofo
Recipe Book is a video data platform where you can semantically search 25M+ clips, iterate on the results, and buy the exact dataset you need to train your model in one sitting. Search -> rate a few clips -> we train a probe on your votes to re-rank the whole catalog to your taste. No contact forms, no sales calls, $3/hour.
What's up PH, I'm Braiden, co-founder of Shofo (YC W26). Every team training a model faces the same data decision: build or buy. Building means spending engineering talent on collection (reverse engineer APIs, deal w proxies, fight scraping defenses, filtering, maintenance) all to collect data sitting in the open. Buying isn't much better. Prices run from $15 to $480 per hour, and your ideal dataset still sits behind a wall of sales calls, samples, and weeks (sometimes months) of collection and iteration. That time can be the difference between SOTA and irrelevant. RecipeBook gives you the ability to search, filter, and build training datasets from millions of publicly available videos in the same day. You search in plain language, narrow with metadata filters (fps, resolution, duration, aspect ratio), and upvote and downvote clips based on needs. Once you vote, we train a classifier on your picks to re-rank the whole catalog to your taste, so you get the exact dataset you need. (Just a few votes can meaningfully change results) Once you're happy with results, you type in how many hours you want. Ask for 1,000 and you get a review sheet: the score distribution, a playable grid of the worst clips in the order so you can see the floor, and a trim slider where the count and price all live. Check out at $3/hour and a manifest CSV with download links hits your inbox within a couple minutes (or instantly depending on size of purchase). Biggest downfall is that metadata isn't super strong right now. No captions or engagement on the corpus (yet), so search is purely visual. (If you think it's something else please tell me!!) Our team has spent years in data collection and curation. We've indexed billions of videos across the internet and built systems to serve that data on request across all modalities. Along the way we learned the data you need to train your model is already public, you just haven't found it yet. If you're training a video or multimodal model, give it a try and tell me what you think. If you know a team sourcing large-scale video data (video gen, VLMs, world models, seed to Series A), intro us and I'll buy you dinner.
3
回复

@braiden_dishman1 Congrats on shipping this. The volume of data processed is impressive

1
回复

the search-then-buy flow is clever, but the thing i'd want answered before using this commercially is rights, not workflow. "publicly available" and "cleared for training a model you're going to sell" are two very different bars, and that gap is exactly what's driving a bunch of the current lawsuits against scrapers. is there any licensing/consent layer behind the 25M clips, or is it on the buyer to figure out whether a given clip is actually safe to train on?

1
回复

@omri_ben_shoham1 Hey Omri, when we collect data we do so in a way that does not require logins, agreeing to TOS, creating fake accounts, etc. Our bar for collection methods is quite high and our approach is supported by public data case law. There are terms on RecipeBook's website that provide a bit more detail as well as a takedown form if that's helpful!

1
回复

Searched for a tricky cooking motion I'd been hunting for and the re-rank after a few votes actually nailed it. The no-sales-call, pay-by-hour setup is refreshing.

0
回复
#20
SUB/WAVE
Self-hosted radio with an AI DJ and one shared stream
104
一句话介绍:SUB/WAVE能将你的本地音乐库变成一台真正的广播电台,通过AI DJ统一选曲、口播和接受自然语言点歌,解决流媒体时代个人听歌孤独、缺乏共享体验的痛点。
Music Open Source Streaming Services
自托管广播 AI DJ 本地音乐库 共享流媒体 开源 私有部署 音乐发现 电台体验 本地LLM
用户评论摘要:用户称赞其让音乐库“活起来”的共享体验;但多数人反馈已不使用本地文件,依赖Spotify等流媒体。有用户希望增加“仅播放久未听曲目”的强制发现模式,开发者回应部分功能已存在,并认可该建议。
AI 锐评

SUB/WAVE在“反算法”和“共享仪式感”上做了一次漂亮的复古创新。它精准瞄准了一个小众但真实存在的群体:拥有NAS、Plex或Navidrome的“数字囤积者”——这些人不是没有歌,而是缺一个让歌“流动起来”的叙事者。AI DJ不是取代人,而是扮演了“电台导播”角色:它懂歌曲过渡的物理特性(渐弱、骤停、分轨混音),能生成不同人格的串场词,甚至允许节目编排。这背后是对“信息分发方式”的重新思考:单个听众无法跳过或私密化播放列表,所有人共享同一段声音时间线,本质上是把音乐从私人消费拉回公共文化空间。

但它的局限性同样明显:完全依赖用户拥有本地高质量曲库,且对本地部署LLM和TTS有一定技术门槛,这注定无法进入大众市场。产品当前的核心价值不是对抗Spotify,而是为自托管生态增加一个“有温度的中枢”——它把原本死板的文件系统变成了有呼吸节奏的广播流。真正亮眼的是其底层架构设计:本地化优先、LLM仅用于内容生成而不传输元数据、对Streaming协议(Subsonic/Sonos)的兼容,这些才是开发者作为“开源基建派”的真实功力。如果未来能支持从YouTube等平台按需下载并源管理,或引入社区公共电台模式,可能会触发更大的裂变。目前来看,它更像一件“数字藏品”——属于那些愿意为自己音乐库建筑一座发射塔的人。

查看原始信息
SUB/WAVE
SUB/WAVE turns your music library into a real radio station: one shared stream, an AI DJ that picks tracks, reads idents and takes requests in plain language. Self-hosted, works with local LLMs, MIT licensed. Web, iOS, Android and desktop players.
Hey Product Hunt 👋 I'm Parminder, and for the last three months I've been building the radio station I always wanted. Streaming apps gave us infinite choice and somehow made music lonely. Everyone sits in their own algorithmic bubble. I missed radio: one broadcast, everyone hearing the same song at the same time, a voice between tracks that makes it feel alive. SUB/WAVE is that, self-hosted. Point it at your music library (Navidrome or any Subsonic server), give it an LLM (a local Ollama model is enough) and you get a 24/7 station: an AI DJ picks the tracks, does station idents, reads the time and weather, and takes song requests in plain language. People tune in from the web player, native iOS and Android apps, a desktop app, or anything that plays an MP3 stream, including Sonos and car radios. Some things I'm proud of: - It's radio, not a playlist. No skip button, no per-listener shuffle. You tune in and hear whatever's on. - Transitions are ending-aware. The DJ knows whether a song fades out or ends cold, and can even blend stems between tracks. - DJ personas, guest co-hosts with banter, and produced programmes with per-episode plans. - Everything can run locally: LLM via Ollama, local TTS, your own music files. No cloud required. - MIT licensed. 48 releases in three months, and it just hit 1.0 this week. There's a live demo station if you want to hear it before installing anything: https://www.getsubwave.com/listen I'd love to hear what you'd want from a station of your own. I'll be around all day answering questions.
3
回复

@perminder_klair Congratulations on the 1.0 launch, Parminder! 🎉 This is one of the most unique takes on AI and music I've seen. Bringing back the shared experience of radio—while keeping everything self-hosted and privacy-friendly—is a compelling idea. The ending-aware transitions and AI DJ personalities are especially impressive. Wishing you a fantastic launch day! I'm curious: what's been the most surprising way people are using SUB/WAVE so far?

0
回复

It looks good, but a problem for me and for most people now i guess is that we don't store our music files, just use spotify or apple music

2
回复

@kamil_infeld Fair, and you're right that it's most people. SUB/WAVE needs a library you actually hold, so if you're all-in on Spotify or Apple Music it isn't for you and I'm not going to pretend otherwise.

The people it's built for are the ones with a Navidrome or Plex server and a few hundred gigs of files they never get through. Smaller crowd, but it's a real one.

If you ever do pull your library down, it's here. Either way, thanks for looking.

0
回复

This app has been the highlight of my summer.

My music library feel alive now.

10/10 excellent project and incredible dev behind it.

1
回复

@tinkermesomething That's genuinely made my week. Thank you.

Curious what persona you settled on too, most people rewrite the built-in ones within a week.

0
回复

This is very cool! Like a self-sovereign Spotify AI DJ :D

1
回复

@conduit_design Thanks. Self-sovereign is a good way to put it. If you run Ollama and local TTS, nothing leaves the box at all. The only thing it phones out for is the weather.

Closer to a station than a DJ though. One stream, everyone hears the same track, no skip button.

0
回复

Love the self-hosted angle and the AI DJ taking plain-language requests is a great touch. One thing I'd love: a "discover from my own library" mode where the DJ only pulls from tracks you haven't played in a while, so it actually helps you rediscover music instead of leaning on the same favorites.

1
回复

@referralpro Thanks. Part of that is already in there: the picker builds its candidate pool from several sources, and deep cuts and rarely-played tracks are one of them. There's also a deep-cut skill where the DJ digs something out of the back of the library and talks about why it's been sitting there.


What doesn't exist yet is the hard version you're describing, something like "nothing I've played in the last 60 days, no exceptions". That's a pool-level filter and it's a good idea.

If you want to hear the current behaviour, the demo station is at getsubwave.com/listen. It'll play things I'd forgotten I owned.

0
回复
I probably wouldn’t replace Spotify with this, but I do like the idea of turning a personal music library into something that feels alive instead of just another playlist. The shared stream and local setup give it a character most music apps have lost. Congrats on the launch!
1
回复

@etiennegarcia Thanks. And you shouldn't replace Spotify with it, that's not the job. It only works if you already have the files.

The alive part is the bit I care about most. One stream, no per-listener anything, no skip button. You tune in and whatever's on is what's on. Turns out that changes how you listen more than I expected.

0
回复