Product Hunt 每日热榜 2026-06-04

PH热榜 | 2026-06-04

#1
Mailwarm 2.0
The email warmup tool, upgraded for deliverability.
498
一句话介绍:Mailwarm 2.0 是一款专为邮件营销人员设计的邮件预热与投递保障系统,通过模拟真实用户互动、监控信誉指标和提供专家诊断,解决新域名或受损域名频繁进入垃圾箱、投递率持续下滑的痛点。
Email Email Marketing
邮件预热 投递率优化 发件人信誉 SaaS工具 邮件营销 垃圾箱规避 域名监控 ESP适配 AI模拟互动 增长引擎
用户评论摘要:用户普遍认可“基础预热已不够”,主要困惑集中在:如何区分预热与真实信号?长期使用还是短期必要?DNS已配好却仍进垃圾箱的隐性原因是什么?以及如何修复已受损域名的信誉。
AI 锐评

Mailwarm 2.0 的“升级”并不在于颠覆性技术,而在于对行业痛点的精准补位——它承认了“基础预热已死”这一残酷现实。从评论反馈看,用户最急迫的需求已从“如何开始预热”转向“为何我什么都做了还是进垃圾箱”。这正是Mailwarm切入的缝隙:用动态非线性的发送曲线、基于不同ESP的差异化交互模式,以及人工专家介入,将运维门槛从“配置工具”拉高到“诊断系统”的高度。

但其价值并非无懈可击。首先,产品的核心护城河并非算法本身(模拟人性化的交互策略可被复制),而是长期积累的19K+客户行为数据和专家响应网络——这是时间壁垒,却也是成本负担。其次,从用户评论中暴露的“隐性痛点”(如IP级节流、ESP级阻断)来看,Mailwarm更像是“灭火系统”而非“防火墙”:它能诊断已被污染的声誉,但无法阻止共享IP池的突发性崩坏。对于中小团队,每月为“专家通话”付费到底值不值?这取决于他们对流失率的敏感程度。

更值得警惕的是产品定位的悖论:用户只在使用初期感知到“价值”(解决进垃圾箱问题),后续的监控与维护却成了隐形服务。正如评论所警醒的,“warmup”这个命名让用户本能把它当成一次性道具,而非长期服务。Mailwarm 2.0试图用监控和专家介入扭转这一认知,但若不能量化“避免了多少潜在的信誉损失”或“提升的投递率换算成多少成交”,它依然难以摆脱“卖铲子”的宿命——工具再精良,用户只关心金子挖得够不够轻松。

查看原始信息
Mailwarm 2.0
Most founders rely on email to grow, but emails don’t land in the inbox by magic. Mailwarm 2.0 is the premium email warmup and deliverability system built to give your emails the best chance of reaching the inbox. It combines automated warmup, real engagement, monitoring, infrastructure checks, and experts call available for every subscriber.

Like most founders, email has been our #1 sales channel since our first startup in Paris.

That’s what led us to build mailwarm and work on email deliverability since 2020.

And one thing became clear very early: your emails don’t land in the inbox by magic.

Back in 2020, we launched Mailwarm here on Product Hunt as one of the first email warmup tools.
It became #1 Product of the Day 🏆 and since then, we’ve helped 10,000+ founders, sales teams, agencies, and businesses improve sender reputation and avoid the spam folder.

But over the years, we learned something important: Basic warmup is not enough anymore.

Teams need real engagement signals, monitoring, infrastructure checks, and sometimes a real deliverability expert to understand what’s happening and what to fix.

That’s why we built Mailwarm 2.0.

Not just the original email warmup tool. A premium email warmup and deliverability system built to give your emails the best chance of reaching the inbox.

If email is part of your growth, tell me in the comments how you’re using it. We’ll take a look and help you improve your inbox placement.

25
回复

@thamibenjelloun Congratulations on the launch! 🚀

I completely agree that email deliverability has evolved far beyond basic warmup. With inbox providers becoming smarter, sender reputation, engagement signals, infrastructure health, and ongoing monitoring have become just as important as the outreach itself.

It's great to see Mailwarm evolve from a warmup tool into a more comprehensive deliverability platform. Helping 10,000+ businesses improve inbox placement is no small achievement.

Curious to know what's the most common deliverability issue you see today among founders and sales teams that they often overlook?

0
回复

@thamibenjelloun How are you separating warmup from true deliverability signals especially for teams that already have decent sending volume but still see inbox placement drift?

6
回复

@thamibenjelloun congrats on the launch team. Other than the usual culprits (dkim/spf/dmarc etc) what are you seeing Mailwarm flag that is otherwise hard to track down?

5
回复

Hey Thami,

Cool launch! Quick question: do you see Mailwarm as something companies should use long term or only when launching on new domain?

7
回复

@bhouy  To be honest, this has been one of our biggest challenges.

When we launched Mailwarm 6 years ago, “email warmup” was barely a known category. But the word warmup makes people think it’s only for the beginning, like a short-term setup step.

After seeing thousands of senders over the years, our view is different: warmup is useful long term to keep positive engagement and protect sender reputation. Or at least, it should be reactivated when reputation drops or before scaling a new campaign.

So I’d say: new domain = warmup is critical. Existing domain = warmup + monitoring is how you avoid silent reputation drops.

Are you asking because you see warmup as a one-time setup, or because you’re thinking about long-term deliverability?

4
回复

@bhouy This sub-thread is the most interesting part of the launch imo. The product delivers value continuously (reputation protection), but the user only feels it once, at setup. That gap between delivered value and felt value is where most SaaS churn lives. Naming the category "warmup" made adoption easy and retention hard

0
回复

Hey Product Hunt 👋

Years ago we launched the first Mailwarm right here. That launch started a beautiful journey for us, and I'm really grateful for it.

Now we are back. With more experience this time. We spent a lot of time listening to our customers, adding features they really asked for, and also removing the ones that only made things complex without real value.

And honestly this last part was our biggest fight inside the team. Where is the line between useful and too much? If you add too little, the product feels empty. If you add too much, it becomes the heavy thing you wanted to avoid in the first place.

So I'm curious, how do you draw this line as builders? Where do you stop adding? And what is one feature in your own product you think you should drop?

Mailwarm 2.0 is the answer we found. Simpler, sharper, built on everything we learned the first time. Would love to hear your feedback 🙏

And a big thank you to @garrytan for hunting us.

7
回复

Hi Product Hunt 👋

Mailwarm 2.0 was interesting because email deliverability looks simple from the outside, but technically there are a lot of moving parts behind it.

Warmup activity, inbox interactions, sending behavior, domain setup, reputation signals, monitoring… everything needs to work together if you want the system to be useful and reliable.

From a dev perspective, the challenge was not just to automate email warmup, but to make the whole process easier to track, understand, and act on.

A lot of deliverability problems are invisible until performance starts dropping, so we focused on building something that helps teams see issues earlier and avoid guessing.

Excited to launch Mailwarm today and hear what people think.

6
回复

This is exactly what I need right now! Been struggling with deliverability issues since launching my new domain. The concept is simple and the problem is real. Congrats on the launch!

6
回复

@ouidad_yassin Thank you. Don't forget to book a call^_^

1
回复

@ouidad_yassin I feel you new domain deliverability is a massive headache right now, so you are definitely not alone in that struggle.

Thanks for the support on the launch, and here’s to getting your emails straight into the primary inbox! Let us know if you need any help navigating it ;)

0
回复

@ouidad_yassin Thank you! We are here to help with the struggles!

2
回复

Hey Product Hunt community 👋

I spent 8 years in roles where I was sending endless follow-up emails, onboarding messages, CRM emails, and customer reminders.

I even worked at a CRM SaaS where I would help users with email-related issues without fully understanding what was happening behind the scenes. Whenever an email disappeared, the advice was often: “Can you check your spam folder?” or “Maybe it’s a DNS issue?”

Working on Mailwarm made me realize how much is actually happening before an email reaches the inbox.

I learned that hitting “send” is only the visible part. Behind every email, there is sender reputation, engagement, domain setup, sending behavior, inbox placement, and a lot of small signals that can decide whether your email lands in the inbox or not.

We wanted to help teams warm up their inboxes properly, build better sender reputation, monitor what’s happening, and get real guidance when something starts going wrong.

Because most founders, sales teams, and marketers don’t want to become deliverability experts. They just want their emails to reach the inbox and know what to fix when they don’t.

The team and I are live today, happy to answer your questions and hear how you’re handling deliverability 🚀

6
回复

@othman_katim congratulations on the launch 🙌, this is a big problem you're tackling as email rules have become way stricter with the new wave of agents!

I'm curious how long does it usually take to warm up, and what's the reasoning behind the default numbers? (did you test on different ones before settling on those?)

5
回复

@skander_karoui Usually 4-6 weeks from a cold start, faster if there's an existing reputation. The defaults come from patterns across +19000 customers. We tested steeper and flatter curves early; both underperformed the current model. Are you warming a fresh IP or one with sending history?

0
回复

Big congratulations to the Mailwarm team on the launch! 🚀

We just started running an email campaign ourselves and discovered that a significant portion of our emails was landing in spam. We had already completed the warm-up process, configured SPF, DKIM, and DMARC correctly, and followed most of the recommended best practices, so it was surprising to see deliverability issues persist sadly.

We searched for a solution that could help diagnose and improve inbox placement, but we couldn't find anything that addressed the problem end-to-end....

One question tho! For users who have already warmed up their domains and configured everything but still experience poor inbox placement, what are the most common issues your platform tends to uncover?

5
回复

@rania_rimali Great question. Surprisingly, in many cases the issue isn’t authentication or warm-up at all.

The most common problems we uncover are low sender reputation inherited from previous activity, poor list quality, engagement signals that mailbox providers don’t like (low opens, replies, or positive interactions), content patterns that trigger filtering, and inconsistencies between sending behavior and recipient expectations.

We also frequently see domains that are technically configured correctly but have underlying reputation issues that aren’t obvious from standard SPF/DKIM/DMARC checks.

That’s exactly why we built the platform, to help identify the less visible factors that impact inbox placement and provide actionable recommendations beyond the usual setup checklist.

Out of curiosity, what percentage of your emails were landing in spam after you completed the warm-up process?

0
回复

@rania_rimali Thanks for the support! You are definitely not alone in this this is the hidden nightmare of modern email marketing. Modern spam filters don't just look at your setup anymore; they track ongoing behavior.

When a team tells us their DNS is perfect but deliverability is still bad, Mailwarm usually uncovers IP-level throttling (your email service provider's shared network is compromised), ESP-specific drops (being blocked by Outlook while Gmail is fine), or velocity triggers (sending too fast within short windows).

Deliverability is a moving target, and sometimes a single bad campaign or a dirty IP pool can undo your hard work. We'd love to help you pinpoint exactly where the leak is

1
回复

@rania_rimali Happy to jump on a call to audit your setup and define the recovery plan!

1
回复

Inbox placement is a deliverability problem, but most warmup tools treat it as a volume problem and just blast engagement signals until the ESP gets suspicious. Curious what actually changed in 2.0, specifically whether you're doing anything smarter around ramp curves and sending patterns, or whether the upgrade is mostly on the dashboard and reporting side. Also wondering how you handle situations where a domain's reputation is already damaged before warmup starts, since that's a different problem than warming a fresh domain.

Congrats for the launch

5
回复
@fberrez1 we continuously change how we interact. We also diversified our kind of interactions to be the most human possible. Honestly, we are freaks about this topic, and believe we are far better than anyone else in the market. But all this is stuff people can’t see. When reputation is damaged, it’s really important to take time, always have alternative domains. But every case can be different, that’s why calls with our experts are included in our subscriptions, they will help and give advice for every case separately. It feels you are speaking from experience, what’s something you feel missing in all placements and deliverability space.
3
回复

@fberrez1 Ooh very good question, you are entirely right: ESPs are incredibly smart now, and treating deliverability as just a volume game is a fast way to getting a domain blacklisted.

The upgrade isn't just cosmetic. 2.0 introduces dynamic, non-linear ramp curves. The algorithm constantly randomizes sending intervals, reply footprints, and engagement spacing. It acts like an organic human inbox pattern, preventing ESPs from identifying a mechanical warmup signature.

& For a fresh domain, you're building a foundation. For a damaged domain, you are in triage. The system prioritizes pulling emails out of spam, marking them as safe, and generating realistic thread replies. It's about feeding the ESP signals to work on the trust baseline before any real volume ramp-up is even allowed to start.

Thank you for your support :))

3
回复

@fberrez1  Exactly. The majority of platforms try to make you think that one-size-fits-all breaks down fast. Gmail's algorithm isn't Outlook's, they weight engagement, authentication, and sending behavior differently. Generic warmup won't catch that.

0
回复

Bonjour Product Hunt 👋

I spent the last 15 years in digital marketing, and a big chunk of that running email at scale at one point, sending over 1.5 million emails a day as an affiliate across more than 30 brands. When you want to operate at that volume, you learn one lesson fast: sender reputation isn't something you set up once. It's something you build, every day, with behavior. And if you don't, perfect DNS, perfect copy, perfect lists won't save you.

That's the gap most teams hit. They configure SPF/DKIM/DMARC, write great emails, build clean lists, and still land in spam. Because reputation isn't a configuration. It's a pattern of activity that inbox providers learn to trust over time. New domains have none of it. Quiet domains lose it. Aggressive senders blow through it.

Mailwarm is what we built to solve that, alongside @bengeekly and @thamibenjelloun. Today is a big one for us; we're launching Mailwarm v2, with real-time reputation monitoring, a smarter warm-up per ESP engine, content warm-up, and more. It's the biggest update we've shipped since launch.

If you have poor email performance, drop your situation in the comments, and I'll dig in.

Excited to launch today 🚀

5
回复

@bengeekly  @thamibenjelloun  @othman_katim Congratulations on the launch!

0
回复

@othman_katim  Bonjour 👋, how long does it typically take before someone sees improvement in their inbox placement rates? Congrats, amazing team !!

0
回复

Hi everyone 😁

Email deliverability is usually something people notice only when results start dropping, and it’s not a good thing as this negativity affects your campaigns.

That’s exactly why we worked on Mailwarm, to help teams warm up their inboxes properly, keep track of what’s happening, and understand what needs to be fixed before deliverability becomes a bigger problem.

We are excited to launch this today, and we are also curious to know how the Product Hunt community is handling email deliverability.

5
回复

Congrats on the launch!

4
回复

@richardzhang Thank you for the support !

1
回复

@richardzhang Thanks, I made a quick check for your domain, happy to share with you some quick wins .

0
回复

Congrats @thamibenjelloun ! Much needed.

3
回复
Good luck
3
回复

@dmitry_zakharov_ai Thank you, same to to you

0
回复

@dmitry_zakharov_ai Just reviewed your infrastructure. Few things worth tweaking that could move the needle on placement.

0
回复

Have been using mailwarm recently, excited to see what has changed here

3
回复

Thank you for your trust @abhishekr_ai. This is the original mailwarm, upgraded. Check it here :

5
回复

@abhishekr_ai Great to have an existing user in the thread! You are going to notice a massive difference with 2.0. Can't wait for you to try it out. Let us know your thoughts once you dive in :))

1
回复

The landing page look is pristine. Quick question on the "infrastructure checks"does the tool actively monitor for configuration drift (like if a teammate accidentally messes up the SPF/DKIM records downstream), or is it just an onboarding check? good job team

3
回复

@vikramp7470 To help you get the best possible results right from the start, we’ve integrated with our sister tool, Themailx, allowing you to check your DNS records instantly!

To back this up, our deliverability experts also conduct a thorough, manual audit of your domain. We verify your DNS setup, check your domain health against major blocklists, and send you a personalized email with clear steps for any needed optimizations.

A successful warm-up relies on a flawless foundation. While MailWarm does the heavy lifting, it works best when your domain infrastructure is perfectly aligned, and we’re fully committed to guiding you every step of the way to maximize your deliverability.

1
回复

@vikramp7470 Thank you so much for the kind words about the landing page! The team worked hard to keep it clean.

Regarding the infrastructure checks: It is 100% active, ongoing monitoring.

We know firsthand how easily a downstream teammate can accidentally overwrite a DNS record months after onboarding. Mailwarm doesn't just check your records at the start; it continuously tracks your SPF, DKIM, and DMARC health in the background.

If configuration drift happens, the tool catches the signal drop and flags it on your dashboard early, saving you from a deliverability crisis.

Appreciate you stopping by and asking such a good question :)

0
回复
@vikramp7470 for SPF/DKIM we verify it every time you click on it. We have live monitoring in our next product! But it can be a good feature to add
4
回复

Hi Product Hunt Community!

I’m Naim, and I handle the technical account management side of things at Mailwarm. If there’s one thing I’ve learned the hard way, it’s this: when your emails start hitting spam, it’s almost never for just one reason.

It’s a ghost in the machine. Maybe a DNS record is slightly off. Maybe a sudden spike in sending volume triggered a filter. Or worse ... your domain reputation has been quietly tanking for weeks, and you only notice when your reply rates hit zero. It is incredibly painful to waste days tweaking random settings in the dark, hoping something works.

That frustration right there is why Mailwarm was built.

The goal is to turn the black box of email deliverability into a transparent dashboard. Mailwarm helps you safely warm up your sender reputation, track the right technical signals, spot infrastructure issues early, and understand exactly what needs attention.

Cold outreach and email marketing are still unmatched growth channels but only if your audience actually sees what you send...

We are back with Mailwarm 2.0 🤘🏻! We’d love your support, your feedback, and for any email questions just drop them in the comments below :))

3
回复

Just curious about how long it typically takes to see a noticeable improvement in deliverability? Asking because most cold outreach tools promise inbox placement, but the warm-up window is where campaigns usually stall.

2
回复

@arjayyy Usually 3-4 weeks for noticeable improvement, depending on your starting reputation. You're right that the warmup window is where campaigns stall, most tools rush it. Gradual ramp clears placement faster than aggressive volume.

0
回复

@arjayyy We get this kind of question every day, and honestly, it’s smart to ask before starting.

It depends a lot on the domain: fresh or aged, past sending history, current reputation, provider, list quality, and how aggressively you start sending while warmup is running.

Some users see strong movement quickly (check this comment with a screenshot : https://www.producthunt.com/products/mailwarm?comment=5426071).

For example, this account went from:

  • First test: 75.64% spam

  • Previous week: 19.72%

  • Current week: 13.61%

Another went from 54.55% spam to 7.29%, then slightly up to 12.14% once campaigns started, which is normal because real sending behavior also affects reputation.

So warmup is not magic. It improves the reputation layer, but your actual sending setup and behavior still matter.

What’s your situation: fresh domain, aged domain, or already sending cold outreach today?

1
回复

How much time we need to wait for warm up to complete before starting email campaign. Is this fully automated?

2
回复

@mavelstech Usually 3-4 weeks for noticeable improvement, depending on your starting reputation. The platform is fully automated, and the setup takes you literally 2 minutes.

0
回复

@mavelstech I wouldn’t think of warmup as something that “finishes” once and then you forget about it. The timeline depends on your domain age, past sending activity, current reputation, provider, and the volume you want to reach.

For a new domain, I’d start with warmup first, then add campaigns progressively while watching inbox placement and spam signals.

Mailwarm automates the warmup itself, but Mailwarm 2.0 also adds monitoring and access to our deliverability team when you need help adjusting the strategy.

Are you preparing a brand-new domain or already sending from this inbox today?

0
回复

We use a similar service. How many warm-up accounts do you have, and how often are they rotated? Because if, for example, the warm-up is done using the same 100 mailboxes over and over, it will lose its effectiveness after a month.

2
回复

@natalia_iankovych You're right, small networks kill the warm-up positive effect; we operate +50 000 across the major providers, and they rotate daily.

0
回复

@natalia_iankovych To be honest, I am happy and proud to write this answer.

And you’re 100% right. If warmup runs on the same small pool of inboxes over and over, it quickly becomes repetitive and loses value.

Mailwarm runs on a network of 50,000+ real and aged inboxes (by the way, mailwarm exist since 2020). The goal is not just to “send warmup emails,” but to create diverse, positive inbox interactions across providers, with enough rotation, email provider mix, and engagement behavior to support sender reputation over time.

That’s also why we built Mailwarm 2.0 beyond basic warmup: monitoring, infrastructure checks, provider visibility, and expert review. We want teams to know if their reputation is actually improving, not just see activity running in the background.

Are you mainly using warmup for cold outreach or for a larger email marketing setup?

0
回复

Congrats on the relaunch! it's bold move going paid-only with no free tier. That makes the first week after payment the real moment of truth. Do you know which early signal makes people stay vs leave or ask for a refund: first email pulled out of spam, the reputation graph ticking up, or something else? Curious how sharply you can see that point, like aha-moment

2
回复

@and_bayleaf Andrew, great question.

For us, the “aha moment” is when the customer sees that Mailwarm is not just sending warmup emails in the background. It actually helps them protect reputation and track improvement.

A result like this (screenshot) is exactly why people stay.

When you see that kind of movement, you understand the value immediately.

The real retention signal is when users stop seeing warmup as a one-time setup and start seeing it as reputation protection + monitoring over time.

That’s why Mailwarm 2.0 adds the expert-led layer too: if numbers don’t move, our team helps you understand what to fix.

6
回复

@and_bayleaf Thanks. Free tiers attract a lot of profiles who never implement anything, and warmup only works when someone's committed to the process. Going paid means we can focus our deliverability team on people genuinely invested in fixing placement, which keeps service quality high for everyone.

On the aha moment: it's the reputation graph ticking up paired with that first batch landing in the inbox instead of spam. Usually shows within the first 20 days. That's when it clicks, they see the trend line moving and the placement data backing it up.

1
回复

super exciting product! congrats

2
回复

@raphael_goldsztejn Thank you! Would absolutely love to hear your thoughts about it.

0
回复

Cool! Crush it! Let's do an integration with Flowlu.com

2
回复

@gb1010  Done!

3
回复

This looks promising. I am going to start cold emails and this looks worth checking out. All the best @thamibenjelloun

2
回复

@neelptl2602 Thanks a lot Neel. Starting cold email is exactly the right moment to think about deliverability, before volume increases and reputation problems become harder to fix.

Happy to have @othman_katim or @manal_essalek1 jump on a quick free call with you to review your setup and help you start clean.

Are you starting with a fresh domain or an existing domain with some sending history?

1
回复

Cold email works best and will be always their. A dedicated warup tool like mailwarm is a must needed one. Looking forward to using it. Congrats on the launch ( :

2
回复

@istiakahmad Thank you so much for the support. Do forward us your feedback once you use it :)

0
回复

Congrats on the launch team 🚀 6 years in and still shipping is the hard part.

Quick notes from the cold outreach trenches: per-ESP reputation monitoring is the right call, Gmail, Outlook and Yahoo weight things so differently that one blended spam score always hid where the real leak was. The non-linear ramp curves are also smart; mechanical warmup signatures are exactly what ESPs got good at catching.

Rooting for the relaunch 👏

2
回复

@saad_el_gueddari Wow, thanks so much! Longevity in this space is hard, but it’s operators like you who keep us pushing to build better tools.

It’s incredibly validating to hear you call out the non-linear curves and per-ESP monitoring. The team spent so much time refining the engine to dodge those mechanical warmup signatures you mentioned, because Google and Microsoft have gotten terrifyingly smart at catching basic automation.

Thanks for the launch day support :))

1
回复

Inbox placement is one of those things people only think about when it suddenly stops working 😄 Love the focus on deliverability beyond basic warmup. Congrats!

2
回复

@alina_tyslenok_ 100% and of course beautiful emails are part from the equation.
The biggest hack we have is to get people answer by asking questions.
Any design hacks you have that make people answer from your experience with stripe.email?

0
回复

Congrats.
Just a question, does it work for cold outreach domains or mainly for newsletters?

2
回复

@felixlandicho Thanks! 😊

Mailwarm works for both cold outreach and newsletters. We support B2B and B2C use cases, so whether you're sending sales emails, outbound campaigns, or newsletter content, the goal is the same: helping you build and maintain a healthy sender reputation for better inbox placement.

Just to be sure, what's your primary use case right now cold outreach, newsletters, or a mix of both?

1
回复

@felixlandicho It applies to all the email senders that consider this channel seriously to generate revenue and deliver a great customer experience

1
回复

the entire landing page and website look absolutely well-prep and psychologically made!! it actually hits my pain point of restarting the warm up circle for every new SaaS launch.

congrats Thami and Othman!!!

1
回复

@nathan_tran2 Appreciate that. The restart-every-launch pain is exactly why we built it to manage everything in auto pilot. Glad it landed. Thanks for the kind words.

0
回复

Mailwarm 2.0 sounds interesting. We rely heavily on email for growth, and anything that gives real deliverability insights is worth checking out

1
回复

@volodymyr_kreschenko It absolutely is. If email drives your growth, flying blind on deliverability is like running a sales team with the headlights turned off. You think you're moving fast, but you have no real visibility into what is actually making it to your prospects.

With the strict filtering rules major inbox providers have enforced recently, old-school "bot-driven" warmup tools just don't cut it anymore. Mailwarm 2.0 caught a lot of traction because it shifts the focus from vanity activity metrics to live, actionable placement data.

The Real Upgrades in 2.0:

  • Live 24/7 Placement Testing: It shows you exactly where your emails are landing (Inbox vs. Spam) across Gmail, Outlook, and Yahoo in real time, rather than giving you a delayed guess.

  • Human-Mimicking Network: It taps into a network of 50,000+ real, aged inboxes to simulate genuine human behavior (opening, marking as important, replying with realistic threads) which satisfies modern, AI-driven spam filters.

  • Infrastructure Auditing: It constantly monitors your SPF, DKIM, and DMARC health alongside major blacklists so a technical glitch doesn't quietly tank an entire outreach sequence.

When email is your primary growth engine, even a minor dip in sender reputation can quietly trigger a massive revenue leak.

What is the biggest deliverability blind spot or specific friction point you are running into with your campaigns right now?

0
回复
#2
Astra Autonomous Pentest
AI agents that find, validate, and fix every vulnerability
336
一句话介绍:Astra Autonomous Pentest 利用AI代理军团,自动化完成从漏洞发现、利用验证到修复的全流程,旨在解决传统渗透测试周期长、误报率高、修复落地难的核心痛点,让软件实现“自我修复”。
SaaS Developer Tools Security
AI渗透测试 自主安全 AI代理 漏洞验证 业务逻辑漏洞 持续安全测试 零误报 AI修复 DevSecOps 应用程序安全
用户评论摘要:用户高度关注AI代理处理复杂业务逻辑和链式漏洞的能力,特别是跨上下文场景。主要疑问包括:如何保持误报率为零?是否能覆盖带有MFA等复杂流程的认证后攻击面?AI代理执行攻击动作时,如何保障生产环境安全及回滚?修复代码是否直接生成PR,及如何确保不引入新缺陷。
AI 锐评

Astra Autonomous Pentest 的登场,与其说是“AI渗透测试工具”,不如说是对传统“安全审计报告”模式的颠覆。它的核心价值不在于“扫描得更快”,而在于构建了一个“发现-验证-修复”的闭环,将安全测试从一次性的“检查点”升级为持续性的“内建流程”。

其真正的锐利之处在于两点:一是将“假阳性”驱近于零的独立验证层,这直接击中了传统DAST工具的七寸——大量误报让安全团队沦为“报告处理机”;二是将修复建议直接输出为Cursor、Copilot等开发者工具的Prompt,实现了从“报告”到“代码修复”的微观落地,将安全责任真正左移到开发者手中。

但挑战同样尖锐。用户评论中关于“业务逻辑边界”和“合规策略”的提问,暴露了当前AI在理解“业务上下文”上的天花板。商业软件的漏洞往往不在于技术实现,而在于“业务规则允许但安全不允许”的逻辑歧义。如果Astra仅能处理“技术漏洞”而无法理解“业务逻辑漏洞”,其“自主穿透”的叙事就会显得单薄。

此外,AI代理在真实生产环境中的“越狱”风险是悬在头上的达摩克利斯之剑。尽管官方强调“只读Payload”和“模拟验证”,但一旦攻击路径失控,其造成的破坏力远超任何传统扫描器。这不仅是技术问题,更是信任问题——何时企业敢让一个不请自来的“AI黑客”在自己的生产环境中自主游荡?

总体而言,Astra为停滞不前的安全测试领域注射了一剂强心针,定义了“自我修复软件”的起点。但它离真正的“自主”仍有距离,尤其是在处理复杂、模糊的业务逻辑时,人类渗透测试专家的经验和直觉,至少在现阶段,仍无法被完全替代。这更多是解放了白帽子,而不是取代了他们。

查看原始信息
Astra Autonomous Pentest
Astra Autonomous Pentesting makes self-healing software the new standard, a category we’re defining after 8 years and 5,000+ real-world pentests. An army of offensive pentesters and bounty hunter agents that discovers complex chained vulnerabilities, an independent validator layer drives false positives to near-zero, and AI-fix agents deliver remediation as native Cursor, Copilot, and Claude Code prompts. The reactive pentest era is over.

Hey Product Hunt 👋

I'm Shikhil, the founder of Astra Security. I did my first pentest 15+ years ago and have been obsessed with offensive security ever since.

Over the years, we built a PTaaS platform, a DAST scanner, API Security platform, a Cloud Vulnerability Scanner - and discovered tens of millions of vulnerabilities along the way. But one belief stayed constant through all of it: business logic vulnerabilities would never be discovered autonomously. Ever.

AI just shattered that limit. And nothing has excited me like this in 15 years of being in infosec. 🤯

So we built Astra Autonomous Pentesting. Not a smarter scanner. An army of AI agents that owns the full pentest cycle:

  • 🔍 Discover - Offensive agents built on insights from 5,000+ real-world pentests hunt complex, chained vulnerabilities.

  • 💥 Exploit - Agents chain and exploit findings to prove real-world impact, not flag theoretical risks.

  • Validate - An independent validator layer drives false positives to near-zero.

  • 🔧 Fix - AI-fix agents that deliver tailored remediation right in your Cursor, Copilot, and Claude Code.

The full cycle. No handoff. No report sitting in someone's inbox. Software that heals itself.

This isn't about replacing pentesters 🙏 Let AI own the grunt work - the cookie flags, the report writing, the endless threat modeling sessions. Let pentesters do what they love: chaining complex vulnerabilities, getting deep into a system. Pentesters at Astra, are central to everything we build. Now AI is their most powerful ally, not their replacement.

We call this the era of self-healing software. And we're just getting started. Would love your questions, brutal takes, and your support today. 🚀

Looking forward to help you with your next Pentest!

— Shikhil, Founder & CEO, Astra Security

25
回复
@shikhilsharma Great product Without left shift checking and reducing false positives and giving a report user.
0
回复

@shikhilsharma @abhishek_krishnan5 Congrats on the new launch! :)

Whoa! I saw this on the leaderboard and immediately sent it to a security testing agency I have been working with. They asked me to pose this question to you... what specific approach or training data enables these agents to understand context-specific business rules that vary across different applications?

6
回复

@shikhilsharma Congrats on the launch! How do you handle the edge case where AI agents chain vulnerabilities across business logic boundaries that require context outside the system (e.g., a workflow that's technically valid but violates a specific compliance rule or business policy)? Do you have a feedback loop where human pentester insights on these edge cases automatically update the agents' chaining logic, or is that still manual?

0
回复

Super excited for this one!

4
回复

Hey everyone 👋

I'm Shelton. I lead marketing at Astra, but I'll skip the pitch and share what actually made this click for me.

Most automated scanners run off a static checklist. They catch the obvious stuff and miss anything that needs context. Astra Autonomous Pentesting builds a threat model from your real application first, then the AI agents target vulnerabilities that only surface when several steps chain together: multi-step attack chains, IDOR, broken access control, business logic flaws, and the full OWASP Top 10. The kind of issues you'd only catch when a human pentester spends a week with your app.

Two details I think matter more than any headline number:

  • Every finding gets vetted by our security team before it lands on your dashboard, so you're not digging through false positives.

  • It runs safely in staging or production with rate limits and controlled attack patterns, no destructive actions, and you set the scope and intensity yourself.

Shikhil already covered the bigger picture, so I'll leave it there. If you've used autonomous or continuous testing before, I'd like to know what it got right for you and where it fell short. And if you think we've missed something, say so.

Thanks for taking a look 🙏

2
回复

Congrats on another launch! Was wondering... discovering and validating is one thing, but you're actually chaining auth bypasses and privilege escalation against a live target to prove impact. That's a real agent taking real destructive actions. What happens the first time it escalates into something it can't cleanly roll back, mid-run on someone's prod?

2
回复

@artstavenka1 Thanks for the support!

Our AI operates with strict boundaries, using a read-only payload mindset to intelligently demonstrate impact and chain logic flaws without triggering destructive mutations or configuration writes in production.

For the Validator layer, we actively simulate exploit paths rather than running irreversible exploit code, allowing us to mathematically verify the risk while keeping your state completely clean.

1
回复

My congrats! Does the autonomous pentest cover authenticated flows out of the box, or does that require manual configuration? Asking because most scanners struggle with post-login attack surfaces.

2
回复

@igorsorokinua Great question, and you are right that most scanners completely fall short here.

Astra's Autonomous pentesting handles authenticated flows out of the box. You provide login credentials and optionally a login recording for complex flows like MFA or CAPTCHA, and the AI agents log in as multiple user roles to crawl and test everything behind the login wall.

This is actually where AP finds the most critical vulnerabilities, things like IDOR, privilege escalation, and BOLA that only surface in authenticated sessions.

1
回复

This is the kind of product that makes security accessible to teams that can't afford a dedicated red team or quarterly pentests. Democratizing offensive security is a big deal, congrats @abhishek_krishnan5 @ananda_getastra @saurabh_miglani

2
回复

Thank you @kate_ramakaieva 🙌🏻 💙

0
回复

Hi Product Hunt 👋
Thank you all the great questions and interest that you folks are showing on our new product. After months of hard-work, we're super excited to finally see this out in the world!
Looking forward to see it in action on all of your applications. Helping you scale, while staying secure!

1
回复
all the best team
1
回复

Thank you @pratyush_r8 - btw, I use Merlin :)

0
回复

Congrats! @shikhilsharma

1
回复

Thank you@danielwayne 🙌🏻

0
回复

Congrats on the launch! The idea of AI agents autonomously discovering, validating, and even suggesting fixes for vulnerabilities is impressive. Excited to see how this shapes the future of pentesting.

1
回复

Thank you@sanket_goyal 🙌🏻

0
回复

Super excited to launch Autonomous Pentest today 🚀, we've set up a 50% discount for the Product Hunt community. Just head over to the link and use the code at checkout. Would love to hear what you think 🙂 The offer is available for a limited time. Experience the future of security testing today. 🚀

1
回复

Congrats on the launch, liked the steps to reproduce and suggested fix approach too

1
回复

Thank you@biswaviraj 🙌🏻

0
回复

Delivering remediation directly as native Cursor, Copilot, and Claude Code prompts is a highly practical workflow. However, how do your 'AI-fix agents' guarantee that the suggested code changes completely resolve the vulnerability without inadvertently breaking existing business logic or introducing new flaws?

0
回复

How does the threat model get generated, is it based on the app's structure discovered during scanning, or does the user define it manually?

0
回复
Love the concept here. The loop from discover to fix looks super smooth on the graphic. Just curious on the remediation side, does the agent actually write the patch or pull request for you, or does it just give you the instructions on how to fix it manually?
0
回复

The remediation-as-Cursor/Copilot/Claude Code prompts angle is interesting. The part I’d want to see in practice is how the validator keeps a clear audit trail from finding → exploit proof → suggested patch, because that handoff is where security workflows usually get messy.

0
回复

Congrats on the launch. This looks really promising. Although you don't currently do auto-remediation, are there plans in the future for that kind of capability?

Does it focus on known vulnerability types or does it also look for new patterns?

0
回复

@edward_g you can set a quick Claude pipeline: run APs with CI/CD, MCP to get fix details, and create a PR. With this pipeline, you can have a human in the loop and control over the fixes.

0
回复

What if I’m a developer and need to quickly audit a client’s website just by providing the site URL? Is that possible? Does it generate a report after the audit? That would be very helpful for selling my services.

0
回复

@natalia_iankovych That is precisely the use case: point Astra at the URL, and the agents do the rest, returning a full report with validated findings, steps to reproduce, and contextual fix recommendations your client can actually act on.

0
回复

Hey team! What's the integration story with GitHub Actions / GitLab CI? Would love to trigger a scan on every PR merge.

0
回复

@mikhail_prasolov Astra integrates seamlessly with GitHub Actions and GitLab CI, allowing you to automatically trigger automated scans right on every PR merge.

0
回复

Love the validation layer approach. How do you keep AI fixes safe in high-sensitivity environments—do you require human approval or enforce policy constraints before any remediation prompt gets applied?

0
回复

@leventbuilds We never execute fixes directly; we provide the precise code blueprint to your dashboard, keeping the ultimate "merge" button firmly in human hands. Our AI operates strictly as an advisor under tight policy guardrails, allowing your engineers to review, test, and safely apply the contextual fixes themselves.

0
回复

Congrats on the launch! Secure web apps is what we need today.

Does your project work with source code only? (to my understanding, in the CI pipeline) Can it also analyze, for example, minified or obfuscated client code on a live or sandboxed website?

0
回复

@nikitaeverywhere Thank you!

No source code needed. You point Astra at a live or sandboxed URL and the agents work from the outside, the way a hacker would.

On minified and obfuscated client code, the agents don't analyse the bundle statically. They interact with the running application, observing API calls, endpoint behaviour, and server responses.

CI/CD integration works via API, trigger a scan against your staging environment before every deploy and get findings before anything reaches production. The only thing source code is used for is the fix delivery layer, where agents read your codebase to generate contextual fixes. The pentest itself needs nothing beyond a URL and user credentials.

0
回复

Congrats on your launch!

0
回复

Thank you @ronakagarwal3434 🙌🏻

0
回复
#3
Empromptu AI
Train Fine Tuned Models With AI Apps You're Already Building
281
一句话介绍:Empromptu AI 通过自动捕获真实用户交互、人工修正和边缘案例,将你在构建的AI应用实时产生的数据转化为你可拥有的微调模型,解决模型优化依赖专家且无法持续从业务经验中学习的痛点。
Developer Tools Artificial Intelligence No-Code
AI模型微调 自动数据捕捉 人工修正反馈 边缘案例学习 模型所有权 模型无关 评估过滤 知识资产化 实时优化 SaaS工具
用户评论摘要:用户关注数据质量(如何过滤坏反馈)、模型演进(基础模型升级后是否失效)、因果剥离(效果来自平台还是模型升级)、技术集成(是否兼容LangChain)、以及专家冲突信号处理。核心建议:需让用户看到特征变化,过滤嘈杂修正而非直接训练。
AI 锐评

Empromptu AI 的巧思在于将“生产环境的AI使用数据”作为微调燃料,这是比传统人工打标更真实的信号源。但它的价值不取决于噱头,而在于两条生命线:其一,如何确保“不断学习”不变成“含错训练”。评论中反复提及的专家冲突、用户噪音、过期知识,是任何持续学习系统的阿喀琉斯之踵。平台所谓的“评估层”听起来能过滤,但真正落地需要将业务标准量化为稳定的评判机制——这通常比微调本身更难,且不同行业的成熟度天差地别。其二,“模型所有权”概念诱人,但必须警惕架构风险。团队回应“你拥有的是数据而非权重”很诚实,用户最终资产是带标签的领域知识,而非一个可直接运行的生产模型。当基础模型迭代(如从Opus 4.x到5.y),用旧数据在新基座上重训的成本和成功率,会直接打破“你拥有的东西永远有效”的叙事。专业买家应关注的是:平台与基础模型供应商的去耦合程度有多深?数据迁移的标准化接口在哪?模型效果提升有多少不能被单纯归因于基础模型更新?目前,Empromptu AI 解决了一个真实且痛苦的问题——让领域知识持续内化至AI,但它更像一个**高效的数据飞轮引擎**,而非一个“买断式”的终极AI。衡量它成败的,最终不是它创造了多少模型,而是它能否帮用户创造出能独立“生存”的领域数据资产。对于非技术驱动的团队,它的低门槛是巨大红利;对于追求深度品控的机构,建议先做好“谁能参与判断定义”的治理框架,否则飞轮转得越快,可能偏得越远。

查看原始信息
Empromptu AI
Most AI apps launch on someone else’s model and stay there forever. Empromptu AI turns live AI features into custom models you own. As your app runs, Empromptu AI captures real-world usage, human corrections, and edge cases from live AI workflows, then uses that signal to train a custom model you own. Improve accuracy, lower inference costs, and stop depending forever on rented intelligence from the same providers moving into your category.

I’m Shanea, co-founder and CEO of @Empromptu AI


We built Empromptu AI's Alchemy because we believe the next phase of AI is not just building apps faster.
It is building AI that learns how your business works.

Right now, everyone is rushing to learn or add AI whether you are someone trying to figure out how to survive as an employee or take your expertise and monetize it. Everyone is plugging into the same frontier models, shipping the same generic workflows, and calling it a moat. But if everyone is using the same intelligence, no one is differentiated for long.

We're changing that.


With Empromptu AI's Alchemy you can fine tune a model with no ml expertise simply by building an AI applications. Our platform automatically captures customer usage, your corrections as a subject matter expert, edge cases, and application feedback. Alchemy turns those signals into a fine tune model that can keep itself up to date. Yes a self learning, self improving AI that doesn't cost trillions.

The simple version:
You've spent the last 5-10 years at your job learning really valuable skills whether its engineering, content, or more insane a highly regulated or specialized field.


Now your AI can learn from you so you can own your expertise.

This matters because the best knowledge usually lives inside people’s heads. The accountant knows the exception. The support lead knows when an escalation is real. The operator knows the edge case. The product team knows what “good” actually looks like.

Alchemy gives everyone a way to turn that expertise into AI that gets self improves with up to 98% accuracy.

Thank you for checking us out today. I’d love your feedback, questions, and brutal honesty.

36
回复

@shanealeven Congrats on the launch team. How do you stop bad user feedback, noisy corrections, or outdated domain assumptions from being absorbed into the tuning loop

9
回复
Does Empromptu work with all frontier models and open source/weights models? What happens to the trained/tuned model in the long term if the frontier model significantly advances? Say from Opus 4.x to 5.y …
19
回复

@lakshminath_dondeti you can only fine tune oss models. The frontier models are deprecating the ability for you to fine tune them. 😬 Which is one of the reasons we're making this available. But Empromptu is totally model agnostic

14
回复

@lakshminath_dondeti To add some architecture context, the fine tuned model you build on Alchemy is your asset regardless of what happens upstream with frontier models. When a new frontier model drops, you're not starting over. Your domain knowledge, your corrections, your edge cases are portable. We can use a newer base and retrain on your existing data. The expertise your team captured doesn't deprecate with the model version.

15
回复

@lakshminath_dondeti Good questions. On the first, the orchestration layer is model-agnostic by design, the point is that you route to whatever's best for the job rather than getting locked to one provider, frontier or open-weights. On the second, that's the more interesting one. Because what you actually own is the labeled data and corrections from your SMEs, not just one frozen checkpoint, a base model leap from 4.x to 5.y isn't a reset. Your accumulated edge-case data is the durable asset, and you re-tune the new base on it. The frontier advancing is a tailwind, not a stranding event. That's a big part of why we frame the data as the thing you own rather than the weights alone.

1
回复
Impressive vision! Turning real-world AI usage, human feedback, and edge cases into custom models that continuously improve is a compelling approach. The focus on ownership, accuracy, and reducing long-term dependency on external models really stands out. Excited to see how Empromptu AI helps teams build AI products that get smarter over time. Congratulations on the launch! 🚀
16
回复

@1mirul Thanks so much! You picked up on the part we care about most, the ownership piece. So much of the AI stack right now is rented, and we wanted teams to actually keep the thing that gets smarter. Really appreciate you taking the time to look. 🚀

1
回复

@1mirul thanks so much. If you own your asset you should be able to decide what you do with it whether you compete or whether you decide to sell that asset but you and everyone else should be able to capitalize on the data you own. Your data is getting scrapped and captured anyway. You should at least be compensated

8
回复

@1mirul The compounding part is what makes it structurally different. Most AI deployments get smarter for the vendor. This one gets smarter for you. That asymmetry is the whole point.

9
回复

Congrats on the launch 🎉
Curious, when the dynamic prompt optimization kicks in after 30 runs, does the user get visibility into what changed, or does it just happen silently in the background?

14
回复

@boyuan_deng1 The user absolutely gets visibility into what runs. You can also do this manually as well for all the tinkerers out there.

12
回复

@boyuan_deng1 Thanks! Good question. Visibility is something we care about a lot, the whole point is that improvements are yours and shouldn't be a black box, so transparency into what changed is the direction we lean. Want to make sure I give you the exact behavior rather than hand-waving though, so let me grab the precise detail and follow up. What's your use case? Curious what you'd want to see surfaced.

1
回复

@boyuan_deng1 And the visibility is intentional from an architecture standpoint. A model you can't inspect is one you can't trust in production. You should always know what signal drove a change before it goes into training.

9
回复

Instrumenting production app usage as a fine-tuning data source is genuinely clever. You avoid the cold start problem of manually curating datasets that don't reflect real user behavior. We hit that exact wall building our AI features and ended up with synthetic data that didn't generalize well. What does your quality filtering pipeline look like between raw app interactions and the training checkpoint?

14
回复

@retain_dev Actually thats kinda the best part if I do say so myself Dr Sean Robinson my cofounder has a PhD is computational astrophysics and he invented a way to get up to 98% accurate outputs out of any model! Its built in to happen completely automatically based on the eval that you write and that last mile is what you label. Thanks for the complement. We know the problem is unless youre a founder, the smes and ai/ml eng are usually separate so you never really get that perfect dataset.

13
回复

@retain_dev To add the engineering layer to what Shanea described, the quality filtering is what makes the self-improving loop actually work in practice rather than in theory.

The eval you write upfront becomes the ground truth signal. Every production interaction gets scored against it automatically. What surfaces for labeling isn't a random sample of outputs, it's specifically the cases that fell outside your accuracy threshold. You're not reviewing everything, you're reviewing the exact delta between what the model did and what your domain requires.

That's why the labeled dataset stays small and high signal over time. The model improves, the eval catches the new edge cases, and you're always training on the real distribution of your own users rather than synthetic approximations.

The part that resonates with what you described about SMEs and ML eng being separate is exactly the failure mode we designed around. The eval layer is built to be owned by the domain expert, not the ML team. The person who knows what correct looks like is the one defining the signal, not translating it through an eng team.

13
回复

@retain_dev You hit the exact wall we built around. Synthetic data that doesn't generalize is the thing real usage fixes, but raw interactions are noisy too, so the filtering matters as much as the capture. The short version: not every interaction becomes training data. The signal comes from where an SME actually touches a case, a correction, a confirmation, an edge case getting flagged, rather than from raw traffic. So the human-in-the-loop step doubles as the first quality filter, and the governance layer is where conflicting or low-confidence labels get resolved before anything hits a checkpoint. Curious what your synthetic approach was, since that cold-start pain is real.

1
回复

The 'bring your own expertise' angle is the right way to think about the next wave of AI. We struggle constantly with customer support AI missing the nuance of our specific software updates. If this plugs directly into customer usage signals to self-improve, it solves a massive operational headache. Amazing job @shanea_leven

14
回复

@shanea_leven  @priya_kushwaha1 That support nuance problem is exactly the kind of thing this is built for. Your team already knows where the AI gets it wrong on your specific updates, and that knowledge is usually the missing piece. Plugging it into real usage signals so it corrects itself over time is the whole idea. Would love to hear how it goes if you try it.

1
回复

@priya_kushwaha1 we agree this power is now really accessible for anyone to be able to fine-tune. And yes fine tuning with your own data does absolutely increase accuracy

10
回复

@shanea_leven  @priya_kushwaha1 And the part that usually surprises people is how small that labeled dataset needs to be when the signal is right. Real production corrections from your own users generalize far better than anything synthetic at 10x the volume.

11
回复

@shanea_leven Congratulations on the launch.

One thing I’m trying to understand from your positioning. If the underlying model providers keep improving rapidly every few months, how do you measure whether the gains your customers see are actually coming from Empromptu’s learning layer versus improvements in the foundation model itself?

It seems like that’s a pretty important distinction because both could lead to better outputs over time, but only one creates a real competitive advantage for the customer. Are you able to quantify that difference in a meaningful way?

13
回复

@shanea_leven  @moh_codokiai yes absolutely you will actually see the performance and accuracy improvements directly in the product. And frontier models have deprecated the ability for you to fine tune their models any more. Also a model trained on your data for your product and your users is always going to be eventually more accurate than a general model.

14
回复

@shanea_leven  @moh_codokiai 

This is the sharpest question in the thread so far!

You're right that both lift outputs but only one is durable, so we care about isolating it too.

The cleanest way to separate the two is to hold the base model constant and measure the delta your own data adds on top of it, the same base, with and without your SMEs' corrections, evaluated on your edge cases. That delta is the part that's actually yours.

When the foundation model jumps a generation, you get that lift for free like everyone else, but you also carry your accumulated data forward and re-measure the delta on the new base. The foundation improving raises the floor for the whole market. Your data is what raises your ceiling above it.

Quantifying that gap precisely on a given customer's eval set is something I'd rather walk through with real numbers than hand-wave, happy to go deeper if you want. As one example, though (and admittedly anecdotally, based on a single launch customer's internal testing), they saw an improvement of ~30% accuracy from training on the first run.

1
回复

@shanea_leven  @moh_codokiai First, we're model agnostic -- the user can specify whatever they'd like to use, and we'll adapt the 'baseline' accordingly; the difference is often, for our users, the other efficiencies, such as the training, vertical integrations and other optimization components that, combined, make an enormous difference in both the quality of life they experience while building (setting up actual databases, auth sequencing, et cetera, with real-world best practices), and doing all of the AI-focused functionality from a template-based system that let's me say something like: "I want to build a growth function, and that's going to be me following up with leads we haven't talked to in more than 3 weeks, and I want that sort of outreach to look like this."

I haven't met a business leader yet who wants to replace someone on their team with a system that does that at 98% accuracy. And that's the difference they walk away remembering.

3
回复
Love seeing tools that bridge the gap between AI experimentation and real-world deployment. This feels built for teams that are serious about shipping AI products.
13
回复

@tanjum Thank you! That's exactly the team we built it for, the ones past the demo stage who actually have to put this in front of customers. Appreciate you taking a look.

1
回复

@tanjum thank you so much yes. We always try to make everything we do accessible.

8
回复

@tanjum Two years of production deployments across healthcare, retail, and financial workflows is what shaped the architecture. The edge cases you only hit in production are exactly what we built around.

8
回复

I like the part abt capturing corrections and edge cases from real usage. That feels more useful than trying to guess everything upfront. One thing I wonder, how do you keep the model from leaaning the wrong patterns when user feedback is inconsistent or when diff experts correct the same situation in diff ways?

13
回复

@busra_seker1 Good question, and it's one of the harder parts. The short answer is we don't treat every correction as automatically true. The governance layer is there partly for this. Conflicting corrections on the same kind of case get surfaced rather than silently averaged into the model, so a human can resolve which pattern is actually right instead of letting noisy feedback train drift in.

The other piece is that expert disagreement is usually signal, not just noise. When two SMEs handle the same situation differently, that often means the policy itself is ambiguous, and catching that early is more valuable than quietly picking one. So we'd rather expose the conflict than bury it. So, the way our system works, you're actually defining what makes your model unique when you help to train it by properly contextualizing these edge cases, and our system has ways to identify good/bad data independently!

Hope you try it out and enjoy it!

1
回复

@busra_seker1 that's a great question. There is one ground truth so SMEs can override customers but SMEs have to agree what is ground truth

10
回复

@busra_seker1 Exactly right on the ground truth architecture. On the inconsistency problem specifically, that's where the eval becomes the arbitration layer. Conflicting corrections don't both make it through, the eval scores against a defined expected outcome so noise and contradictions get filtered before they touch training. The model learns from signal that passed a quality bar, not raw feedback volume.

10
回复
Love the name Alchemy, it makes sense to own the outputs and control your own data.
9
回复

@hellovidya thanks! We all know that data in today's world is gold! We feel that everyone should be able to turn data into gold 😁

8
回复

@hellovidya Thank you! The name felt right, turning your everyday usage into something valuable that's actually yours. Glad it resonates. 🙌

1
回复

@hellovidya Thank you!

6
回复

This is awesome Shanea! Wish you all the best on this impressive launch

9
回复

@german_merlo1 Thank you so much. Excited to get this out to the world.

7
回复

@german_merlo1 Thank you Germán!

8
回复

@german_merlo1 Thank you so much! Really appreciate it. 🙏

1
回复

Can you walk us through what AI apps you're already building means in practice are you integrating with specific frameworks like LangChain or LlamaIndex?

8
回复

@ana_popescu2 Good question. In practice it means you're not starting from a blank canvas, you pick the function you want automated and the template drops you onto the orchestration layer that handles routing, context, evaluation, governance, and monitoring underneath, so "building an AI app" is really configuring that for your use case rather than wiring it all yourself. On the framework question, let me give you a precise answer rather than hand-wave the integration story, what are you currently running? Curious whether you're coming from a LangChain/LlamaIndex setup or starting fresh, since that shapes the answer.

In general, we want to abstract away complexity, so for example, you could take a git repo from most other projects (built in Lovable, Claude Code, etc) and be able to transition any component of what you're building (that uses AI/needs what we offer) to help improve accuracy through edge case training automatically over time.

1
回复

@ana_popescu2 no. We built our own proprietary frameworks specifically designed to get highly accurate AI outputs. So many of our users build their entire products on our platform like we have had users build Soc 2 platforms, AI tax products, full healthcare products for their communities. We have a grant program winner who is building AI workflows for airlines! This is sophisticated tools that ml and ai engineers use made accessible for everyone

8
回复

@ana_popescu2 To double click on Shanea's point, the proprietary architecture is what makes those wildly different use cases possible on the same platform. The fine tuning loop needs to close reliably whether you're building a healthcare product or an airline workflow, and that's a very different design requirement than general purpose orchestration frameworks are built for.

6
回复

It’s absolutely amazing everything you can build with Empromptu! Custom models are the future — own my data, better accuracy, and cheaper!?

8
回复

@joshua_leven Yes we think it's pretty wild too. We think this gap is the missing link for AI to really take off.

9
回复

@joshua_leven And the cost curve is the part that surprises people most. A smaller model trained on your domain consistently outperforms a general frontier model on your specific tasks at a fraction of the inference cost. You get more accurate and cheaper at the same time!

8
回复

@joshua_leven That's exactly the pitch! Own the data, get accuracy on your actual edge cases, and stop renting it all forever. Glad it clicks for you. 🙌

1
回复

the feedback loop approach is smart. the part that usually trips teams up isnt the training pipeline though, its the quality of the corrections feeding it. if the humans correcting the AI output dont have a systematic way to evaluate whats actually wrong you end up fine-tuning on noise. curious how you handle that signal quality problem

7
回复

@ozandag The correction itself is not the signal, the correction scored against a defined expected outcome is the signal. Without that layer you are just fine tuning on whoever had an opinion that day.

The eval is defined upfront by the SMEs who know what correct looks like. Every correction gets scored against it before it touches training, filtered if it doesn't meet the bar, flagged for review if it's ambiguous. When experts disagree the eval arbitrates. You never train on conflicting signal, only verified ground truth.

4
回复

@ozandag +1 to what Sean said. We allow the user to first define what good looks like. and we actually remove the hard parts so they can define it in natural language. Often with a single statement. No configs or files so anyone can do it. Then we measure accuracy towards that goal.

3
回复

@ozandag The pipeline's the easy half. We score corrections before they hit fine-tuning, so a sloppy "this is wrong" doesn't weigh the same as a structured one. Bad corrections are their own failure mode and most teams don't instrument for it. What surfaced this for you?

4
回复

I keep thinking about how much institutional knowledge disappears when someone leaves a company. Most organizations have years of expertise locked inside conversations, corrections, and unwritten rules. The idea of turning those signals into a continuously improving system feels like a much bigger opportunity than no-code app building itself.

7
回复

@mehmet_s_taskesen You're pointing at something we think is the bigger story too. Most companies have years of hard-won judgment living in people's heads and in one-off corrections that never get written down, and when someone leaves, a lot of it walks out the door. Capturing that as your experts work, and turning it into something that compounds instead of evaporates, is honestly the part that excites us most. The app building is the on-ramp. The knowledge becoming a durable asset is the point.

1
回复

@mehmet_s_taskesen This is exactly the framing that drove the original architecture decision. The problem was never that foundation models were bad. It was that every organization was essentially starting from zero every time because there was no infrastructure to capture what their best people actually knew.

The accountant who knows the exception. The support lead who knows when an escalation is real. That knowledge compounds inside a person over years and then walks out the door. Alchemy is fundamentally an infrastructure problem solved, not an AI feature added. The signal was always there in the corrections, the edge cases, the judgment calls. It just had nowhere to land permanently.

The no code angle is actually secondary to us. What we are really building is the layer that turns institutional knowledge into a durable asset that survives the people who created it.

5
回复

@mehmet_s_taskesen we totally and completely agree. Better yet. What if you could monetize it as well?

3
回复

Fine tuning from live app behavior raises an interesting data quality challenge what mechanisms does Empromptu use to filter out bad or edge case interactions that could quietly degrade model performance over time?

7
回复

@antonio_manuel1 The key is that not every interaction becomes training data. Raw app traffic is noisy, so the signal comes from where an SME actually touches a case, a correction or a confirmation, rather than from volume. That human-in-the-loop step is the first quality filter. From there the governance layer is where low-confidence or conflicting labels get resolved before anything reaches a training checkpoint, so a noisy or genuinely edge-case interaction gets surfaced rather than silently absorbed. The goal is that bad data gets caught at the gate instead of quietly degrading the model down the line.

1
回复

@antonio_manuel1 this is really a New paradigm. Subject matter experts can self-correct data to always ensure performance is high for the first time alongside a tool that feels like you're simply vibe coding

5
回复

@antonio_manuel1 This is exactly the right question to ask and honestly one of the harder engineering problems we solved. Raw production data is noisy by nature so we never let it flow directly into training.

Every interaction gets scored against the eval you defined upfront. That eval is your ground truth and anything that doesn't meet the quality threshold gets filtered before it touches the training pipeline. Edge cases don't get discarded though, they get flagged for SME review because edge cases are often where the most valuable signal lives. The difference is they go through a human in the loop step before they become training data.

The other layer is the SME override architecture Shanea mentioned earlier. When experts disagree on a correction, the eval arbitrates. You're never training on conflicting signal, you're training on verified ground truth.

The result is a labeled dataset that stays small and clean over time rather than growing noisy with volume.

6
回复

The idea of turning your own app's usage into fine tuning data is genuinely clever but how do you handle the cold start problem for teams whose apps don't yet have enough interaction volume to generate meaningful training signal?

7
回复

@andrew_paul11 that's the beautiful thing about the app creating the data. The very first time you run the app it creates data!

6
回复

@andrew_paul11  The short answer is you don't need volume to start, because the signal isn't raw traffic, it's SME corrections on the cases that matter. A handful of expert-labeled edge cases is worth more than thousands of unlabeled interactions, so even a low-volume app can start accruing useful training signal early. It also helps that you're picking a specific function to automate rather than boiling the ocean, so the interactions you do have are concentrated on the cases you actually care about getting right. Volume helps over time, but it's not the gate to getting value on day one.

This is also why we say things like: your 'edge cases actually define what makes your model unique and uniquely valuable' when we talk to prospects because the nuances aren't in the data volume, but the training itself.

1
回复

@andrew_paul11 The part that makes that work under the hood is that even a small number of high quality runs beats a large volume of synthetic data. Because the signal comes from real interactions against your actual eval, even early data is already shaped around your domain. You're not waiting for volume, you're waiting for signal and those are very different thresholds.

6
回复

@jordan_hanson1 congrats guys, about the 98% figure. Is that something customers typically achieve after the model has learned from their application data, or is that the starting point?

It would be interesting to understand what the baseline was before the learning process.

7
回复

@jordan_hanson1  @xavair 

That's an after-learning number, the gain that shows up once the model has been corrected on a customer's own edge cases, not a starting point. The baseline is the more interesting part of the story honestly, since the whole value is in the delta between where it starts and where it lands once your SMEs have been in the loop. I don't want to quote you a baseline off the cuff since it varies by use case, but I'm happy to walk through a real before/after if you want to see the shape of it.

the pre-training/learning baseline sort of varies depending on a few things, like (a) Did you DIY in your product and are connecting to us after the fact, or (b) did you start building FROM Empromptu initially? The reason this matters is that we have a significant 'impact' on outcomes in both scenarios, but we obviously control more of the 'development lifecycle' in the (b) case, and so the immediate results may be less dramatic (because it starts with the benefit of the vertically integrated optimizations our platform offers).

1
回复

@jordan_hanson1  @xavair yes! It is. Whats unique about our platform is that we can agenticly get really high accuracy rates as we built a custom model ourselves to make corrections in real time dynamically. But the last percent you can correct responses. So because these tools are integrated into one platform you really get very high accuracy rates

7
回复

@jordan_hanson1  @xavair Building on what Shanea said, the 98% is what the system converges toward over time, not where it starts. The baseline depends on how far the foundation model is from your specific domain out of the box. The more specialized your use case, the bigger the delta and honestly the more dramatic the improvement curve once your own production data starts feeding the loop.

6
回复

How does Empromptu approach the tricky intersection of user privacy and training data collection specifically how do you help developers stay compliant when end users haven't explicitly consented to having their interactions used for model training?

6
回复

@new_user___10520260379921a76fc2d64 Great question!! We've actually built in a data anonymizer! to randomize any PII before it goes into model training.

3
回复

@new_user___10520260379921a76fc2d64 we have several mechanisms to securely protect end-user data, and also offer the ability to mechanically substitute the 'data' from the 'identifying components' through 'john doe' substitutes that makes PII attribution impossible, especially in training applications.

3
回复

@new_user___10520260379921a76fc2d64 First, our systems are built with HIPAA compliance in mind, which means there are built-in safeguards regulating access to and use of all data that belongs to a customer, but even within that, stricter controls on the unification and association of medical records with personal identifiers, structurally preventing these sorts of issues.

We work with customers in healthcare and would be happy to share specific examples if you or your team would like to meet and discuss how your particular function sets would work?

1
回复

Continuous fine tuning from live data sounds powerful but it also risks model drift over time how does Empromptu protect against a model that gradually shifts away from its intended behavior as usage patterns evolve?

6
回复

@daniel_juan2 The protection is that learning isn't continuous in the sense of the model silently shifting on its own, updates happen in controlled cycles your team reviews and promotes, so the model doesn't drift away from intended behavior without someone seeing it.

Each cycle gets evaluated against your edge cases before it's promoted, so you're checking that an update actually improves behavior rather than quietly changing it. And because you own your checkpoints, if a cycle does move in the wrong direction you can roll back.

The goal is the model evolving deliberately on your terms, not drifting out from under you.

1
回复

@daniel_juan2 The eval is the anchor. No matter how much production data flows through, the model can only update in directions that pass the ground truth your SMEs defined upfront. Usage patterns evolve but correct stays fixed until you deliberately change it.

The versioned checkpoint architecture handles the rest. Every training cycle is inspectable and fully reversible so drift never accumulates silently. Healthcare and financial workflows shaped these requirements from day one, silent drift was never an option.

3
回复

@daniel_juan2 automatic drift detection for the win! We think about all of these things on a daily basis

3
回复

Woo, love seeing this ship! Already mulling through some of the fun stuff I could add to my companies in terms of being able to fine tune some models, hah.

6
回复

@holman Woo, thanks! Love that you're already scheming. The fine-tuning side is honestly the fun part once it clicks. If you end up building something, let me know how it goes, always curious what people reach for first.

1
回复

@holman amazing! Would be super curious to see what you build!!

5
回复

@holman let us know how we can support you!

5
回复

The positioning of apps you're already building is really compelling what does the actual developer integration look like and how invasive is the instrumentation required to start capturing usable training data?

6
回复

@chen_hao3 The short version is the instrumentation isn't a heavy separate layer you bolt on, the capture happens through the orchestration layer your app is already running on, so the signal comes from the work itself rather than from you standing up a parallel data pipeline.

The point of building on the platform is that instrumentation isn't a second project, it's a property of running there.

On the exact integration surface, I'd rather give you a precise answer than hand-wave it, what's your stack look like? Happy to walk through what getting started actually requires for your setup; feel free to book a meeting with our team!

1
回复

@chen_hao3 To make that concrete, the instrumentation is intentionally lightweight. You are building on Empromptu's platform so the capture layer is built in, not bolted on. There is no separate SDK to integrate, no custom logging pipeline to stand up, no data pipeline to maintain.

The signals Alchemy needs, corrections, edge cases, application feedback, are byproducts of normal app usage on the platform. The developer experience is just building the app. The training data infrastructure runs underneath it automatically.

That was a deliberate architecture decision. If instrumentation requires meaningful engineering work it becomes a project that competes with shipping. We needed it to be invisible so teams could focus on the application layer and let the learning layer take care of itself.

4
回复

@chen_hao3 said in another way it feels like you're vibe coding but then you start to notice all of the things devs need are there

5
回复

Most fine tuning tools treat evaluation as an afterthought does Empromptu have a built in framework for measuring whether a fine tuned model is actually outperforming the base model in production not just on a held out test set?

6
回复

@carter_son yes! absolutely. this is how we were able to help one of our Alchemy pilot customers measure and contextualize a 30% improvement in accuracy n their very first training run! It was absolutely incredible and even we didn't expect such a dramatic result.

If you'd like, we're always free for calls to walk people through how Alchemy does this, specific to your workflow or use case.

1
回复

@carter_son Worth addressing directly because you're pointing at a real gap in how most fine tuning tools are built. Held out test set performance is a proxy metric. What actually matters is whether the model makes better decisions in your specific production environment and those are often very different things.

The eval you define in Alchemy is not a one time benchmark, it runs continuously against live production interactions. So you are always measuring against real user behavior, real edge cases, real domain conditions rather than a static slice of data you curated before deployment. The comparison to the base model is ongoing not a one time checkpoint.

The other piece is that because the eval is domain specific and defined by your SMEs, the performance measurement is actually meaningful to your business. You are not optimizing for a generic accuracy score, you are optimizing for the exact outcomes your team defined as correct.

5
回复

@carter_son yes you can actually see this live in our dashboard!

2
回复


the self improving AI angle is really interesting. how do you balance continuous learning with maintaining model stability and consistency? can customers roll back changes if needed?


6
回复

@easton_carter Continuous learning is only useful if it doesn't make the model unpredictable, so the answer isn't "learn constantly," it's learning in controlled cycles your team reviews rather than a model quietly shifting under you.

That review gate is what keeps stability and consistency intact, you decide when an update is good enough to promote.

On rollback, yes, that matters for the same reason, if an update doesn't behave the way you expected you need to be able to go back, and since what you own is your data and your checkpoints rather than a single frozen black box, reverting is part of staying in control.

The whole point is the model improving on your terms, not despite you.

1
回复

@easton_carter From the technical side, continuous learning and stability are actually in tension by design so we had to solve for both explicitly. The eval is what keeps the model from drifting. Every update gets scored against the same ground truth before it ships so the model can only improve in directions your domain actually validates.

On rollback, yes. Every training checkpoint is versioned so if a learning cycle produces something unexpected you can revert to a previous state. You are never locked into a bad update. The combination of eval gating on the way in and versioned checkpoints on the way out is what makes continuous learning safe enough to run in production environments where consistency actually matters.

5
回复

@easton_carter yes automatic drift detection for the win!

4
回复

One concern enterprises always raise around fine tuning pipelines is data residency can you speak to how Empromptu isolates customer training data and whether it ever touches shared infrastructure?

6
回复

@carlos_leonardo1 The principle is that your training data and the models it produces are yours and governed by the same layer that produces them, isolation and residency are core to that, not an afterthought. But I don't want to hand-wave the specifics on something this important. If you want, I can get you our actual data handling and residency documentation, or connect you with someone on our side who can walk through the architecture properly for your requirements.

Please feel free to book a meeting and we'll walk you through our security protocols!

1
回复

@carlos_leonardo1 Data isolation is non negotiable for the customer segments we serve. Healthcare organizations and financial workflows were in production on Empromptu before we ever talked to a general market so the architecture was built around those requirements from day one, not retrofitted later.

Every customer's training data is fully isolated. Your corrections, your edge cases, your labeled dataset never touch shared infrastructure and they never inform another customer's model. The model you train is trained exclusively on your data and the weights are yours.

The broader point is that data residency is actually core to the product thesis. The reason Alchemy exists is that your data should build your asset, not someone else's. Letting customer training data bleed across tenants would contradict the entire value proposition.

3
回复

@carlos_leonardo1 we take privacy and data very seriously

4
回复

Congratulations on the launch!
how does alchemy decide which user corrections are valuable enough to incorporate into future model updates?can teams review or approve those learning cycles before deployment?

6
回复

@joshua_cooper2 Building on the ground truth architecture we designed around, every correction gets scored against the eval before it ever touches training. The eval is what decides whether a correction is signal or noise, not volume, not recency, just whether it meets the quality bar your team defined.

On the review question, yes completely. Nothing deploys silently. Teams get full visibility into what corrections are queued, can approve or reject learning cycles manually, and can also trigger retraining on demand for the tinkerers who want that control. The autonomy is configurable depending on how much oversight your team wants in the loop.

4
回复

@joshua_cooper2 

Thank you! Two good questions.

On the first: not every correction carries equal weight, the ones that matter come from your SMEs touching a case that's genuinely ambiguous or wrong, and the governance layer is where low-confidence or conflicting signals get resolved before anything is treated as training-worthy.

On the second: yes, and that's deliberate. Teams stay in control of the learning cycles rather than having updates happen silently, the whole ownership idea falls apart if the model changes out from under you without review. Human approval before deployment is the model we lean toward, not autopilot. Your training judgment is what makes your alchemy model uniquely valuable, so of course we want to protect that for each of our customers, independently.

1
回复

@joshua_cooper2 yes absolutely we give teams full autonomy to correct responses. We actually encourage this as it increases accuracy

3
回复

As a tool in the 'Vibe Coding' space, how much control does a user retain over the underlying architecture? If the conversational builder creates the full stack, is there a way to export the code or infrastructure configurations, or are users locked into the Empromptu ecosystem?

1
回复
#4
Google Gemma 4 12B
Run multimodal AI locally with an encoder-free architecture
246
一句话介绍:Google Gemma 4 12B 是一款无需独立编码器即可在本地 16GB VRAM 上原生处理文本、图像和音频的多模态 AI 模型,解决了开发者依赖云 API、本地硬件受限且不想在编码栈上浪费内存的痛点。
Open Source Developer Tools GitHub
多模态大模型 本地推理 编码器无关架构 16GB VRAM 开源模型 端侧AI Google DeepMind LoRA适配 消费级硬件 Apache 2.0
用户评论摘要:用户普遍认可其16GB低门槛和编码器无关架构的价值,认为这为本地多模态应用提供了可能性。主要疑问集中在:与26B MoE版的具体性能差距?对LoRA微调及适配器方法是否友好?以及serverless部署LoRA的可行性。
AI 锐评

砍掉编码器这一步,Google DeepMind 玩得很聪明,也很大胆。Gemma 4 12B 真正的价值不在于它在基准测试上“接近”26B版——那是一种营销话术,实际任务差异可能很大——而在于它从根本上改变了本地AI开发的成本结构。长期以来,多模态等于“吃硬件”的认知被它推翻:16GB VRAM 就能跑通文本、视觉、音频三模态,意味着一个MacBook Pro或中端游戏本就能成为独立的AI工作站,这将对无数云端API服务商造成降维打击。

但冷静下来看,它并非万能。编码器无关架构固然优雅,但这种“原生融合”的视觉和音频理解能力,天然在细节特征提取上可能弱于专用编码器,比如高精度OCR、细粒度音频分类等任务上可能会打折扣。评论中反复提及的LoRA微调问题,恰恰是这种新架构的潜在坑——如果视觉嵌入和音频信号直接被压缩到token空间,那么标准的微调方法是否还能保证所有模态的适配效果,需要开发者亲测。此外,“16GB适用”本身意味着这是在性能和资源间做出的妥协,若想运行复杂智能体应用,实际体验大概率会有延迟瓶颈。

总的来说,它是本地AI开发者翘首以盼的“解放者”,但也是一把尚未淬火的剑。它为独立开发者、隐私敏感场景打开了大门,但性能取舍和微调生态尚待验证。别被“接近26B”的表述冲昏,把它当作一个低成本的、灵活的多模态基础玩具来评估,你会得到更实际的回报。

查看原始信息
Google Gemma 4 12B
Gemma 4 12B processes text, vision, and audio natively without separate encoders, running on 16GB VRAM. For developers building local agentic applications who need multimodal capability without cloud dependency.

Gemma 4 12B is Google DeepMind's latest open-source model that processes text, images, and audio natively on consumer hardware, running on just 16GB of VRAM.

Most multimodal models carry a hidden memory tax: separate encoder stacks for vision and audio that inflate overhead before a single token is generated. Gemma 4 12B removes the encoders entirely. Vision runs through a lightweight embedding module, audio is projected as raw signal directly into the token space, and the LLM backbone handles the rest.

The result is a model that benchmarks close to Google's larger 26B MoE variant while fitting comfortably on a consumer laptop.

Key capabilities include:

  • 🧠 Encoder-free architecture for native text, vision, and audio processing

  • 💻 Runs locally on 16GB VRAM or unified memory

  • 🤖 Reasoning performance nearing the 26B MoE Gemma model

  • ⚡ Multi-Token Prediction drafters for reduced local inference latency

  • 📦 Apache 2.0 license, available now on Hugging Face and Kaggle

  • 🛠️ Compatible with Ollama, LM Studio, llama.cpp, vLLM, and HF Transformers

It is built for ML engineers and AI developers building on-device or edge applications that need multimodal capability without a cloud API dependency. Download the weights on Hugging Face or Kaggle and start building today.

P.S. I hunt the latest and greatest launches in tech, SaaS and AI, follow to be notified @rohanrecommends

4
回复

16GB VRAM for text, vision, and audio without separate encoders is a big deal for anyone building local AI tools. most multimodal setups eat half your memory just loading the encoder stack before you even start doing anything useful. curious how this compares to gemma 26B in real-world tasks though... benchmarks say "close" but close can mean very different things depending on what you're actually building with it

3
回复

The 16GB requirement is probably the most interesting number here. It feels like we're getting closer to a point where indie developers can build genuinely useful multimodal products without depending on external APIs for every interaction. Excited to see what people build when privacy, offline access, and multimodal AI can coexist on consumer devices.

1
回复

I have been witing for something like this for quite a long time!

1
回复

Encoder-free multimodal architecture is a bold call. Folding visual understanding directly into the transformer avoids the latency of separate encoder forward passes and reduces memory significantly. We've been wrestling with the privacy tradeoffs of calling external APIs for customer data, so local inference changes that calculus entirely. How does this affect LoRA fine-tuning? Can you still do efficient adapter-based approaches across both modalities?

1
回复
The only one local model that fits to my MBP (36gb ram) and can be used in IDE
0
回复
serverless deployment for LoRa adapter? 🥹
0
回复
#5
Build Club Campus
Virtual AI School: Upskill in AI and Become Great at it Fast
194
一句话介绍:Build Club Campus 是一款免费的虚拟AI学校,通过游戏化、社区驱动和动手实践的方式,解决AI工具和知识更新过快导致传统课程过时的痛点,帮助用户快速掌握真实可用的AI技能。
Education Developer Tools Artificial Intelligence
AI教育 虚拟学校 游戏化学习 社区驱动 动手实践 技能认证 职业路径 免费工具 自学平台 AI素养
用户评论摘要:用户普遍认可其“动手+社区”的核心理念,点赞免费模式。主要疑问包括:如何持续更新内容以应对AI快速迭代?游戏化机制能否长期留住用户?是否面向非技术人群及老年人?开发者表示已考虑入门友好,并将扩展开发者课程。
AI 锐评

Build Club Campus抓住了当前AI教育市场一个真实且尖锐的痛点——课程的“保质期”太短。传统的录播课或认证体系在面对以周为单位进化的AI工具时,几乎瞬间失效。其“动态、动手、社区驱动”的定位,本质上是在试图从“知识库”转向“知识流”,让学习不再是消费静态内容,而是参与一个持续更新的实践社群。

这种思路很聪明,但也非常难做。游戏化(徽章、点数)和社区粘性是典型的“初夜甜蜜,长期费力”。评论中已有用户尖锐指出:“如何在最初的兴趣消退后,让用户持续回归?” 这是一个灵魂拷问。如果社区内容最终沦为对官方发布笔记的简单搬运,或者动手项目深度不够、流于表面,那么Campus与一个带有积分系统的Discord频道并无本质区别。目前产品刚发布,投票数194,评论多来自内部团队或友好用户,缺乏残酷的外部压力验证。

其真正的价值不在于“教”AI,而在于“驯化”AI更新的焦虑。它为用户提供了一个“组织感”:将混乱的AI信息流结构化、任务化,并给予社交认同(社区点赞、认证)。但长期来看,能否成功取决于两件事:第一,社区能否自发产生高质量、领先于官方的实践内容(而非官方内容分销);第二,认证在用人市场上的含金量(是否被企业认可为能力凭证,而非仅仅是“完成率”)。目前看来,这是一个极具潜力的MVP,但离“AI时代的Coursera”仍有鸿沟。警惕“为了上架而上架”,避免内容深度被碎片化稀释。

查看原始信息
Build Club Campus
Build Club Campus is a fun, gamified and community-driven virtual AI school for learning AI by building with it. Static courses get old fast in AI, so Campus helps you stay current through bite-sized courses, real projects, role-based use cases and community templates that evolve with the tools. Earn certifications in OpenAI, Claude, Copilot and more, and become great at AI fast - for work, startups or side hustles. It’s 100% free, as part of our mission to enable anyone to build with AI.

Hey Product Hunt 👋!

Annie here, co-founder of Build Club.

Today we’re launching Build Club Campus - a free virtual AI school for anyone who wants to learn AI by actually using it.

We built Campus because AI learning is moving too fast for static courses.

New models launch every week. Tools change overnight. Features get released, renamed, replaced or deprecated. The “best way” to use AI rarely stays the best way for long.

So learning AI can’t just be a library of old tutorials 👀.

It needs to be:

Hands-on, because you only understand AI by building with it
Dynamic, because the tools keep changing
Community-driven, because the fastest way to stay current is to learn with people experimenting in real time

That’s what Campus is.

We learnt this while scaling Build Club into one of the world’s largest AI communities - 70K+ members across 60+ cities.

People are jumping between YouTube tutorials, docs, Discord threads, LinkedIn posts and random tool updates. It’s exciting, but it’s also overwhelming.

Campus gives people one place to keep learning as AI changes.

With Campus, you can:

→ Take practical AI courses and projects
→ Learn tools like Manus, OpenAI, Claude, Copilot and more
→ Follow role-based pathways across sales, ops, GTM and more
→ Earn certifications
→ Track your progress
→ Learn with a community that is building, sharing and updating together

The belief is simple:

Building beats theory.
Community beats static content.
AI is best learnt by making something real. 🤝

We’re the team behind Build Club, Manus Academy, Solaris, partner academies, accelerators and enterprise AI upskilling programs.

Campus brings that same practical, project-based learning experience into a free AI school anyone can use ✨

Explore Campus: https://campus.buildclub.ai/
Join Solaris: https://solaris.buildclub.ai/
Manus Academy: https://academy.manus.im/

Question for you:

What’s the first AI skill you want to feel genuinely confident in?

Roasts, course requests and feedback are very welcome 🔥

With love,
The Build Club team - Annie, Kevin, Talin, Clinton, David and Andrew

9
回复

@annie_liao Congrats on the launch team. You made the point about the pace of models and tools moving very quickly. How do you keep content up to date to deal with that?

1
回复

@annie_liao Congrats team! Love the gamification of AI learning. We all need it

0
回复

@annie_liao Congrats!

0
回复

SUPER hyped about this! Hands on, accessible AI education for free. This is amazing :))

5
回复

@talin_roche lets go! so excited to see builders level up on this platform and get certified!

1
回复

@annie_liao great approach to this! Quick question - have you considered an offering for the older generation that might be intimidated by the whole topic?

3
回复

@tom_palmer_ux thank you! the platform is aimed to be beginner friendly and learning is easy to digest so older generations can also benefit!

1
回复

@annie_liao  @tom_palmer_ux Good question!

0
回复

Gamified and community-driven - two things that sound great in a pitch and are genuinely hard to sustain. Points and badges work for the first few weeks. What's the mechanic that keeps someone coming back after the initial streak breaks? Congrats on the launch!

2
回复

@jared_salois thank you appreciate this comment! AI develops daily so we see this as a hub that centralises staying up to date. Community is also a core sticking point we are betting big on

0
回复
What’s your pricing? Are these self paced?
2
回复

@lakshminath_dondeti its free! :)

0
回复

I understood this as AI training for non-technical people? I really like the idea of short AI courses, and I’d like to see something like that for developers. And in multiple languages at once (we currently have people working from 5 countries).

1
回复

@natalia_iankovych 100% - more in store for developers coming soon!

0
回复

Completely agree with the premise. AI changes too quickly for traditional courses to keep up. The best way to learn is still to build things. Congrats on the launch!

1
回复

@alina_tyslenok_ thanks Alina! It's a fast moving field! If there are any specific topics you would like to see let us know :)

0
回复

AI is one of the fastest moving technologies ever. Campus is much-needed product to help declutter the noise and leverage it well.

Super hyped!

1
回复

@kevinzhu 100% - glad to be on the journey with you buddy

1
回复
#6
AppWizzy
Rent a private VM with Codex to build production apps
181
一句话介绍:AppWizzy 提供一台预装 Codex 的私有云虚拟机,让开发者通过对话式AI直接在云端构建、运行并托管生产级Web应用,解决AI生成原型容易、但部署和运维真实生产环境困难的核心痛点。
Software Engineering Developer Tools Artificial Intelligence
AI编程 云端开发环境 私有VM 生产部署 代码托管 开发者工具 生产力工具 应用托管 后端支持 SaaS构建
用户评论摘要:用户认可其持久化工作区和环境统一的理念,并关注VM内版本回滚、状态恢复和上下文管理问题。多数提问围绕为何仅限Codex而非支持Claude,以及服务器地域和长会话的文件索引机制。有用户反馈初始生成过程无进度显示。
AI 锐评

AppWizzy的“环境即服务”模式,精准切中了“Vibe Coding”热潮中被严重低估的痛点:AI生成的代码“能跑”和“能上线”之间存在巨大的鸿沟。创始人Philip用16年开发经验提炼出的认知很清醒——绝大多数AI编程平台只解决了“产生代码”的前半段,却将“运维、部署、数据持久化”这些真正构成生产级应用的硬骨头甩给了用户。AppWizzy将开发agent、运行环境、数据库与托管服务全部打包进一台私有持久化VM中,本质上是用基础设施即代码(IaC)的思路,为AI编程套上了一层“生产级合规”的缰绳。

然而,风险同样显而易见。首先,强制绑定Codex(基于开源的考量务实但局限)意味着在模型迭代速度上可能落后于Claude或GPT-5。其次,“所有东西都在VM里”虽然减少了环境不一致,但也制造了新的单点故障和扩展瓶颈。当用户需要横向扩展或引入微服务架构时,这台VM的边界将成为新的枷锁。此外,产品目前仍在通过内部工具向独立产品过渡的初期阶段,用户反馈的“生成过程无进度提示”等细节问题,折射出在工程化打磨上的短板。

真正的核心考验在于:当用户规模上来后,AppWizzy能否在不破坏“零运维”用户体验的前提下,提供按需伸缩和细粒度资源控制?如果能做到这一点,它就有机会成为AI时代的Heroku;如果不能,它终将沦为又一款“好用的脚手架生成器”。

查看原始信息
AppWizzy
AppWizzy gives you a private VM with Codex installed where you build, run, and host production web apps by chatting with AI. Your code is yours, the workspace persists, and the app lives in the same environment where it was created. Pay only for AI usage, hosting days, and optional templates
Hi Product Hunt I am Philip Daineka, founder of AppWizzy. I have been building software for around 16 years and running Flatlogic, our software development company, for 13+ years. Over that time, we have delivered many client projects and spent hundreds of thousands of hours building real business software: SaaS products, CRMs, ERPs, admin panels, portals, internal tools, and custom web apps. One thing I have learned very clearly: Making a prototype is easy. Shipping and maintaining production software is the hard part. That's why I've always been skeptical of many vibe-coding platforms. They are impressive for quick demos, but I would not start a serious client project on something that only gives me a front-end preview, hides the infrastructure, or locks the app into a platform that is hard to control later. So we built AppWizzy. AppWizzy gives you a private dedicated VM with Codex installed, where you can build, run, and host production web apps by chatting with AI. The key difference is that your app does not just get generated and exported somewhere. It lives in the same cloud environment where it was built. You get a real development machine, real backend capability, database support, hosting, source code access, different environment, and much more control over the stack - basically, everything you can do when you do indeed have full control over your VM. AppWizzy started from our internal tooling at Flatlogic and has now become a standalone product. We are also working on vertical production-ready templates for different use cases, so users can start from a strong foundation instead of a blank page and ship faster from day one. For the Product Hunt launch, we are offering 20% off. Use promocode: APPWIZZY-LAUNCH2026 I would love feedback from founders, developers, agencies, and anyone who has tried vibe-coding tools but eventually hit the question: "Okay, nice demo... but how do I turn this into a real production app?" Thanks for checking out AppWizzy! Philip
34
回复

@okendoken very interesting

7
回复

@okendoken Clean idea . Keeping coding , execution , and hosting in the same persistent environment feels like a natural step for AI-driven dev workflows. Curious how you're handling state recovery and versioning inside the VM when multiple iterations of the app are generated.

7
回复

@okendoken It's really appreciable

1
回复

Cool project! I wanted to know, is there any reason why Codex was chosen over Claude? Why not both options?

10
回复

Hi @ashishkingdom ! Thanks for your question!

Because Codex cli is open-source, unlike Claude. We love being able to tweak, control, and understand how things work under the hood, and it also gives us a better path to support many different models in the future

5
回复

Really proud to be part of the team behind Appwizzy! 🚀

I have seen how much work, care and energy went into building this product, and it's amazing to finally share it with the community.

We created Appwizzy to make app creation faster and less overwhelming for founders and teams, and I'm excited to see how people use it.

Would love your feedback and thanks so much for the support!

8
回复

@erik_kalmykov 

Congratulations on the launch! 🚀 Appwizzy looks like a great way to simplify app creation. Wishing the team lots of success and excited to see what people build with it!

0
回复

Really happy to see AppWizzy live today.

The part I personally like most is that it is not just about generating a quick demo and then figuring out what to do next. You get a real workspace, your own code, a private VM, backend/database support, and a place where the app can actually keep running.

Curious to hear what people think, especially if you have tried building something serious with AI coding tools before.

7
回复

I've jumped between dozens of builders, always changing, never settling. AppWizzy feels different. It adapts to my projects, not the other way around. I can finally build exactly what I need, not just what fits a demo. Shape-shifting just got easier.

4
回复

@sidraarifali Thank you!

0
回复
Never thought I'd feel grateful for a dev tool. AppWizzy genuinely took stress off my team. Real care wrapped in a clean UI.
4
回复

@tanjum Thank you!

0
回复
Thought AppWizzy would be another over-promised vaporware. Turns out, annoyingly, it delivers exactly what it promises. Competent, clear, and reliable-infuriatingly useful. You've spoiled my cynicism.
4
回复

@1mirul thank you so much!

0
回复

Why only Codex? We usually use Claude Code. And second question: where are the hosting servers located?

3
回复

Hello @natalia_iankovych ,
Thank you for your questions!

We chose Codex because it's open-source, which gives us more flexibility, transparency, and control under the hood. It also makes it easier for us to support multiple models in the future.

As for hosting, our servers are currently located in the USA.

If you can share with us how you usually work with Claude Code it will be very interesting.

Thank you.

1
回复

Good luck

3
回复

Congrats to the @AppWizzy team - I look forward to seeing the solutions built with Appwizzy in the coming days and weeks. I'm curious to see what others are building. I built this site with @AppWizzy . Easy and high-quality tools.

2
回复

@kevin_mckeand1 Thank you Kevin! Excited to see what you and others ship in the next days and weeks. Send us what you built, we’d love to feature it.

1
回复

AppWizzy actually lets me build, not just assemble random UI blocks. When did software development tools forget this basic requirement?

2
回复

@monir_ Exactly. Somewhere along the way, "build software" became "drag blocks until it looks like software".

We’re trying to bring back the obvious requirement: real app, real code, real runtime, real ownership.

0
回复

AppWizzy's approach of co-locating the coding agent, runtime, and host in one VM is smart. Separating build environment from deployment creates an entire class of 'works on my machine' bugs in AI-generated apps. How does Codex handle long sessions where the agent needs to reference files it created hours earlier: does it keep a persistent context index, or re-scan the workspace on each request?

2
回复

@anand_thakkar1 Exactly. For us the VM/workspace is the durable memory: files, Git history, logs, configs, and build output. Codex can resume sessions, but we don’t rely on LLM context as a perfect long-term memory, the agent re-inspects the repo when needed.

0
回复

It gave me this:
--
Starting generation now — you can refine data fields, UI screens, and matching logic after the scaffold is created.
--
then it looks like the process has stalled, at least I see no visible progress any further

1
回复

Hello @egorka ,
Thank you for trying this and let your review!

As far as I see from our side, your app is working and everything should be ok.

You could contact with me via email: e.kalmykov@flatlogic.com at any time and we will help you with any questions.

Thank you!

0
回复
@okendoken The “works in the same environment it was built in” distinction is underrated. Every vibe-coding demo I’ve seen eventually hits the “okay but where does it actually run” wall. The co-located VM approach is a clean answer to that. Curious how you handle rollbacks - if the AI makes a destructive change mid-session, is that recoverable from within the same environment or does it require external Git discipline?
1
回复

@okendoken  @skyninety Yes, recoverable inside the same environment. We treat the VM as a versioned workspace, not just a runtime. The AI works through files and Git/checkpoints, so if it breaks something, we can inspect the diff and roll back from there. External Git is still useful, but rollback should not depend on the user having perfect engineering discipline.

1
回复

Cool project! I wanted to know, is there any reason why Codex was chosen over Claude? Why not both options?

0
回复
Finally found a platform that doesn't force me to follow arbitrary rules. Real environment, no cage, pure freedom. Why isn't every AI tool built like this? Is that too much to ask?
0
回复

@faisal_2420010 Honestly, same. Once an AI can work where the app actually lives, it's hard to go back.

Not "no rules ever" — just real files, real logs, real builds, and a way to roll back.

0
回复
#7
Deliveryman.ai
Cold email infrastructure on autopilot without Gsuite
166
一句话介绍:Deliveryman.ai 是一款自动化冷邮件基础设施管理工具,帮助用户免去手动配置域名、DNS、预热、验证等繁琐环节,无需Gsuite即可快速搭建并维护可交付的邮件发送系统。
Sales Email Marketing Marketing
冷邮件 邮件基础设施 邮件投递率 DNS配置自动化 邮箱预热 黑名单监控 邮件验证 SaaS工具 销售外联 自动化
用户评论摘要:用户普遍认可该工具解决了DNS配置、预热、验证等痛点。主要问题:域名休眠后如何自动恢复预热?邮件中途故障能否自动切换备用域名?目前需手动路由,用户期望自动failover。此外,有用户询问是否支持自带域名、如何防止垃圾箱,以及是否有智能建议功能(如黑名单后的操作指南)。
AI 锐评

Deliveryman.ai 切中了一个真实且足够痛的场景:冷邮件基础设施建设,堪称“销售团队的DNS噩梦终结者”。从产品来看,它并非简单的“又一个邮件发送工具”,而是把域名配置、预热、验证、黑名单监控、回复路由等一堆琐碎且专业性极强的事情打包成一条自动化流水线。

但值得警惕的是,这类产品最大的陷阱在于“自动化”与“黑盒”之间的平衡。用户评论中已经暴露了关键风险:当域名出现问题,系统尚不能自动故障切换至备用邮箱,需要用户手动介入。在冷邮件场景下,一旦投递率急降或域名被标记,哪怕几分钟的延迟都意味着大量线索信件的废掉。作为“基础设施”级别的产品,自动化带来便利的同时,用户必须清楚:你是在把“抗风险能力”交给一个自动化系统,而它在关键时刻仍然需要你手动救火。

另外,“不依赖Gsuite”是亮点,也是双刃剑。它降低了门槛,但对非Google系邮箱的兼容性和长期稳定性仍需市场检验。用户提到“健康分数”数据来源有限,这又是一个潜在脆弱点——如果无法准确感知发件域健康状况,自动化反而会加速风险蔓延。

总的来说,Deliveryman.ai 对初创团队和中小型销售团队有明确价值:省掉1-2周的配置时间,让非技术创始人能快速跑通外联流程。但它离“真正的自动托管”还有距离,更准确地说,它是一件“帮你把70%的脏活干完,剩下30%仍需要你盯着的工具”。建议用户把它的监控和告警功能用好,而不是完全当甩手掌柜。

查看原始信息
Deliveryman.ai
Deliveryman.ai automates the hardest parts of cold email infrastructure mailbox setup, DNS records, warmup, email verification, sending systems, blacklist monitoring, reply management, and routing. Instead of spending weeks setting up your own cold email infrastructure and fixing deliverability issues, you can launch faster, scale confidently, and focus on what actually drives revenue. No G-suite required. Warmup, email list verification, Blacklist monitoring, etc. all included in one tool.
Hey Product Hunt! 👋 We built deliveryman.ai because cold email infrastructure is absurdly broken. Setting up a serious outbound system means juggling domains, configuring SPF/DKIM/DMARC, warming up dozens of mailboxes, monitoring sender reputation, and stitching together 4 different tools before you've sent a single email. Deliveryman.ai automates all of that: Connect your domains, and we handle the technical setup, email warmup, lead list verifications, and ongoing deliverability management so your emails land in inboxes, not spam folders. We're especially curious to hear from anyone doing high-volume outbound: What's your current setup, and what breaks most often? Drop it below. We read everything.
13
回复

@junaidansari @aminmemon Congrats on the launch! I have been an early customer, but then the platform went into maintenance. Is it fully back now? Happy to resubscribe whenever my clients need it next. :)

4
回复

@junaidansari congrats on the launch Junaid. Can you expand on the "no gsuite"?

2
回复

Most founders setting up cold email for the first time spend their first week fighting DNS records, SPF misconfigurations, warmup sequences, and blacklist monitoring — before they've sent a single real email. DeliverymanAI's bet that you can hand over that entire infrastructure layer — sub-domain creation, dedicated IP pools, reputation management, reply routing, all automated — and focus entirely on the message and the list, is the right abstraction for a founder trying to move fast. For seed-stage founders doing their first real outbound motion, this is exactly the kind of tool they should find before they spend a week in DNS hell. Added Deliveryman.ai to SoftRankings under the seed-stage sales stack. @junaidansari — what's the most common infrastructure mistake founders make before they find Deliveryman?

0
回复

This would've saved me hours when I was building our outbound operation.

Can I use my existing domains or do I need new ones?

2
回复

@aksayyed Thank you! That's exactly the kind of problem we built Deliveryman AI to solve.

Yes, you can use your existing domains.

In fact, established domains perform better than brand-new ones because they've had more time to build trust and reputation.

Simply connect your domain to Deliveryman AI, and we'll handle mailbox creation, warmup, monitoring, and the rest of the infrastructure for you.

0
回复

Hey Product Hunt community 👋

cold email should be simple.

write a great offer.
send it to the right people.
get replies. grow revenue.

but that’s not how it works today :(

We've built something that helps you focus on the above & less with the setups without costing you a fortune.

With Deliveryman.ai
- You connect a domain, we handle the rest.
- we create inboxes for you at no cost. (fully automated)
- we set up SPF, DKIM, DMARC & other DNS setup. (fully automated)
- we build your sender reputation before you send. (fully automated)
- we manage your sending thresholds. (fully automated)
- we monitor deliverability & blacklists behind the scenes. (fully automated)
- we route all your lead replies to your preferred email. (fully automated)
- we manage unsubscribes, auto-classify replies, & keep you compliant. (fully automated)
- we remove bounced emails & negative replied leads from your list. (fully automated)
- we let you send at scale, without fear, without hacks, without chaos. (fully automated)

Would love to know your thoughts on how I can improve it or what's missing from it.

P.S. Not vibecoded or AI slop.

2
回复

Looks interesting, but why should I pay for this instead of setting everything up myself?

2
回复

@ria_179 Thank you :)

You can absolutely do it yourself.

You can manage domains, buy mailboxes, configure DNS, spend $$ on warmups, spend $$ on email list verification, and monitoring delivery yourself.

But if your time is better spent closing deals, Deliveryman.ai handles that infrastructure for you.

1
回复

Congrats on the launch @junaidansari @aminmemon ! this is good one, upvoted :)

Curious though - so this is handling the mailbox + warming up domain emails and then I will bring the leads list from say Apollo? Or that is handled as well?

And how are you making sure that my emails will still not go to spam?

1
回复

@aminmemon  @aiswarya_s Thank you for the support! 🙌

Yes, that's correct. Creating & handling the mailbox + warming up domain emails (+ verifying lead list emails + monitoring popular blacklists) is all taken care by Deliveryman AI.

Deliveryman AI handles the infrastructure side of cold email of domains, mailboxes, DNS setup, warmup, email verification, blacklist monitoring, and inbox management. You can bring your leads from Apollo or any other lead source and use your preferred sending workflow.


As for spam prevention, no platform can guarantee that emails will never land in spam. Deliverability depends on several factors including infrastructure quality, list quality, email content, sending behavior, domain age, and recipient engagement. Our motive is to give you the healthiest possible foundation by automating the technical side and following deliverability best practices, but campaign quality still plays a huge role.

Out of curiosity, what's your current outbound stack today?

0
回复

This looks useful. Does it handle domain warming automatically or is that a separate setup step?

1
回复

@dhiraj_patel5 Thank you!

Yes, the warmup is built into Deliveryman AI. Once your domain is connected and your mailboxes are created, the warmup starts automatically.

There's no need to configure warmup or buy a separate warmup tool or manage additional integrations. The entire process of warmup is done for you automatically.

0
回复

Awesome that you unify DNS, warmup, verification, blacklists. How do you stop the system from over-warming dormant domains, and what’s your playbook when a mailbox health score drops suddenly?

1
回复

@leventbuilds Thank you. Happy to know that you found our product useful. Great question.

For warmup, we gradually increase sending activity rather than aggressively ramping up volume. We also coordinate warmup behavior with actual campaign activity so mailboxes aren't unnecessarily overworked when they're already sending outreach.

As for mailbox or domain health, that's honestly one of the hardest problems in deliverability. Every email provider behaves differently, and there isn't a single reliable "health score" source. We've integrated with Google Postmaster to gather insights, but the data isn't always comprehensive or updated frequently. In fact, newer versions of Postmaster no longer expose some of the domain reputation signals that were previously available.

Rather than showing a fancy score that may not reflect reality, we're actively working on ways to combine multiple signals and build a more accurate picture of deliverability health. It's an area we're continuing to invest heavily in.

0
回复

Nice launch! Curious from a deliverability POV if a sending domain gets into trouble mid‑campaign, does @Deliveryman.ai auto‑failover to warmed backups, or do I need to step in and re-route things manually?

1
回复

@munis_abbas Thank you, Munis.

At this very moment, you will have to re-route things manually.

But, we will be building a simpler workflow in a couple of months that will make it possible with just a few clicks.


Normally, users don't keep their domains idle. All connected domains keep running campaigns all the time.

Maybe once the development of this new process is live, we can provide you with an option to auto add backup domains if the current domain gets into trouble.

P.S. If you have any feature requests, we have a public todo list as well: https://deliveryman.ai/todo/

0
回复

Anyone who has ever configured SPF, DKIM, and DMARC knows that sending the first email is somehow the hardest part 😄 Congrats on the launch!

1
回复

@alina_tyslenok_ Couldn't agree more.

The first email is the hardest one because of everything that needs to be setup before sending.
Our goal with Deliveryman AI is to make that setup process feel effortless.
Thank you for the support! ❤️

0
回复

We were planning to start cold outreach next month, so the timing is perfect.

Signed up a few minutes ago. Do you have any recommendations/guide to setup?

1
回复

Kudos for your launch! This looks promising, am signing up for this to give it a shot

1
回复

@iamanantgupta Thank you so much!

Excited to have you on board. If you run into any questions or have ideas for improvement, don't hesitate to reach out.

0
回复
@junaidansari Congrats on the launch! How does the warmup handle domains that go quiet for 30+ days mid-campaign (paused outreach, seasonal gaps)? Do you ramp them back up automatically or treat them as fresh?
0
回复

@skyninety Thanks. Regarding the question...

With Deliveryman AI, domains don't really go "cold" when you're not actively running campaigns.


We continuously balance campaign sending volume and warmup activity. As your campaign volume decreases, warmup volume automatically increases to maintain healthy sending patterns and positive engagement signals. When you resume sending campaigns, warmup volume is adjusted back down accordingly.

This helps ensure your domain maintains a consistent reputation with email providers, even during seasonal gaps or periods of low activity.

We also factor in metrics such as bounce rates, inbox health, and overall sending performance to dynamically adjust warmup behavior over time.

The result is that your domain remains healthy and ready to send, without needing to be treated as a fresh domain after a period of inactivity.

How are you currently handling warmup for domains when campaigns are paused for a few weeks or months?

1
回复

Cool! We currently use 5 separate services for this.

Does it provide recommendations as well? For example, if you get added to a blacklist, does the service automatically suggest best-practice actions, such as stopping email campaigns for 30 days?

0
回复

@natalia_iankovych Thanks! That's exactly the problem we're trying to solve.

Most teams end up stitching together multiple tools just to manage their email infrastructure.

When a domain is detected on a blacklist, Deliveryman.ai automatically pauses email campaigns when necessary (depending on the severity of the blacklist) to help protect your sender reputation.

We also provide step-by-step guides for removing domains from the relevant blacklists. Once the issue is resolved and the domain is delisted, you can safely resume your campaigns.

Our goal is to not only monitor deliverability issues but also help users take the right actions to recover quickly and keep their outreach running smoothly.

Out of curiosity, which 5 tools are you currently using to manage your email infrastructure?

0
回复
#8
Novus
Catch and fix usability issues automatically as you ship
129
一句话介绍:Novus通过自动埋点、实时监控和代码级修复,解决了团队高速迭代时“产品分析滞后、问题发现慢、修复周期长”的痛点,让产品体验优化与开发同步。
Analytics SaaS Artificial Intelligence
产品代理 自动化分析 代码埋点 可用性监控 异常检测 智能修复 AI开发者工具 PM效率工具 用户行为分析 自动化QA
用户评论摘要:用户高度肯定自动发现和修复问题的能力,尤其赞赏GitHub集成和信号降噪处理。核心建议包括:如何区分关键问题与噪音(已回复:基于代码库、配置和实时数据交叉分析);异常检测频率(每日检测,支持自动PR修复);以及是否覆盖生产环境异常(已确认支持)。整体反馈正面,无负面意见。
AI 锐评

Novus的定位非常精准:它试图解决AI加速开发后遗症的“最后一公里”。当Copilot让写代码变得廉价,代码量指数级增长,传统手动埋点和事后看仪表盘的方式彻底失效。Novus的破局点在于,它不是一个更酷的“分析平台”,而是一个自动闭环的“产品代理”。它将分析动作从“事后人工复盘”前置到“每次构建”,甚至主动生成PR修复。

从评论看,其“代码级源数据”+“自动修复”的范式比同类工具更具竞争力。但危险信号同样明显:产品高度依赖对代码库的深度理解和自动“断案”,这在高复杂度业务逻辑和非常规交互面前,误判和错误修复的风险不低。当前评论几乎全是合作方和内部员工,缺乏独立用户的负面挑战,其“信号/噪音”的平衡能力在非Pendo生态、非技术栈统一的场景下存疑。

真正的价值在于:它把“产品体验”从PM的主观判断和仪表盘的滞后指标,变成了一个可监控、可回滚、可修复的工程化对象。但Novus的成功前提是,它必须比团队本身更懂他们自己的业务意图,而不仅仅是代码语法。如果做不到这一点,“自动修复”将沦为比“不埋点”更可怕的错误放大器。

查看原始信息
Novus
Novus is the product agent built for teams that ship fast. Connect your codebase and Novus automatically instruments, analyzes, and improves your product — no manual setup. It monitors continuously, flags usability issues proactively, and delivers intelligence for engineers and PMs.

At almost every company I've been part of, product analytics has been the wedge between a good product and a great one.

The challenge is that most teams are flying partially blind. Critical user journeys aren't instrumented, key signals are missing, and engineering teams spend valuable time manually adding tracking, maintaining analytics, and digging through dashboards to understand what's happening.

Even when you identify a problem, someone still has to determine the root cause and implement the fix.

That's why we built Novus.

Novus connects to your codebase and automatically instruments your product as you ship. It collects the signals you didn't know were missing, continuously monitors for issues users are encountering, investigates what changed, and proposes fixes directly in code via pull requests for your review.

The vision isn't another analytics platform. It's an autonomous product teammate that closes the loop from instrumentation → insight → fix.

43
回复

Hi Builders!

I'm Valeria, Head of Community @ Mind the Product, working closely with the Novus team to help support the new wave of building. We've partnered with the Novus team to bring you our World Product Day Hackathon that kicked off two weeks ago: https://mindtheproduct.devpost.com/

Also use Novus daily to check on all things MTP. Excited to see what y'all think!

34
回复

@valeria_khokhlova If you want more info on Novus, you can also check out our landing page (it's pretty snazzy) at Novus.ai or ask a question here!

15
回复

We've been using Novus for a few months and it's changed how we think about analytics. Analytics is only as good as the story it tells, and if you're just looking at the numbers you can make them say almost anything. Novus grounds that story in reality by treating your codebase as a source of truth. This means that your analytics metrics now have the nuance and understanding of the product which makes insights exponentially more valuable and less prone to misinterpretation. This is a really novel and interesting approach and I feel like one of those things we will look back in a year and say "I can't imagine going back to the old way of doing it".


As an engineer, the GitHub integration is one of my favorite features. Novus watches every PR, warns us when a change might break our tagging, suggests new things worth tracking, and flags the UX impact of what we're shipping. I've been impressed with the signal to noise ratio and it has become a really valuable step in our toolchain, especially as AI keeps speeding up how fast we build.

30
回复

Fun confession from the build: when we first put up novus.ai to gauge interest, our sign-up forms weren't instrumented correctly. I had shipped them in a hurry and the events just weren't firing right. Then we installed Novus on the site. It flagged the broken instrumentation almost immediately, told me exactly what was wrong, and offered to fix it—opening a detailed PR for me to review and merge. The power of Novus instantly clicked for me, even for a marketing website.

25
回复

@ryan_cates1 That's the best kind of product story. The tool caught its own team's mistake before users did.

0
回复

I like the idea of catching usability issues before they turn into user complaints. That's usually hard to see when a team is shipping fast. How does Novus decide which issues are important enough to flag?

18
回复

@busra_seker1 Thanks Büşra, great question!! Novus is absolutely great for catching the issues your users face but don't (yet) report and getting ahead of those.

Novus grounds issues from first party data of every track event with built in statistical filters of what's relevant, what's significant, and how important it is. It also watches the real user sessions (through session replay) to understand what's going wrong in every journey.

The agent even understands your business and application in it's contextual Cortex it generates when you onboard to be able to figure out what problems are critical to your product and what could wait.

All of this data helps Novus to cut through the noise and give you real actionable insights, that you can then investigate and solve with Novus itself (it's pretty awesome).

7
回复

congrats on the launch team. I'd be interested to understand how, practically; Novus identifies signal v noise on real issues (a real bug v a housekeeping issue)

18
回复

@zolani_matebese Thanks so much!! Great question, Zolani.

Novus cross-references your GitHub repo, Novus configuration, and live usage data simultaneously, looking for gaps between what's instrumented and what's actually observed.

Every flagged issue gets a severity rating based on data fidelity impact, how central the affected journey is, and whether there's a concrete fix available. And before anything changes, Novus presents an investigation plan for your approval. You see what it found, why it flagged it, and exactly what the proposed fix will do. No PR gets created until you sign off.

17
回复

Hey folks, I lead product for Novus and more generally design & research at Pendo. Why is a designer responsible for such a developer-centric product experience?

My view is every product & design org right now is move closer to technology again, and the best are shipping code every day regardless of their role.

Sure we can now all build more but you don't know if it worked until you have the ground truth of whether it drove adoption.

AMA!

15
回复

Hey builders!

I'm Sean, Lead Engineer at Mind the Product (the world's largest PM community and creators of Vennie.AI). These days thanks to AI we're shipping at 100x speed, but understanding what our users do after we ship is getting harder to keep up (...so many more features to track!).

With Novus, we're able to automatically tag, track and understand our users at a scale that's never been possible as a tiny team within Pendo. We can ask Novus in Slack of any question related to our product, or use the MCP to bring it into our coding agents like Claude and give it the full important context of our usage data.

I'm excited for you all to try it out yourselves. Let us know what you think, and ask any questions. We're here to help 🫡

13
回复

@seandotexe Love this! 👏

Amazing to see how fast teams can build with an idea and the right tools, and even better when they can actually understand what users are doing with what they ship.

Excited to see what how community builds and shapes the product space with Novus. Congrats to the team, and thanks for sharing the journey! 🚀

0
回复

Love the idea of a product agent that proactively finds issues instead of waiting for someone to notice them. Feels like the natural next step after coding agents. Congrats on the launch!

11
回复

@alina_tyslenok_ We're glad you love Novus, thank you!

That's exactly the problem we're trying to solve. Creating an agent that not only tracks data but also watches it for you and finds those insights... rather than you having to sit in a dashboard all day 😎

5
回复

@alina_tyslenok_ Thanks so much and what you said is exactly the awesomness of it too! It finds it, fixes it, and you get to review and merge it 🤘

2
回复

Hey Product Hunt, I’m Brett, Director of Marketing Engineering at Pendo. Tracking new website features has always been challenging. Classes change, elements move, flows get refactored, making manually instrumented analytics painful to maintain. Now my team is leveraging AI to ship at 10x speed which only makes that harder. Using AI tools to ship more doesn’t mean much if you can’t tell whether the features actually add value. That’s why I’m excited about Novus. By connecting directly to your codebase, Novus understands your code updates, then keeps analytics current as your product changes. Excited to hear what you all think and happy to answer any questions I can!

10
回复

Congrats on the launch! Does it also alerts if some anomaly happens in production? (because of some edge cases) Also, how frequently it checks for app health?

10
回复

@ashishkingdom great question! Yes it checks every day for anomalies, including funnel drop offs or being rage prompting your AI, and surfaces that as signals. Novus will then run an investigation and post a PR to fix it too!

15
回复

Hey everyone 👋

I’ve spent the last few months helping with the Novus beta program, hosting a webinar with one of our design partners, and recently getting the Novus community ready for launch (https://community.pendo.io/home/clubs/novus).

Fun meta moment: while building the community, I used Novus to create an engagement guide to our users.

Excited to finally share Novus with the Product Hunt community today 🚀

If you’re building something, I’d love to hear about it in the comments. AMA!

9
回复

Hey Product Hunt builders! 👋

I'm Mike -- Head of Product Evangelism at Pendo and a longtime member of the global product community as Co-Founder of Product Collective and, more recently, Director of Mind the Product. I spend most of my time talking with product people around the world about how our work is changing... and the theme I keep hearing everywhere is the same: we're all shipping faster, but staying on top of what's actually working is becoming its own full-time job.

That's why I'm pretty pumped about Novus. The hardest part of product right now isn't building features -- it's closing the loop between what you ship and whether it's moving the needle. Novus automates that loop, from instrumentation to insight to fix -- and honestly, I haven't seen much out there that does that.

Whether you're a PM rethinking your workflow in the AI era, a founder trying to understand if this fits your team, or just curious about where product intelligence is heading -- feel free to drop a question below! Happy to chat.

9
回复

Great work on Novus, Pendo folks! The ability to identify usability issues and not only flag them but produce code to fix them is pretty awesome.

8
回复

Hi folks!

I'm the product lead for Novus, and what excites me most isn't just the automation. It's that Novus closes the full loop: from automatically instrumenting your code, to proactively surfacing signals, to implementing fixes via pull requests. No tool has done this before, and it changes what's possible for product teams of any size.

Can't wait to see what our beta customers build with it — AMA!

7
回复

One of the things I love most about Novus is the ability it is helping builders ship! There is no reason to get caught up and taken off of our momentum when building a great product to have to fix something you missed and that is exactly what I am seeing Novus help people with and its awesome. Improving your product as you build is just something a small team used to not be able to do and now you keep building and have an amazing tool polishing up that final product.

6
回复

I use Novus on Novus as one of the designers on the product team! Now that I'm bridging both design and engineering — actually building the features that I design — Novus has become an integral part of how I understand both how customers are using what I shipped, and where the gaps are. The PR reviews give me that extra layer of support and confidence, and signals help me close the loop by monitoring metrics and resolving issues, all in one place.

6
回复

This is a strong direction. Teams are shipping faster, but usability issues often show up after users struggle, churn, or complain. I like the idea of a product agent that watches continuously instead of waiting for manual QA or user feedback.

What kind of issues does Novus catch best today: confusing flows, broken states, or low-conversion moments?

6
回复

@thamibenjelloun thank you for your question and congratulations on your launch today!

Low conversion and frustration signals are strongest because conversion is key to a successful business while frustration is a quick early signal you see as soon as the day of release

6
回复
#9
TimeTuna.com
If Calendly had gorgeous video backgrounds
125
一句话介绍:TimeTuna是一个为创始人、自由职业者和工作室设计的预约页面工具,通过支持自定义视频背景(如YouTube视频),让原本枯燥的预约体验变得富有情感和个性,解决同质化严重的SaaS预约工具“千篇一律”的痛点。
Productivity Meetings Video Art
预约工具 Calendly替代 视频背景 品牌个性化 设计导向 SaaS 自由职业者 客户转化 Bookme 自定义域名
用户评论摘要:用户高度认可视频背景带来的情感冲击和独特感,但主要担忧集中在可访问性(如老年人或屏幕阅读器用户)和视频是否会对转化率造成干扰。创始团队回应称已提供参数禁用背景,并承认无障碍功能尚未完善,将改进。
AI 锐评

TimeTuna的入场策略非常聪明:它没有试图在日程管理、团队协作或API集成这些“堆功能”的泥潭里与Calendly一较高下,而是直接转向了体验设计这个空白地带。“If Calendly had gorgeous video backgrounds”这句标语把野心和谦逊都摆在了桌面上——它承认自己是Calendly的替代品,但又明明白白告诉你,我卖的是一种“情绪价值”。

从获客角度看,它精准切中了三大高净值用户群:创始人、自由职业者和工作室。这群人的共同痛点是,他们的预约链接往往是对外输出的第一道“品牌门面”,而Calendly的默认UI恰恰是反品牌的——千篇一律的白色方块。TimeTuna让预约页面从“事务性工具”变成了“表达性媒介”,本质上是在帮用户做“品牌溢价”。

但产品最致命的软肋,在其核心矛盾:**“美”和“效率”的冲突。** 视频背景无论多好看,本质上是一种潜在的注意力耗散。用户评论中直接有人质疑“会不会降低转化率”,这不是吹毛求疵。对于冷启动的陌生人预约,一个炫技的页面可能让客户犹豫;对于已经建立信任的熟客,视频又显得冗余。创始团队给出的“cold outreach用强视频,已约见用弱视频”的区分建议,听起来合理,实则把操作复杂度甩给了用户。

此外,可访问性问题绝非小修小补。“用参数禁用背景”这种开发者思维的解决方案,对普通用户极不友好。忽视无障碍标准在欧美市场是合规风险,更是直接切断了老年用户或残障人士这一大块潜在客户群。

综合来看,TimeTuna是一个“小而美”的尖刀产品,价值在于为预约链接赋予了情感温度及品牌识别度。但它若要破圈,必须解决两个核心问题:一是提供更智能的“上下文感知”切换(比如自动识别用户设备或来源,决定是否展示视频),二是将无障碍设计从“事后补救”提升为“默认就绪”。否则它大概率会停留在审美小众的圈子,无法威胁Calendly的基本盘。

查看原始信息
TimeTuna.com
If Calendly was for people who care about design. We help founders, freelancers, and studios create a memorable impression on their prospective clients through booking pages with beautiful video backgrounds. Your video, your domain, your taste. You have a character. Let you booking page have a character too. Free to start, $10/month for the rest.

scheduling space is so crowded, but this is the first time in a long time I've seen a tool create a distinct emotional reaction. A video background makes the booking experience feel like an invitation rather than a transaction.

Since you mentioned harsh feedback is the most useful: my only worry is readability for older clients or people using accessibility screen readers over a moving background. Do you have high-contrast text options or fallback solid colors for accessibility compliance? If you've solved that, this is flawless.

3
回复

@priya_kushwaha1 Good point on the accessibility. Currently you can pass parameters such as: https://timetuna.com/pavel?translucent=false&background=false&theme=light to remove background, to remove translucency, etc to increase readibility if we send to a person that we know might have accessibility requirements.

1
回复

@priya_kushwaha1 good point Priya! We have not implemented many accessibility features yet. We’ll definitely look into this! 👍

Thanks for your feedback!

3
回复

Lets goooooo TimeTuna!!

2
回复

Hi,

It is the 3rd time launching TimeTuna on Product Hunt.
The previous two ended at #6 (July 2025) and #1 (Dec 2025) of the day, which still feels surreal, so first: thank you for that.

Scheduling is the most boring problem in software. A link, a slot, a confirmation. Every tool in this space ends up looking the same: a beige or a white rectangle, a font that screams SaaS. Should it be this way?

While we developed a lot of features that many other scheduling tools have (e.g. multi-account support, custom domain, meeting types, accepting payments, webhooks, website embed, team support), today we primarily highlight the support of video backgrounds, and Youtube backgrounds specifically. Let your booking page have some character! Live a little :)

Most of the existing active users of TimeTuna are: startup founders, investors, freelancers, agencies and studios. They choose TimeTuna mostly for the unique look with solid features.

If you have a booking link out in the world, paste it in the comments. I'll redesign it as a TimeTuna page and post both side by side. Curious what you think looks more like you.

Free plan covers most use cases. Pro is $10/month. Harsh feedback is the most useful thing you can leave here. Grab your 30% discount code.

1
回复

Really cool tool, I were using this tool like for 5 months and it's awesome ))

1
回复

@ivan_senkin that is so kind of you!

0
回复

@ivan_senkin Great to hear Ivan!

0
回复

First time coming across TimeTuna as a Calendly alternative and the YouTube video background really is cool... but I'm wondering, wouldn't it negatively impact conversions? Don't you want to keep people on-task in your scheduling tool? I worry the YouTube video would distract.

1
回复

@denitsapenchevavaltchanova I believe that for different use cases different booking pages could work. If this is a cold outreach and the person does not know you much - having a booking page with a video explaining what you do is a strong case, if you have already agreed to meet, you could share a booking page with a more subtle video background - just to give a warm feel, but not to distract.

1
回复

@denitsapenchevavaltchanova Definitely valid point, Denitsa!

If you're a real estate agent, you want to get people into the vibe. You could even create a new page per home you're selling.

Something like this would get people into the mood of booking a viewing, no? ;-)

https://timetuna.com/yannick-veys-102025

1
回复

👏

0
回复

Booking just got simple. Love this.

0
回复

The "paste your booking link, I'll redesign it side by side" move is a great engagement mechanic, instant before/after right in the comments.
Question on the $10 tier: with a free plan that covers most use cases, what's the moment a free user decides to upgrade? First compliment from a client on their page, or something else?
Can you actually see that trigger or have to guess.

0
回复

@and_bayleaf We don't have enough data on that right now.

We expect most people already have a meeting scheduling tool and just move. They move to the $10 plan because of the custom domain and the removal of TimeTuna's watermark on booking pages.

0
回复
#10
Keen Code
A context-efficient CLI coding agent built by agents
116
一句话介绍:Keen Code 是一款由AI代理自主编写、专注于上下文效率的开源CLI编码助手,通过“轮次记忆”和“惰性加载MCP技能”大幅降低长会话中的token消耗,解决多轮交互中上下文窗口快速膨胀导致的性能瓶颈。
Open Source Developer Tools Artificial Intelligence GitHub
开源 CLI 编码代理 上下文管理 MCP服务器 轮次记忆 惰性加载 AI驱动开发 Go语言 工具链
用户评论摘要:多数评论认可其上下文节约设计(轮次记忆、惰性加载MCP),并关注策略细节:如何决定保留或丢弃内容?如何进行多文件优先级处理?作者回应:仅保留文件变更与失败命令的简单结构,工具结果按需重跑;暂无多文件优先级逻辑,后续计划优化MCP工具支持。
AI 锐评

Keen Code 的价值不在于“又一个AI编码工具”,而在于它精准击中了当前CLI代理最隐蔽的痛点——上下文窗口浪费。多数同类产品在长会话中线性膨胀token消耗,而Keen通过“轮次记忆+惰性加载”的组合拳,将多轮交互的上下文开销压制到近乎单次调用的水平,这本身就是对“AI工具”设计范式的一次务实重构。

但需冷静的是,它本质是一个“极简主义”的妥协方案:放弃工具结果的长期可追溯性,要求代理在后续轮次中重新执行工具(如读文件),这虽合理却增加了延迟与失败风险。其“文件变更+失败命令”的轮次记忆过于简陋,面对复杂重构或跨文件依赖时,模型可能因上下文断片而产生幻觉。

此外,开源无商业模式意味着长期维护力存疑;而“由代理自建”的噱头虽有趣,却未解决其自身代码质量是否优于手工构建的本质问题。Keen Code 是一把锋利的手术刀,专治上下文肥胖症,但别指望它包治百病——适合对token开销极度敏感、能容忍轻量状态遗忘的开发者,而非需要深度上下文推理的重度场景。

查看原始信息
Keen Code
Keen Code is an open-source, context-aware and efficient CLI coding agent written in Go. Three aspects stand it out from other similar products: - It was built from scratch by coding agents, with the full prompt/design trail preserved and shared in the repo. - It uses turn memory to keep multi-turn sessions lean which saves context significantly. - It maps MCP servers to lazy-loaded Skills instead of stuffing large schemas into context upfront. This again saves context in mult-MCP setting.
Hi Product Hunt! I’m happy to share Keen Code, an open-source CLI coding agent written in Go. I’ve been building it solo since February as a side project, and I used it as an opportunity to experiment with context efficiency and agent-driven development. Three things make Keen Code different from other similar products: 1. Built by agents Keen Code was built from scratch using state-of-the-art coding agents. My role was to act as the human orchestrator: writing prompts and requirements, then reviewing the designs and code produced by agents. To keep this transparent, the repo includes an ai-interactions folder with prompts and output docs. More: https://mochow13.github.io/keen-... 2. Turn memory To avoid filling the context window during multi-turn loops, Keen discards raw tool inputs and outputs after each turn. It keeps a distilled “turn memory” instead: a simple deterministic Go struct passed into the next turn. More here: https://mochow13.github.io/keen-... 3. Skills-driven MCP servers Instead of loading large MCP server schemas into context upfront, Keen abstracts MCP tools into local markdown Skills. It only retrieves the exact JSON schema when the LLM requests a specific tool at runtime. Details: https://mochow13.github.io/keen-... I’ve been using Keen to develop Keen itself, as well as in my other projects. I’m looking forward to questions, feedback, suggestions, and reviews. I’m committed to improving the project over the long term. Thanks in advance!
3
回复

Building a CLI agent that manages its own context window is a genuinely hard problem. We've dealt with similar tradeoffs in long-running background jobs where keeping relevant context without blowing token budgets required careful chunking. What's your eviction strategy when the agent's working set grows mid-task: do you prioritize recency or semantic relevance?

1
回复

@anand_thakkar1 Hey! Good question. Right now Keen prefers keeping user-agent interaction over tool calls. And in my experience, tool call results have been the single most costly part of the context window.

So if an agent loop grows too much that it cannot fit in the context any more, the simple strategy is to evict oldest tool results for the turn and so on. Since Keen doesn't retain tool results beyond a single turn, the context is already lean when a new turn starts. As a result, this strategy to evict oldest tool results is enough to reduce in-turn context window size.

1
回复

One thing we've noticed with agent workflows is that reducing context size and improving context quality are often treated as the same problem when they're actually different.

Have you found Keep Code's biggest win coming from compression, retrieval, or helping agents maintain a more consistent mental model of the codebase across longer sessions?

0
回复

@zaid_mallik1 Good question. Not exactly mental model, but I have been thinking to experiment with a pre-calculated repository mapping for agents. I don't know if it will surely improve efficiency or performance, but need to actually implement and get a feel for it.

For Keen, yes the focus is indeed context efficiency through compression or lazy loading where possible. I think this is what the biggest win would be.

0
回复

Sounds interesting, but aren’t you worried about competition? There are already many AI coding tools, and new ones appear every day.

0
回复

@natalia_iankovych Hey! This is an open source product and there is no business model associated with it. Yes there is competition which I don't mind at all. If people use it and find it useful, that's good enough for me and worth the effort. :)

0
回复

The context-efficiency angle is what stands out to me here — lazy-loading MCP skills instead of front-loading schemas is a genuinely smart architectural choice that most CLI agents skip. Fellow solo builder launching nearby (Sensemaker, a thinking-to-writing tool), and the "built by agents, orchestrated by human" workflow you described is almost exactly how I shipped my own MVP. Curious: how did you handle the moments where the agent went off-spec mid-session? Did you find the prompt/design trail in the repo was enough to course-correct, or did you need other guardrails?

0
回复

@eran_shayshon Hey, yes going off-spec mid-session has not been a problem. I think it's up to you to carefully design how big of a bite you would want the agent to take. As long as it's reasonably planned, it works pretty well.

0
回复

the skills-driven MCP approach is the part that caught my attention. every multi-tool agent setup I've seen just dumps the entire schema into context upfront and then wonders why performance drops off a cliff three turns in. lazy-loading only what the model actually needs at runtime feels obvious in hindsight but almost nobody does it. also respect for shipping the full prompt trail in the repo... that's genuinely useful for anyone trying to learn how agent-driven development actually works in practice. congrats on the launch and doing this solo is no joke

0
回复

@tina_chhabra Thanks for your feedback!

0
回复

The "lazy-loaded Skills instead of stuffing MCP schemas into context upfront" choice is the detail that stood out to me - that's usually where multi-MCP setups quietly fall apart, the context budget gets eaten before you do any real work.

As someone still figuring out how to keep agent sessions lean, I'm curious: with turn memory compacting older turns, how do you decide what's safe to drop vs keep? Do you summarize older turns or hard-truncate them? Wondering where the line is before the agent starts "forgetting" a constraint from earlier in the session.

Also love that the whole prompt/design trail is in the repo - genuinely useful to learn from. Respect for building it solo. Congrats on the launch.

0
回复

@grace_lee26 Turn memory is pretty simple and dumb. Right now it's a Go struct with two fields: files_changed and failed_bash_commands. This is to give later turns a simple sticky note: this is what changed in the previous turns and this is what failed. That's it.

The reasoning behind this is that if agent needs to refer to some tool result before, it should simply run that tool again. In fact, SoTA agents do that frequently: they read some file, does some change. Later, if you ask them to do something else, they read the same file again to ensure the latest state.

So Keen's thought process is: just don't retain tool call results from read_file or grep or other tools after an agent turn loop finishes.

0
回复

Context efficiency is the right constraint to optimize for in a coding agent. Most agents bloat the context window with irrelevant file chunks and then thrash on eviction decisions. We've hit this exact failure mode building multi-file reasoning features and it's where agent reliability falls apart. How does Keen Code handle context prioritization across a multi-step tool use chain when multiple files are relevant?

0
回复

@retain_dev I am not entirely sure I understood your question :D

Right now, there is nothing done specifically for prioritisation between multiple files in Keen.

0
回复

Really cool, and respect for building it solo. The turn memory idea for keeping context lean is smart. How much context does it actually save in a long session?

0
回复

@ianhxu In my experience, it has been significant. I have done some very basic benchmarking between Keen vs OpenCode. In those benchmarks, I have clearly seen Keen's context window drops 1/10th from one agent turn loop to the next.

Context window size for OpenCode grows linearly. The more a multi-turn conversation progresses, it keeps going up. With Keen, it grows within a turn, but comes back down significantly after a turn. The next turn starts with a far smaller context compared to previous turn's end.

The tradeoff that Keen has here is pretty obvious: if agent needs to refer to some tool call results done a turn or more ago, that result is no more available. The counter-argument here is that models can always re-run those tools, assuming those tools are idempotent. There has to be some shenanigans done for MCP tools, which I intend to address in next releases.

Here is a Claud-written analysis: https://mochow13.github.io/keen-code/docs/turn-memory-analysis.html. It's a bit dramatic but good for understanding the idea.

0
回复
#11
Perplexity Personal Computer for Windows
Run AI agents across your local files and apps on Windows
114
一句话介绍:Perplexity Personal Computer for Windows 是一款将AI代理系统深度集成到Windows本地的工具,通过混合云-本地架构,让用户无需在多个应用和网页间手动切换,即可在本地文件、原生应用与云端服务之间自动执行多步骤工作流,解决知识工作者在复杂跨平台任务中的效率痛点。
Productivity Task Management Search
AI代理 Windows本地 工作流自动化 混合架构 企业效率 多模态模型 跨应用操作 知识工作者 深度集成 隐私保护
用户评论摘要:用户对Windows版发布表示强烈期待和欢迎,如“终于等到Windows版”“已排很久队”。核心问题聚焦技术细节:有用户询问多步骤工作流中如何智能路由不同子任务到对应的前沿模型,官方未回应。
AI 锐评

Perplexity Personal Computer for Windows 的发布,本质上是在打一场“AI入口权”的战争。当ChatGPT、Gemini等对话式AI仍被困在浏览器标签页里,只能被动回答问题时,Perplexity选择直接钻进用户的Windows桌面,变成一只能读写本地文件、指挥原生App、调用云端模型的“数字手”。这种“偷家”策略确实犀利——它试图将AI从“问与答”的工具升级为“做与成”的执行者。

然而,114票的冷淡数据(远非爆款)暗示了市场的审慎态度。真正值得质疑的是:这到底是革命性生产效率工具,还是又一个“看起来很酷的空中楼阁”?

**第一,混合云-本地架构是亮点也是风险点。** 它声称“敏感文件不出机器”,但实际运行时,哪些数据必须上云、哪些本地完成,决策逻辑完全依赖Perplexity的“智慧”。对企业用户而言,这是把信息安全命门交到了一个黑盒模型手中,合规审计将成为巨大阻碍。**第二,生态绑定陷阱。** 它深度整合Gmail、Slack、Notion等SaaS工具,但一旦用户形成依赖,Perplexity便掌握了工作流的核心路由权。未来若涨价或功能降级,切换成本极高。**第三,体验鸿沟。** 当前仅面向Max和Enterprise Max订阅者,且需排队——这意味着普通用户被刻意排除。这种饥饿营销对一款需要“用户规模效应”来优化模型路由策略的AI产品而言,可能适得其反,加速早期用户的耐心流失。

真正的价值在于:Perplexity首次证明了AI Agent从“浏览器AI”向“操作系统级AI”进化的可行性。但若要成为Windows上的“AI调度中心”,它还需要回答一个致命问题:当用户把本地文件、隐私数据和关键工作流托付给你时,你的容错率、透明度和可审计性,能否比得过手工操作?目前来看,答案尚未写在代码里。

查看原始信息
Perplexity Personal Computer for Windows
Personal Computer extends Perplexity's multi-model agent system to your Windows machine, running tasks across local files, native apps, and the web for Max and Enterprise Max subscribers on the waitlist.

Windows users have been waiting for this since Perplexity shipped Personal Computer on Mac in April.

Personal Computer is an AI orchestration layer that runs directly on your Windows machine, working across your local files, native apps, and the web to complete multi-step tasks without you manually switching between them.

Most AI agents live in a browser tab and have no access to what's actually on your machine. Personal Computer closes that loop by acting on your machine, not just answering in a chat window.

What makes it worth paying attention to is the hybrid local-cloud model. It decides in real time which parts of a task stay on your device and which go to frontier models in the cloud, so sensitive files don't leave your machine unless they need to.

Here's what it can do:

  • Read and write across local files and native Windows apps

  • Orchestrate tasks across Gmail, Slack, GitHub, Notion, Salesforce, and more

  • Route subtasks to the right frontier model automatically

  • Run continuous workflows in the background while you're away from your desk

Built for Windows-based knowledge workers and operators who manage complex workflows across local files and multiple apps. Access is rolling out first to Max and Enterprise Max subscribers on the waitlist. If that fits your plan, worth signing up now.

P.S. I hunt the latest and greatest launches in tech, SaaS and AI, follow to be notified @rohanrecommends

2
回复

@rohanrecommends when orchestrating multi-step workflows across local files, native apps, and cloud services, how does personal computer decide which fronteir model to route each subtask to?

2
回复

finally there's a launch for windows. I have signed up for the waitlist for sooooo long ( ̄▽ ̄)

1
回复

Finally, i was waiting for this from any AI LLM

0
回复
#12
Koji by Brilliant
A world-class personal tutor for every home
112
一句话介绍:Koji 是一款面向家庭数学和编程学习的超智能一对一私教工具,通过在屏幕上实时圈画、语音交互和自适应难度调节,解决传统真人家教昂贵、难以触达且无法随时响应用户卡点的问题。
Education Artificial Intelligence Online Learning
AI 教育 自适应学习 智能家教 数学辅导 编程学习 语音交互 屏幕涂鸦 可汗替代 k12 教育 超级应用
用户评论摘要:用户普遍认可屏幕涂鸦和语音多语言功能,认为比纯文本讲解更贴近真人辅导。一个关键追问是遇到学生反复犯错时,Koji 能否切换策略而非重复相同讲解。官方回复确认会换路径尝试,并强调 AI 拥有“无限耐心”的优势。
AI 锐评

Brilliant 过去以“高质量互动课程”闻名,如今推出 Koji,本质上是一次从“内容平台”向“AI 教练”的惊险跳跃。涂鸦、语音、自适应这三板斧确实比市面上绝大多数“聊天框+”式 AI 助手更有教学感,尤其是“指着屏幕讲题”这个物理互动细节,直击远程教育最缺失的临场感。但官方宣称“这不是聊天机器人”,实际体验却难逃对话式追问的结构——再好用的涂鸦也只是 UI 层,核心依然是模型对问题的拆解能力。评论里那句“遇到反复犯错会不会换策略”才是真正试金石:当前大模型在数学推导中的自纠能力依旧脆弱,一不小心就陷入“礼貌但无用”的循环。另外,产品覆盖六年级到大学,跨度太大,单个“超级智能体”能否在抽象代数与循环语句之间自如切换,值得怀疑。从商业角度看,Koji 准确切中了“高性价比一对一”这个刚需,尤其对欧美家庭教育预算收缩的当下极具吸引力。但它的真实护城河不在模型本身,而在 Brilliant 多年积累的互动课程体系和学生行为数据——没有这个底座,Koji 就是又一个聪明的空壳。它要真正成为“世界级家教”,还需证明自己在复杂数学逻辑推导中的稳健性,以及面对不同学习风格时的策略多样性。目前来看,它更像一个设计极佳的原型,而非成熟的替代方案。

查看原始信息
Koji by Brilliant
This isn’t just a feature launch, it’s an evolution of our entire product experience. Brilliant is now a superintelligent tutor designed to work like the best human tutors. It asks the right questions, sketches right on your screen, and adapts to how you learn best.
Hi Product Hunt! We’re so excited to share something we’ve been building toward for a long time. One-on-one tutoring has always been the gold standard of learning. Everyone knows it. The problem is that a great tutor costs hundreds of dollars an hour, is hard to schedule, and is out of reach for most families. With this launch, we’re changing that. Koji is Brilliant’s new superintelligent tutor for math and coding. This is not a chatbot. It’s not a multiple-choice question generator. It’s something that we think is totally magical. Here’s why: • Koji can point, sketch, and annotate directly on your screen as you work through problems together • Koji's coaching adapts continuously to your level — breaking down hard problems into steps when you’re stuck, scaling up the challenge when you’re ready • You can talk to Koji naturally by voice or text, in any language This is the tutor we always wanted. And it's layered on top of our full interactive curriculum, covering grade 5 through college and beyond. We’d love to hear what you think. Take Koji for a spin and let us know!
5
回复

the sketching directly on your screen while you work through problems is what makes this feel different from every other AI tutor out there. most of them just dump text explanations and hope you figure it out. having something that actually points at the part you're stuck on and walks you through it visually is so much closer to how a real tutor works. also the voice in any language part is huge for accessibility. curious how it handles the moment where someone is clearly frustrated and keeps getting the same thing wrong... does it switch approach or just repeat the same explanation slower

1
回复

@tina_chhabra Yes! Making it so Koji can point to and draw on problems directly is core to what makes the tutor so effective. Koji can add points to a graph or ask questions about different parts of a problem, like "Which of these operations is the next step?" Our goal is to help the learner get "unstuck" just enough so they can solve the problem on their own. Koji also will try a different tack if the learner isn't getting it one way. This is one advantage of an AI tutor over a human – infinite patience! :)

1
回复
#13
Boxes.dev
Run Claude Code and Codex in your own cloud environment
107
一句话介绍:Boxes.dev为每个AI编码代理(如Claude Code、Codex)在云端提供独立计算机,解决本地资源竞争、环境混乱和移动办公痛点,让开发者从任何设备安全地并行运行推理代码。
Developer Tools Artificial Intelligence Vibe coding
云端开发环境 AI编码代理 Claude Code Codex 隔离计算 移动开发 并行代理工作流 DevOps自动化 MCP支持 开发安全
用户评论摘要:用户点赞隔离执行和移动便利性,主要问题包括:能否运行Playwright测试(支持)、如何处理环境变量和密钥(提供安全机制)、并行性(支持独立VM副本)、长途旅行延迟(支持区域部署)。另有关注长期工作空间连续性问题(通过模板快照实现)。
AI 锐评

Boxes.dev精准切入了一个被许多人忽视但日益紧迫的需求:AI代理不是单纯的工具,而是失控的“租户”。当开发者允许一个编码代理在本地裸奔安装依赖、修改配置、运行测试时,其破坏力不亚于任意代码执行攻击。Boxes的“每代理一台独立Linux主机”方案,本质上是在AI能力与开发者信任之间构建了一层“强制隔离墙”,彻底解决了“代理把开发机器搞瘫痪”或“多代理抢资源”的物理摩擦。

从产品设计看,它并未试图替代本地IDE或沦为远程桌面,而是聚焦于“将AI代理视为全栈操作员”的场景。模板快照与并行VM分支机制尤其聪明,这等同于为每个代理任务提供了一次性的、可重复的“沙盒化构建环境”,完美匹配CI/CD中“并行流水线”的思维。此外,它悄然完成了从“开发者手动维护环境”到“AI自动同步并快照环境”的权力转移,这比任何RPA工具都更贴近原生云开发范式的本质。

不过,其面临两大挑战:一是成本与规模效应,每个长时间运行的代理都消耗真金白银的云资源,这比本地零边际成本模式更难让个人开发者长期接受;二是“AI代理本地调试循环”的延迟问题——纵使其宣称“仅发指令看图”低敏延迟,但开发者一旦需要终端介入(如调试复杂错误),网络丢包立即将体验打回原形。总体而言,Boxes.dev是资本效率导向的“AI I/O代理基础设施”,专为高价值、高负载的团队化编码场景(而非个人草稿)而生,其实际价值将取决于它能多快碾压本地方案的直觉和成本偏见。

查看原始信息
Boxes.dev
Cloud dev environments for agentic coding. Run each Claude Code or Codex chat on its own computer in the cloud, connect from mobile and desktop, and code from anywhere.

Hi Product Hunt! It's great to be back with a new launch. Fond memories of launching my last company, Gem, here way back in 2017!

We’re Nick and Drew, and we’re building boxes.dev – the first cloud-only agentic dev environment (ADE) that gives every Codex and Claude Code agent its own cloud computer.

We spent the last year coding almost exclusively with Codex / Claude Code, but got tired of leaving our laptops cracked open and dealing with clunky git worktrees. Developing on localhost was holding us back, so we decided to build the cloud-based ADE that we wished existed.

We’re obviously biased, but we’ve been building boxes.dev with boxes.dev for months and it’s been a gamechanger – it’s hard to imagine going back.

Key features:

  • Super easy setup (our agent scans your local dev setup and ports it to the cloud)

  • Full-featured desktop, CLI, and mobile app (not just handoffs or remote control)

  • Uses your Claude Code / Codex subscription (native harnesses & familiar UX)

  • Your coding agents can run/test your full app end-to-end in isolation

  • Scheduled automations & Slack integration

  • Optimized for parallel agent workflows

All new users get 10 free box-hours to test it out at boxes.dev. Hope you’ll give it a try – would love any and all feedback!

5
回复

@nbushak Looks great! Can the agents spin up playwright for tests?

2
回复

Congrats on the launch Nick and Drew!

Early user of Boxes, and a big big fan. Some of my workloads are multi-hour so I’d previously have to either delay kicking it off or walking around with my laptop open (if you’ve been around South Park in SF you may have seen me doing this 😅).

Boxes has made this so much more convenient. But also broadly I’ve just started spending way less time at my computer, doing my thinking on whiteboards and notebooks, compiling requirements that way before kicking off tasks via my phone instead.

And for those wondering “why not use the cloud hosted versions of Claude Code or Codex” - for me a lot of it is down to needing local instances of a db for testing, wanting playwright screenshots for certain tasks when completed, etc. it’s literally my computer on the cloud.

2
回复

@raveesh Thanks so much for being an early user and believer in us!

0
回复

This looks great Nick, congrats on launching! Having a way to give Claude its own computer rather than trusting it with your own (terrifying) is a real smart idea and the space is heating up. I'm rooting for y'all and will be sure to try it out!

2
回复

@seandotexe Thanks Sean! Yes, it feels freeing not having agents running wild on my personal laptop, with our VMs only having access to development credentials.

0
回复

Does it have access to mcp and skills? can it run things like playwright mcp? how does it work?
I could be very interested to use it in my startup

1
回复

@fberrez1 Yes, it does! You get full control over a linux box in the cloud for each claude code and codex thread (and a template box that gets snapshotted and cloned to produce the VM for every thread). It's great at running chromium headless, running playwright, and even sending you screenshots to the thread (visible in the desktop app and mobile). And because each thread/VM is independent of the others you can have a bunch of them running multiple copies of your app and chromium in parallel without slowing down the others. If it works on your laptop via the codex/claude CLI, it will almost certainly work on boxes.dev.

1
回复

Nice! One thing I use all the time with claude code on my local is the secrets (.dev.vars, .envrc) so I can test things like OAuth flows or 3p integrations. How do you handle that with boxes.dev? Is there a secure way for me to put secrets on the boxes?

0
回复

Congrats on the launch, using it now!

0
回复

@iloveluce Thanks Luciano!

0
回复

One thing that seems under-discussed with cloud coding environments is continuity.

Are users primarily spinning up fresh environments for each task, or do the most active users end up building long-lived workspaces that accumulate context over time?

0
回复

@samyak_sanklecha We handle this with the notion of a "main box" or template box, which you can imagine to be your development laptop in the cloud. When you set up boxes.dev, our coding agents find your environment variables and undeclared dependencies locally, and then get your app fully running in a cloud VM. After it's set up, you can always jump into the terminal or make any changes you want on the template box.

When you create a new task, that main box is snapshotted (filesystem and RAM), and then a new independent fork is created from the snapshot.

So whenever you want to add context or update your setup on boxes.dev, change it on the main box and update the snapshot, and all future threads will get created from that snapshot.

0
回复

Congrats on the launch! I've been playing around with the platform for a while now and have been blown away

0
回复

@gil_feig Thanks a ton Gil!

0
回复

This looks great. I spend a lot of time running coding agents so giving each one its own machine makes a lot of sense. Can you run several agents in parallel on the same repo?

0
回复

@ianhxu Yes! This was the primary motivation for us -- you can have multiple threads/agents running in parallel, each with their own totally independent machine. No more running out of resources locally or having to juggle git worktrees. Since we have a template/main box that you have full control of, there's also no setup required per-thread/per-machine, each thread's VM starts from a perfect snapshot of the template box.

0
回复

Awesome idea. Have several my own boxes in USA running under my TV. However now I am traveling in Asia and latency makes work uncomfortable. Hope your boxes will be fast.

0
回复

@sergebulaev Over the long term, we'll support spinning up VMs in local point of presence datacenters, so any threads/agents you create or wake up while traveling can be colocated with you.

But one thing I've noticed using boxes.dev on a plane is that since I'm just sending messages to a thread and looking at screenshots from the agent testing its own work, the user experience isn't that bad even with a lot of latency. Of course, if I have to use the terminal, that's slower, but on a plane I end up just pushing the agent to do whatever I wanted to do manually in a terminal window.

0
回复
#14
Sun
Collaborative voice API for agents
98
一句话介绍:Sun是一款专为多人实时协作场景打造的语音AI API,解决了现有语音模型(如ChatGPT Realtime)只支持一对一对话、在多 speaker 会议或群组讨论中无法处理轮流发言、打断和身份识别等痛点。
Meetings Developer Tools Artificial Intelligence
语音API 多说话人感知 实时协作 AI代理 语音交互 会议智能 多代理辩论 上下文窗口 打断控制 教学场景
用户评论摘要:用户主要关切:是否能用于会议助手(如Fireflies/Otter)实现主动发言;如何处理打断时机,支持触发词唤醒;能否中途添加上下文;支持多少说话人。团队回应称支持无限说话人(测试过25人会场)、触发词唤醒、及用“context.update”动态注入信息。
AI 锐评

Sun的定位精准切中了当前语音AI领域的空白——绝大多数实时语音API(OpenAI Realtime、Gemini Live)本质仍是“一对一对话”的延伸,而Sun试图用“多说话人感知+10倍上下文窗口+代理感知打断”构建一个真正的协作基础设施建设。

从产品设计来看,Sun并非简单地给现有模型打补丁。其核心创新在于“agent-aware barge-in”取代了传统的VAD(语音活动检测),这意味着AI代理不再因声浪被动响应,而是能基于语义和角色逻辑决定何时插话。这条路径如果走通,将彻底改变会议纪要工具、在线教育、客服多路通话等场景的范式——从“监听”跃升为“参与”。

但风险同样明显:技术实现难度极高。多说话人场景下的延时、重叠语音处理、上下文一致性,是业界公认的“三座大山”。用户评论中提到的“老年人语速不均”“超过5人时的延迟”正是致命考验。目前Sun依赖文本输入(转录),这意味着对上游ASR的准确性强耦合,任何转录错误都会直接导致打断判断失真。此外,测试数据声称支持25人,但实际生产环境能否稳定支撑“三到五人轮番插话并保持逻辑流畅”,才是检验产品力的标尺。

市场策略上,Sun选择从开发者社区切入,用“免费Playground+公开征集压力测试”的做法很务实。这种开放姿态能快速积累灰犀牛案例,但也暴露出产品尚未经过大规模商业验证。集成支持(LiveKit、Twilio等)虽多,但真正能决定生死的是与现有会议生态(Zoom、Teams等)的深度绑定——若仅作为独立API,则容易被平台自研能力反噬。

一句话总结:Sun踩准了“从对话到协作”的转折点,但破局的关键不是堆功能,而是用极致的低延迟和高鲁棒性,把“知道该在哪个分母上开口”这件事做得比人类还准。否则,它只会沦为又一个聪明的噪音源。

查看原始信息
Sun
Sun is a voice first AI model built specifically for real‑time collaborative voice interaction, not just one‑on‑one chat. ChatGPT Realtime and Gemini Live were built for one user talking to AI. Sun is built for collaboration — meetings, group calls, multi-agent debates, classrooms. One API, multi-speaker awareness, 10× the context window.
Hey Product Hunt 👋 I'm Anand, co-founder of Sun (https://getsun.io). Every realtime voice API today — OpenAI Realtime, Gemini Live, Hume — was built for one user talking to one AI. That breaks the moment a third voice enters the room. Sales calls, classroom debates, multi-agent workflows, group brainstorms — they all need voice infra that knows who's talking, when to interrupt, and how to let three speakers share a turn. 🌞 Sun is built for that: • Multi-speaker turn-taking • 10× the context window of ChatGPT Realtime and Gemini Live • Agent-aware barge-in (not just VAD) • Multi-agent in one room — run two AIs against each other on a real audio channel Today's PH offer: • Live playground — try it in your browser, no Credit Card → https://getsun.io • Live demo at https://demo.getsun.io Two things I'd love your help with: 1. Tell me where this breaks. We've stress-tested ~20 multi-speaker apps; we want yours to be #21. 2. What integrations would unlock you? LiveKit, Daily, Vonage, Twilio, custom WebRTC — drop a comment. Huge thanks to Anoop for hunting. Happy to answer anything in the comments today. 🌞
16
回复

So, is this like speech engine for note takers like otter or fireflies ? like making them talk back. fireflies does that now, but it takes forever to make it work.

1
回复

@arjun_reghu Yes, absolutely. This will help with in-meeting intelligence (like meeting assistant) and is faster than the fireflies version. It can respond in voice in a very short time. For comparison, for a smaller answer, fireflies will start speaking after the sun agent has completed the answer, or is halfway. Take a look into our playground where you can try it out.

0
回复

Does it take audio in or should we give transcribed text as inputs ?

1
回复

@jaison1993 Sun takes in text input (both partial and final transcripts - streaming). It gives out both text and audio output.

1
回复

how does it know when to interrupt and when to not? Also, is there way to make it speak only when explicitly asked?

1
回复

@ashishkingdom Yes, you can actually give trigger words or names to the agent. The agent will then respond only when called with the name (and is asked a question), not when passively mentioned in a conversation. It also looks at voice activity to see if someone is actively speaking to avoid interrupting them.

1
回复

Does it mean that by using this, we can make our internal meeting agent talk in meetings about what is happening in meetings ? Like talking fireflies or otter ?

1
回复

@regaldreamtech Exactly — that's the core use case Sun was built for.


Fireflies and Otter listen silently. With Sun, you can build an agent that listens, knows who's saying what, and actually speaks up at the right moment — to summarize, answer a question, surface a related decision from earlier in the meeting, or even moderate.

2
回复

The classroom use case is the one that caught my attention. An AI that can follow multiple speakers and participate at the right moment could be genuinely useful during group discussions.

Have you tested it with actual teachers or students yet? I'd be curious to hear how it performs in a real classroom setting.

Congrats on the launch!

0
回复

The "when to interrupt" part is exactly where most voice infra falls down. I'm building voice AI for older adults, and barge-in plus handling slow or overlapping speech is the hardest piece, models either talk over people or freeze. Curious how Sun handles turn-taking when speakers have very different pacing, and whether latency holds up with 4-5 voices in the room?

0
回复
In the playground, there is a trigger words list. Why is it there even though there is already agent name field ? Is it for alternate names ?
0
回复

@aisha_m_a Yes, the trigger words list is for alternate names of the agent. Agent will see it as a wake up call to answer. It is also helpful if the transcription service used sometimes misses to transcribe properly. For instance, John can be sometimes written as Jon and this could help in those scenarios as well.

1
回复
Can we add or give information to the agent mid-session? Or is it possible only at the beggining of the session?
0
回复

@aswanth_viswanathan8 Yes, it is possible to add context mid-session using context.update. It could be useful if you want the agent to have certain information, like sales number or a context about the ongoing discussion, in the middle of discussion.

2
回复

This is interesting, how many speakers/voices does it tolerate?

0
回复

@real_digidavid The model doesn't detect speakers, but it takes in speaker name and what they said as a streaming text input. It can therefore support infinite speakers theoretically. We have tested up to 25 participant meetings. The output from the model is both text and audio. The output audio has 4 different voices which you can switch while the session is in progress as well. 2 male and 2 female voices.

1
回复
#15
Carbon Voice Speed Dial
Get your whole team (humans and agents!) on speed dial
97
一句话介绍:Carbon Voice Speed Dial 是一款为团队打造的“语音快捷拨号”工具,让用户通过一个按键或一次点击即可与同事或AI智能体(如Hermes、n8n等)进行异步语音通话,解决会议频繁、沟通效率低下的痛点。
Messaging Artificial Intelligence Bots
语音快捷拨号 AI智能体协作 异步沟通 团队协作工具 桌面热键 移动应用 工作流自动化 降本增效 Product Hunt
用户评论摘要:用户赞赏异步语音+热键组合能减少会议时间、降低上下文切换成本。有用户询问是否支持Telegram等第三方平台,官方回复称可通过n8n/Zapier中转。技术用户关心语音延迟和转录对专业术语的准确性。
AI 锐评

Carbon Voice看似是一个“语音版通讯录”,但其真实价值在于解构了“会议”这一现代职场的效率黑洞。它用“异步沟通”取代“必须同步”的会议——一个按键完成消息投递,无需预约、等待、寒暄,尤其适合早已厌倦Slack轰炸和日程绑架的远程团队。

产品的野心不仅限于人际沟通,更聪明地将AI代理纳入“团队通讯录”。当Hermes、Claude Code等工具能像同事一样被“一键呼出”时,语言交互由“人去适应机器”变为“机器像人一样被使用”。这本质上是把AI从对话框拉入了工作流的核心。不过,这种便利高度依赖生态集成和转录准确率——技术团队口中的变量名、API报错等术语若频繁出错,便会让用户退回打字确认,破坏“不假思索”的体验。

犀利之处在于:它不是在“更好开会”,而是在“消灭会议”。但这也是一把双刃剑——异步沟通固然高效,却也容易导致信息过载和“读语音”的不耐烦。若不能平衡好推送的“轻”与信息密度的“重”,很可能成为又一个被屏蔽的噪音源。对于已经用惯了Zoom或Teams的团队而言,改变心智模型比安装一个工具困难得多。它需要是“战术性武器”而非“玩具”,而这取决于Carbon Voice能否在早期技术用户的基础上,用真实的ROI说服管理层。目前97票的社区热度,离“颠覆会议”还差一个完整的飞轮。

查看原始信息
Carbon Voice Speed Dial
Talk to your whole team — people and AI agents — with one tap. Carbon Voice gives you a speed dial for your team. One keystroke on desktop, one tap on mobile, and you're talking to anyone — a teammate, a group, or an AI agent like OpenClaw, Hermes, n8n, Claude Code, or Tasklet. Teams see 60% less time in meetings. Now your agents are just as easy to reach.

Hey Product Hunt 👋


I'm Travis, CEO of Carbon Voice.


Carbon Voice gives you a speed dial to talk to your whole team, people and AI agents, with one tap.

Built for builders and teams who are running AI agents alongside their people and want to communicate at the speed of thought.


🎙️ Here's how it works:

  1. Assign your team, people or agents, to a speed dial hotkey

  2. One keystroke on desktop or one tap on mobile and you're talking

  3. Listen to replies or read the transcript, your choice

🤩 Why you want Carbon Voice:

  • Your whole team in one place - people and agents, one tap away

  • 60% less meeting time - talking is faster than scheduling

  • Works with everything - Hermes, OpenClaw, n8n, Claude Code, Tasklet, any webhook

  • Async by design - speak your thought, move on, get pinged when ready

  • Desktop, mobile, watch - one tap away wherever you are

What agents are you running that you'd want to talk to? Drop them below, we’ll add it next if we don’t have it already.

— Travis

2
回复

Voice is often enough for context, but meetings turn it into a 30-minute block for everyone.

I like the idea of async voice + transcript because it keeps the human tone without forcing everyone into a calendar slot.

What use case is working best so far: internal updates, client feedback, or team decisions?

2
回复
@thamibenjelloun Exactly right. Those magic words of “let’s talk about this” shouldn’t have to become a synchronous event. I find most discussions are done with sync voice before they could have ever gotten scheduled and they take a fraction of the time.
0
回复

Big fan of Carbon Voice - have been using it for over 2 years. Great to talking async to my team and friends - there are many small improvements over sending voice notes through other apps - and they all add up to a much better experience.

Love the addition of now also hooking up the ability to talk to my AI agents. Nice that you chose Hermes for your demo!

1
回复
@nickholzherr Hermes feels it most natively is trying to distill down its response for a message vs a data dump of information. It feels more like I’m talking to a teammate rather than a researcher sharing all their knowledge. 
0
回复

The async voice + hotkey combo addresses a real pain point I see with my remote technical teams; context switching kills productivity, especially when you need to loop in specific people in different time zones.

Curious how you're handling voice processing latency and whether the transcription accuracy holds up with technical jargon and domain-specific terminology that dev teams typically use.

0
回复

Hi!

Just to make sure I understand it correctly: Could I create a hotkey for a specific contact in Telegram and instantly send them a voice/text message?

Or does it work only with AI agents and agent integrations?

0
回复

@ilya_makarov2 Good question. You can set a hotkey to point to a person or group of people and that will deliver to their email as a transcript + link to voice. If they themselves are Carbon Voice users, it will deliver right into their conversastions.

Or you can have a hotkey pointing to an agent or any webhook. There isn't a direct feed to a Telegram user yet, but could hit a n8n or Zapier workflow that would direct it to that user on Telegram or any other platform.

0
回复
#16
Basedash Semantic Layer
Define metrics once. Use them everywhere.
97
一句话介绍:Basedash Semantic Layer是一个语义层工具,让团队在数据源中定义一次可复用的SQL指标(如月经常性收入),AI随后即可在聊天、图表、仪表盘等场景统一调用,解决AI分析中指标定义不一致和重复编写SQL的痛点。
Artificial Intelligence Data & Analytics Business Intelligence
语义层 AI分析 SQL复用 指标治理 商业智能 数据平台 团队协作 自动化报表 数据分析 企业级SaaS
用户评论摘要:用户关注语义层在多团队场景下的实际落地,核心疑问包括:同一指标在不同上下文(如货币换算、时区、业务单元)需要微调时AI如何处理;不同团队定义指标略有差异时的解决方案。创始团队回应支持团队级指标集及以他人指标为基底的嵌套定义,并指出多数企业由中心化团队负责统一指标;有用户认为技术定义容易,组织达成共识更难。
AI 锐评

Basedash Semantic Layer切中的确实是AI+BI领域一个真实且持续的痛点:当AI开始“听懂”数据语言时,它最常犯的错误不是不会分析,而是用错误的逻辑算出漂亮数字。这个产品通过强制引入“语义层”作为中间件,本质上是在赋予AI“纪律性”——告诉它哪些SQL是经团队背书的“官方版本”,而不仅仅是靠模型推断。

从实施角度看,这个方案的价值在于它把“指标治理”从文档文化变成了可执行的代码。过去,团队靠文档或记忆来对齐“什么叫激活率”,现在直接锁死在定义层。这降低了审计和新人融入成本,尤其是对于中小团队来说,避免了每次做报表都要追溯原始SQL的混乱。

但必须指出,这个设计的挑战不在技术,而在组织。正如评论所暗示的:当销售认为的“收入”和财务认为的“收入”根本不是同一个数字时,语义层只能帮你存储冲突,不能帮你消除冲突。产品目前支持团队级差异,这看似灵活,实则可能导致企业内形成互相孤立的指标孤岛,背离“统一”的初衷。本质上,这是一个强组织流程驱动的产品,如果你所在的公司连月报命名都统一不了,那Semantic Layer只会让你更快地跑出两套正确但矛盾的数据。

另外,AI引用语义层的执行效率也是一个需要验证的短板——尤其是当定义复杂嵌套SQL时,是否会导致查询性能下降?这些问题在实际场景中会比Demo更加致命。总评:方向正确,工具扎实,但别指望它能解决“人的问题”。适合已有规范指标体系的团队做自动化“加杠杆”,不适合从零搞数据治理的混沌组织。

查看原始信息
Basedash Semantic Layer
The Basedash semantic layer lets teams create reusable SQL metrics and models that AI can reference across chat, charts, dashboards, insights, and automations.
Hey everyone, Max here from Basedash. Today we're launching the Basedash semantic layer: reusable SQL definitions for AI analytics in Basedash. These definitions are reusable SQL queries attached to a data source. You can define "monthly recurring revenue", "activation rate", or "qualified pipeline" once, give it a reference name and description, and Basedash can use that exact SQL anywhere. This matters because AI is great at exploring data, but teams still need deterministic calculations for the metrics they run the business on. With definitions, the AI can build charts, answer chat questions, generate insights, and run automations while reusing the same approved SQL every time. We've been using this internally for metrics that show up across dashboards and reporting workflows. It removes the copy-paste SQL problem and makes the AI's work easier to audit. The semantic layer is available in Basedash today. PH community gets an extra week on the trial this week. Happy to answer anything.
2
回复

@maxmusing Congrats on the launch. Just a short question: when the AI uses your sematic layer definitions, how does it handle edge cases where the same metric might need slight variations across different contexts, say currency conversions, time zones or business units?

0
回复

Congratulations on your launch. I like the idea.

1
回复
0
回复

This feels practical for BI. AI can help people ask questions faster but the important business metrics still need one agreed definition. I like the idea of defining smth like MRR or activation rate once, then letting AI reuse that same logic everywhere. How do you handle cases where diff teams define the same metric slightly differently?

1
回复

@ada_johnsen great question. Each team can (optionally) have their own set of metrics, and you can even define a metric using another team’s metric as a base. We support team-level AI context to teach the AI exactly how your team works.

1
回复

Semantic layers seem straightforward until different teams start defining the same metric differently.

Have you found the technical challenge is the easy part, and the harder problem is getting organizations to agree on shared definitions in the first place?

0
回复

@samyak_sanklecha we’ve found that most great companies have centralized teams responsible for defining these kinds of metrics for the other teams.

0
回复

Wanted to add the why behind this one!

People love what AI does with their data right up until the AI gets something wrong. Which happens all too often with other tools, unfortunately.

That's why we're so focused around better context for AI agents, and why we think we're building one of the most accurate data agent platforms in the world today. Definitions take that even further. The metric gets written once, reviewed, and the AI reuses that exact SQL every time it touches it. So when finance and the AI both report revenue, it's the same revenue.

Give it a spin and tell us where it breaks :D

0
回复

@kris_lachance context is all you need

0
回复
#17
Gather
Save it once, never lose it again
95
一句话介绍:Gather是一款视觉灵感管理工具,帮助设计师将散落在浏览器、截图和书签中的设计参考统一保存,并通过自然语言描述(如“那张复古汽车广告”)快速搜索找回,解决“存得容易、找得难”的痛点。
Design Tools Productivity Artificial Intelligence
设计参考管理 视觉灵感库 搜索工具 截图管理 浏览器插件 第二大脑 设计师工具 AI提示词 自然语言检索 书签整理
用户评论摘要:用户普遍反映设计参考零散存储在Twitter书签、Figma、截图和Notion中,查找困难。积极评价其统一管理和搜索功能。用户建议增加MCP(模型上下文协议)以对接Claude等AI工具,提升引用效率。
AI 锐评

Gather切中了一个真实但被低估的痛点:设计师的“数字杂物间”。在Figma、Pinterest、Twitter、本地截图等多端间反复横跳找参考,确实是高频且低效的噩梦。其核心亮点并非存储——这已是红海——而是“以自然语言描述搜索”的查找逻辑。这比传统的文件夹分类或加标签更符合人的记忆习惯,尤其当参考量积累到几百上千时,这种“模糊检索”能极大降低心智负担。

但从产品现状看,它更像一个“精巧的原型”,而非成熟的解决方案。50条的免费额度太紧,几乎只是体验券;而4美元/月的Pro定价,在同类工具(如Eagle、Milanote甚至Notion)面前缺乏碾压性优势。用户期待的“MCP对接AI”功能目前仍在计划中,这意味着在“AI作为默认设计助手”的当下,Gather尚未形成闭环。

真正的价值在于:Gather能否从一个“图床+搜索器”进化为“视觉资产的启发引擎”。如果它能利用AI理解图片的设计风格、色彩组合、布局特征,并主动推荐关联参考,甚至反向生成提示词(prompt)供Midjourney或DALL·E使用,那它将不只是设计师的记事本,而是创意生产链中的关键节点。目前来看,它还差这口气。

查看原始信息
Gather
Easily find any design references! Gather is your second brain for visual inspiration — save screenshots, photos, and links, then search them the way you'd describe them to a friend: "that muted retro car ad."

Congrats on the launch, I'm constantly saving inspiration in X bookmarks (much like you!) and Notion pages but having it all in one central place (and visual, of course, I'm a designer at heart) will be such a gamechanger.

I would love to know what plans you have to make this even more awesome, maybe an MCP to pull references into Claude?

2
回复

Thanks @seandotexe! Yes, I was thinking about creating an MCP or a CLI you can use in any of those AI tools, which would make it even easier to use

1
回复

Hey Product Hunt, I built Gather to make it easy to find design references I was bookmarking on Twitter

For a long time, I saved design references in my Twitter bookmarks. It was a pain to find the one I wanted whenever I needed a reference for a friend, or to prove a point, or now, to give to an AI as a design reference.

What pushed me to create it was when I started copying it from Twitter, pasting on Claude Code and then explaining the image to the AI so that it can understand what I want from a design perspective. So I thought "there must be an way to make this automatic and easier"

So I built Gather, a website aimed at having references from multiple places at the same time where it's easy to add and easy to find them. With the chrome extension, it's a matter of right clicking on a reference, saving to Gather, and then later when you want to find it just talk to it "website with animation skills" and there you go, from there you can download it, copy it or copy a prompt for your AI.

Free tier is 50 references forever, Pro is $4/mo. It's the tool I now use daily to save all my references.

Share how you save your design references today and what's your workflow.

1
回复

@samu_monteiro I tried it, I liked it, congratulations on the launch

0
回复

Congrats on the launch!! 👏🏻
My design refs are a beautiful mess across Figma and screenshots. Turning that into something searchable is exactly what I needed.

1
回复

@munis_abbas thank you! Hope you enjoy using it

1
回复

I probably spend more time searching for old design references than collecting them 😄 Love the idea of turning a pile of bookmarks into something actually searchable. Congrats on the launch!

1
回复
0
回复
#18
Extella.AI
Agentic platform that evolves & builds reusable systems
90
一句话介绍:Extella.AI是一个自进化的智能体平台,通过四层记忆架构(规则、知识库、专家系统、加密存储)将用户的一次性任务转化为永久可复用的自动化系统,解决AI工具“用完即忘、无法沉淀能力”的核心痛点。
Productivity Artificial Intelligence No-Code
AI智能体平台 自进化系统 可复用专家库 域特定语言 多模型路由 本地优先 工具链集成 知识管理 自动化工作流 无代码开发
用户评论摘要:正面:非开发者用户称赞其“无需命令行即可构建永久自动化专家”,酒吧老板用它整合CRM、薪资、财务等业务。质疑集中于:初始学习成本高、记忆冲突处理机制、本地/云端模型路由的自主控制权。团队回应称冲突会显式暴露由用户决策,默认所有专家本地运行。
AI 锐评

Extella.AI的叙事极具野心:宣称自己是“第一个真正自进化的AI”,且技术细节(CSPL、FPGA合成、四层记忆)确实比市面套壳产品硬核。但Lauch首日仅90票和评论区的强烈防御性回帖(尤其长篇对比Claude Code)暴露了两个问题:第一,产品目前更接近“高级自动化框架”而非成熟平台,用户需要投入大量时间“调教”系统才能感受到价值,这与多数人期待的“开箱即用”相悖。第二,所有竞争优势都建立在用户深度参与构建生态的前提下(如自建CSPL、自定义Router),这在早期阶段是极高的使用门槛——评论中甚至有用户因不理解“目标设备”概念而困惑。

真正的隐忧在于:Extella宣称的“10x-1000x效率提升”本质上是工具链的帕累托优化,而非模型能力的根本性突破。其核心价值是为高阶用户(开发者、技术极客)提供将AI能力“资产化”的生产力工具,但普适性存疑。如果CSPL和自建专家库无法形成足够大的社区生态,产品很可能陷入“核动力自行车”的窘境——架构领先但操作复杂,最终只服务于少数建筑工。值得肯定的是,其对本地优先和数据主权的坚持,在当下“所有数据都上云”的AI狂潮中,确实切中了企业级用户的真实痛点。

查看原始信息
Extella.AI
Extella is an agentic AI platform that self-evolves. It remembers what works and reuses it, so each task runs faster and cheaper. That memory lives in four layers: Rules adapt to you, Concepts build a knowledge base, Experts turn work into automations, and a KV store keeps your keys encrypted. Bring your own LLM, connect any tool (Slack, Notion, Gmail, GitHub), any GitHub library, ML model, or your own — all managed from a single workspace. By Day 30, it's a system only you have.

Hi Product Hunt,

I'm Vlad, one of the co-founders of Extella. Timur and I started down this road back in 2014 with a goal that sounds almost too simple when you say it out loud: build a calculator that could solve any human problem.

It's not a chatbot, nor an assistant. It's a calculator in the most literal sense — you give it a problem, it gives you back a correct solution. Getting there meant three things had to work together. It had to be modular, so it could grow without limits and build new abilities out of small pieces instead of being rewritten every time. It had to actually evolve, so every problem it solved made it permanently smarter. And it needed real intelligence on top, so it could understand a problem before solving it, pick the right tools, and get better at its own thinking over time.

That took the better part of a decade. Patents, university pilots, a long list of dead ends (we even tinkered with FPGA chips at one point because the hardware just couldn't keep up).

For most of those years the tech wasn't ready for what we were trying to do.

When LLMs finally showed up, all the pieces we'd been quietly assembling just clicked into place. Now Extella is that calculator.

One thing I'll be straight about, because most launches won't: Extella isn't magic on day one. It needs a little time to learn how you work and what "good" looks like to you. The first week, you're mostly teaching it. Around month two it's saving you real hours. By month six it's doing things you forgot you ever showed it, and your cost per task keeps dropping as your Expert library carries more of the load. You bring your own models, local or cloud, and Extella routes each task to whichever one fits, by cost and accuracy. Most AI tools wow you on day one and bore you by day thirty — this one's the other way around. Quiet at first, then it compounds.

It took us ten years to build a platform that continuously self-evolves and adapts to each user individually. Try it now on macOS 13+, Windows 11+, or Linux.

P.S. check the last slides for Product Hunt Community Perks. We are giving away 50,000 credits and launching the Build Challenge with a one-time opportunity to win some unlimited features for life. Details in the slides.
Sign up for the Build Challenge here: https://forms.gle/mtAauocmYzCcLnUW7

10
回复

What does your platform do that Claude can't?

5
回复

@perizatka_03 Claude Code is an excellent tool. But it's a session. When it's done, it's done.

Extella persists.

Every solution becomes a permanent Expert — a verified, reusable module callable from anywhere: chat, API, another expert, a robot. Claude Code writes code for you. Extella builds a growing library of capabilities that outlive every conversation.

Extella runs on anything.

Claude Code runs in your terminal. Extella runs on any device you connect — your laptop, a server, a Raspberry Pi, a robot. Through Targets, the same expert executes wherever the task requires.

Extella has memory that evolves.

Concepts accumulate knowledge permanently. Rules shape behavior. Through Sleep — fine-tuning — everything the system learns gets embedded into the model's weights. The next session starts smarter than the last one.

Extella is multi-agent with infinite context.

Claude Code is one model, one context window. Extella decomposes complex tasks across specialized agents recursively — each with a clean context, its own LLM, its own memory and rules. No context overflow. No ceiling.

Now the part no other platform comes close to: CSPL.

When Anthropic ships a new feature to Claude Code, an engineer wrote it, reviewed it, deployed it. You waited for it. You had no say in what got built.

In Extella, if the execution layer doesn't do what you need — you define a new one. Not a plugin. Not a config. A new CSPL container that teaches Extella how to interpret an entirely new class of experts.

Want experts that execute as Rust? Define a CSPL. Want a domain-specific language that describes a financial strategy, a molecular structure, a neural architecture, a CNC toolpath — and have Extella execute it natively? Build that CSPL. Now every expert written in that language runs correctly by construction. The language itself prevents incorrect code.

And this is where the 10x–1000x efficiency comes from.

A general-purpose language like Python carries the full weight of everything it can do. For any specific domain, most of that is irrelevant noise — and noise means tokens, means ambiguity, means errors, means LLM cycles spent on syntax instead of meaning.

A DSL built for one domain eliminates that noise entirely. If your CSPL describes a financial strategy, there is no syntax for doing anything other than a financial strategy. The model works only with meaning. No boilerplate. No misinterpretation. No retries.

The more constrained the language, the less the model needs to do — and the more reliably it does it. That's where the multiplier comes from. A 10-line DSL description replaces 300 lines of Python. A domain-specific instruction replaces five rounds of prompt iteration. At scale — across thousands of experts, dozens of agents, running continuously — that difference compounds into orders of magnitude.

At the extreme end, CSPL reaches hardware. Verilog through Extella doesn't just generate code — it synthesizes directly to FPGA and dynamically builds processor architecture per task. That's Patent US 12,099,462. The architecture adapts to the algorithm, not the other way around. The limit stops being software. It becomes physics: speed of light, transistor density.

This is the distinction that matters:

Anthropic updates the product for the user — a team decides what gets built, you receive it, you adapt to it.

Extella is updated by the user from a single request. You describe what you need. Extella builds the Expert, the CSPL, the Rules. The capability becomes permanent. The system is more capable than it was this morning — not because a company shipped a release, but because you had a conversation.

Extella is the first genuinely self-evolving AI
— not in the marketing sense, but in the literal engineering sense. The system's own components — memory, reasoning, execution, hardware interaction — are all implemented as Experts. Which means they all inherit the same three planes of evolution: the code grows more complex, the description evolves, the execution layer gets replaced with something better.

The system improves its own way of improving.

Every other AI platform has a ceiling defined by its developers. Extella's ceiling is defined by what is physically computable.

They are tools. Extella is the first system that builds its own tools — and then builds better tools for building tools.

3
回复

Congrats on the launch! This looks very promising. So does that mean that due to the architecture it basically.... doesn't forget anything? What if things get contradictory and/or confusing, how does the system handle it?

5
回复
@aigerim_nogaibayeva Great question. By default — in zero-state — the system remembers everything that matters through Concepts, which use semantic search to surface relevant knowledge for each task. It's structured, searchable, and grows smarter over time. Not a memory dump. But also not just a vector search. Here's what that means in practice: if the default memory logic doesn't fit your workflow, you don't reconfigure settings — you tell Extella. It evolves to match your logic through Experts, CSPL, and Rules. For example, we built concept_search_pro and concept_write_pro — custom memory experts with re-ranking, LLM synthesis, quality scoring, deduplication, auto-tagging, and semantic merging on write. None of that existed out of the box. We just described what we needed, and Extella built it — and now uses it as its own memory layer instead of the default one. That's the point: you can say *"make concept search smarter for my use case"* and the system will evolve that layer specifically for you. Want memory to behave differently for a specific task? Say so. The architecture — Experts, CSPL, Rules — is the mechanism by which Extella adapts itself to your logic, not the other way around. On contradictions in Rules: the agent surfaces them during interaction, and the user decides how to resolve them. You stay in control. The default behavior is a solid starting point. What it becomes is entirely up to you.
5
回复

I tried it and it's really impressive.

It took me some time to figure it out, but it works much better than Gemini for me.

I asked it to find an old autosave file that I had lost with specific changes in it, and it found it!
I would have spent hours doing it myself.
It seems that you can connect all your tools and file systems and use it as a personal operating system with memory.

I wonder if you are planning to make a mobile version, which would be more convenient)
Good luck with ur launch!

3
回复

@ella_bye Thank you! Finding lost files is one of those quiet wins that actually saves your day.

You can already use Extella from your phone via the web version and run Experts as long as one of your desktop devices is online. So it's usable today — you just need a machine running in the background.

A full mobile app is on the roadmap — where your phone will also act as a target device. 🙂

2
回复

I’m fine with the slow start if month two actually pays off. The routing piece is what I’d want to understand first: can I set a local-only policy for certain Experts, or does Extella decide between local and cloud models per task?

3
回复

@novamaker01 Both. And you control which.

You set the policy. Extella enforces it.

Every Expert runs on a Target — a specific device you've registered. You can pin any Expert to a local Target explicitly. That Expert will never leave your infrastructure, regardless of what the RL Router would otherwise prefer. Local-only is not a global toggle — it's a per-Expert, per-Agent, or per-Rule decision. As granular as you need.

At the Agent level, each agent has its own LLM configuration. You can point an agent at a local model — Llama via Ollama, Mistral, whatever you're running — and that agent will never make an outbound model call. Other agents in the same profile can use cloud models. They coexist in the same pipeline.

Rules and Concepts orchestrate the whole thing.

Rules are the highest-priority layer. The Router optimizes within the envelope you set, not outside it. But orchestration goes deeper than routing. Rules define when to switch models, when to escalate from a reflex to a reasoning chain, when to delegate to a different agent. Concepts carry the accumulated knowledge of what worked before — which model performed better on which task type, which pipeline failed and how it was fixed. That knowledge feeds back into every future decision. The system isn't just following your rules. It's learning from its own execution history and getting more precise about when to apply which resource.

CSPL enables auto model switching at the execution layer.

You can define a CSPL that dynamically selects its execution backend based on the task at runtime — local inference for latency-sensitive work, cloud for complex reasoning, FPGA synthesis for compute-bound algorithms. The switching logic lives in the container, not in a config file. Which means it can itself be evolved, versioned, and replaced without touching anything else in the system.

So in practice:

- Sensitive data pipeline → local Target + local LLM agent + Rule enforcing local-only → nothing leaves
- General reasoning tasks → RL Router picks the best available model, cloud or local, based on cost and accuracy
- Hardware-adjacent tasks → pinned to the device that has the FPGA or the sensors

The architecture was built for exactly this split — because in enterprise and robotics contexts, data locality isn't optional. You don't have to trust the system to make the right call on sensitive workloads. You just write the Rule, and the decision stops being a decision.

But step back for a second — because the routing question is actually the smaller story.

Claude Code is a session. When it's done, it's done. Extella persists. Every solution becomes a permanent Expert — verified, reusable, callable from chat, API, another expert, or a robot. Claude Code writes code for you. Extella builds a growing library of capabilities that outlive every conversation.

Extella also runs on anything. The same expert executes on your laptop, a server, a Raspberry Pi, or a robot — wherever the task requires. Memory accumulates permanently through Concepts. Through Sleep — fine-tuning — everything the system learns gets embedded into the model's weights. The next session starts smarter than the last.

And through multi-agent profiles, complex tasks decompose across specialized agents recursively — each with a clean context, its own LLM, its own memory and rules. No context overflow. No ceiling.

Now the part no other platform comes close to: CSPL and the 10x–1000x multiplier.

A general-purpose language carries the full weight of everything it can do. For any specific domain, most of that is irrelevant noise — tokens, ambiguity, errors, LLM cycles spent on syntax instead of meaning.

A DSL built for one domain eliminates that noise entirely. If your CSPL describes a financial strategy, there is no syntax for doing anything other than a financial strategy. The model works only with meaning. No boilerplate. No misinterpretation. No retries. A 10-line DSL description replaces 300 lines of Python. A domain-specific instruction replaces five rounds of prompt iteration.

At scale — across thousands of experts, dozens of agents, running continuously — that difference compounds into orders of magnitude: 10x, 100x, 1000x and beyond depending on domain complexity.

At the extreme end, CSPL reaches hardware. Verilog through Extella doesn't just generate code — it synthesizes directly to FPGA and dynamically builds processor architecture per task. That's Patent US 12,099,462. The architecture adapts to the algorithm, not the other way around. The limit stops being software. It becomes physics: speed of light, transistor density.

And everything above is the baseline. Zero-state behavior.

The system is adaptive by design and evolves along three axes simultaneously:

From errors and new knowledge.
Every failed execution, every corrected output, every new Concept added feeds back into the system. The Expert that failed gets updated. The Rule that caused the wrong routing gets refined. The Concept that lacked context gets enriched. The system doesn't repeat mistakes it has already resolved.

From users.
Every request that pushes the system past its current capabilities results in a new Expert, a new CSPL, a new Rule — permanently. One user's solution becomes infrastructure for every subsequent similar task. You don't file a feature request and wait for a release. You describe what you need, and the system builds it.

From researchers.
When a new paper drops — a better fine-tuning method, a new routing algorithm, a more efficient inference architecture — you don't wait for Anthropic to ship it. You implement it as an Expert or a new CSPL container. The Sleep expert gets replaced with the new method. The RL Router gets updated with the new algorithm. Research becomes capability in hours, not quarters.

This is the distinction that matters:

Anthropic updates the product for the user — a team decides what gets built, you receive it, you adapt to it.

Extella is updated by the user from a single request.
The system is more capable than it was this morning — not because a company shipped a release, but because you had a conversation.

Extella is the first genuinely self-evolving AI
— not in the marketing sense, but in the literal engineering sense. Every component of the system — memory, reasoning, execution, routing, hardware interaction — is implemented as an Expert. Which means every component inherits the same three planes of evolution: the code grows more complex, the description evolves, the execution layer gets replaced with something better.

The system improves its own way of improving.

Every other AI platform has a ceiling defined by its developers. Extella's ceiling is defined by what is physically computable.

They are tools. Extella is the first system that builds its own tools — and then builds better tools for building tools.

2
回复

@novamaker01 

Felix, good question — and it goes straight to the architecture.

By default, every Expert runs locally on the device where Extella Desktop is installed. There's no automatic "local vs. cloud" routing happening under the hood — the platform doesn't make that call for you.

If you want a hard local-only guarantee for certain Experts, you simply don't assign a target when running them. A target is the mechanism for routing execution to a different device (e.g., a remote server or second machine). No target = always local, full stop.

As for AI model choice (local LLM like Ollama vs. a cloud API) - Extella doesn't dynamically pick a model per task. You implement that routing yourself (e.g. when setting an agent in Agent Builder of the Pro Plan), which actually gives you more control, not less.

So to directly answer: yes, local-only policy per Expert is trivially enforced — it's the default. The platform gives you the routing primitives; the decision logic is yours.

0
回复

I'm not a developer. Tried Cursor, tried Claude Code — both assumed I know what a terminal is. And I mean, sure, I can open it. I've opened it. But then it says something like "activate your virtual environment" and I just stare at the screen waiting for it to explain what sin I committed to deserve this. Every tutorial starts with "it's simple, just run..." and then lists six commands that somehow break my entire computer.

So I'd get this beautiful, confident AI explanation of exactly what to do, copy the commands, paste them, get a red error, paste the error back, get new commands, get a new error, and somewhere around the fourth loop realize I have no idea what I'm doing or why. And even when it worked — next session, gone. Nothing carried over. The AI remembered nothing, my machine had nothing, I had a folder full of scripts I was too scared to touch.

extella was the first thing that felt like it was built for someone like me. I described what I wanted. It figured out the rest — dependencies, how to run it, whether it should run once or keep running in the background. I didn't pick any of that. It picked it. And now I have these things called "experts" that just sit there and work. I click, it runs. No terminal. No environment. No "did you remember to install..." — it handles all of that silently before it even starts.

I don't fully understand what's happening under the hood and honestly that's the point. This is the first AI tool that treated my ignorance as a constraint to work around, not a problem I needed to fix before I was allowed to use it.
Thanks!

2
回复
@artyom_kuznetsov Thanks!
0
回复

If you want to test Pro today – here's how it works:

1. Apply code PRO-PLAN to activate the plan first

2. After that, top up credits inside the app using code PHPROMO50 (first 50 users)

If you already have an Anthropic or OpenAI key, connect it instead of buying credits.
Either works, but the Pro activation has to go first.

1
回复

Great app. I used it to build an AI layer on top of my bar’s CRM. It automatically calculates salaries and bonuses for employees, generates financial reports, monitors suspicious receipt deletions in the accounting system, assigns tasks to the manager, and tracks their deadlines. It also helps with purchasing orders and monitors the markup.

Wishing the team the best of luck!

1
回复

@azizb This is exactly what Extella is built for — not just automating one task, but becoming the operational brain of an entire business. Thanks for sharing!

0
回复

This is a great companion when you are on a cross between creativity and technology!
I built some tools for my TTRPGs and automated script writing with this tool

0
回复
@vladimir_ananyev TTRPG tools and automated scripts — love it. Thanks for sharing! 🙌
0
回复

Intereting. Just curious how this handles conflicts or inconsistencies when it starts learning across different workflows. Good luck!

0
回复

@henry_habib Good question — and an honest one.

Conflicts surface explicitly, not silently. When two Rules contradict, the agent flags it during interaction. You decide which stays, or you merge them into a new Rule that resolves the tension. Nothing conflicting quietly accumulates in the background.

For knowledge across workflows, Concepts handle it the same way. When similar knowledge comes in from different sources, the system deduplicates and merges rather than stacking duplicates. If there's a genuine inconsistency — two workflows that learned opposite things — it gets surfaced for human resolution, not auto-resolved by the system.

But here's where it gets more interesting.

You can teach the agent exactly how to handle ambiguity — not just flag it, but resolve it. You define Rules for how to research a contested claim, how to fact-check, how to assess recency and relevance, how to weigh conflicting sources by importance. The agent learns your epistemology, not just your preferences. Once those Rules exist, the same judgment you'd apply manually gets applied automatically — every time, at scale.

You can also teach the agent to recognize the boundaries of its own knowledge. Rules and Concepts can encode a self-assessment protocol: when the agent hits a gap, it doesn't guess — it identifies what it doesn't know, flags it, and can trigger an automated research or learning workflow to close that gap. It learns to know what it doesn't know.

And then the final step: you automate the configuration work itself. Once the agent understands your standards for conflict resolution, fact-checking, and self-improvement — you build Experts that handle that process automatically. New knowledge comes in, gets verified against your criteria, gets merged or flagged, gets added to memory — without you touching it.

The goal of human involvement is not to be always present. It's to teach the system well enough once that it stops needing you for that class of decisions entirely. Every hour you spend teaching the agent reduces the next thousand hours of manual intervention to zero.

That's the trajectory the architecture is designed for. You start in the loop. You work yourself out of it.

Thank you for asking!

1
回复
#19
Intelligent Terminal
Windows Terminal with native agent integration
89
一句话介绍:Intelligent Terminal 是微软开源的 Windows Terminal 实验分支,通过原生集成智能代理(Agent)面板,自动检测命令错误、管理会话并提供上下文感知建议,解决了开发者手动复制粘贴错误信息到 CLI 代理的痛点。
Open Source Developer Tools Artificial Intelligence
Windows Terminal 智能终端 代理集成 错误检测 会话管理 开源 命令行工具 AI辅助开发 开发者工具 微软
用户评论摘要:用户关心独立应用如何保持窗口状态持久化,以及是否有计划将功能合并到主终端,暗示对安全性、合规性及使用便捷性的顾虑。建议明确分离原因与未来整合路径。
AI 锐评

Intelligent Terminal 的核心价值并不在于“多了一个 Agent 面板”,而在于它重新定义了终端与 AI 代理的交互范式:从“人复制错误 → 粘贴给代理 → 等待反馈”的断裂流程,进化为“代理直接实时读取 shell 输出 → 自动检测错误 → 主动提供修复建议”的连续闭环。这消除了开发者频繁切换上下文的认知负担,尤其在调试、日志分析、多步骤部署等高频场景中,效率提升显著。

然而,微软选择将其包装为独立应用而非主终端的内置功能,这种“实验性隔离”既是保险策略,也暴露了产品形态上的不成熟。评论中用户的底层担忧很犀利:如果始终无法合并,意味着底层架构或安全模型与主终端存在根本冲突,开发者将被迫在“更好用的独立工具”与“生态统一的官方终端”之间做选择,这反而可能削弱长期采用意愿。

另外,当前默认绑定 GitHub Copilot CLI 是一把双刃剑——它既能借微软生态快速冷启动,也暗示了对特定代理的深度适配,可能导致对其他 ACP 代理的支持流于表面。真正的价值不在于有多少代理可用,而在于代理能否真正理解终端上下文并给出精准建议,而这需要更完善的状态感知能力和会话管理机制。总体而言,方向正确,但距离“原生体验的智能终端”还有一段工程化距离。

查看原始信息
Intelligent Terminal
Intelligent Terminal is an open-source experimental fork of Windows Terminal with native agent integration. It adds an agent status bar, context-aware agent pane, automatic error detection, session management, and command palette prompts for ACP-compatible agent CLIs.

Hi everyone!

MS has an experimental fork of Windows Terminal. It adds a docked agent pane that has context from your shell output, can detect command errors, and lets you manage multiple agent sessions.

Currently defaults to @Github Copilot CLI but supports other ACP-compatible agents (connected mine with @Gemini). One notable thing is that they’re shipping it as a completely separate app instead of putting it into the main terminal.

Your agent reads the shell output directly, removing the need to copy and paste errors back and forth. Is that a good reason to try, or will you just stick to standalone CLI agents?

2
回复

@zaczuo Since it ships as a separate app, how do you handle window state/persistence across sessions and is there a path to merge the pain into the main terminal center, or is the separation intentional for safety and compliance?

0
回复
#20
Chloe by Close
AI agent built into your CRM who works leads for you
87
一句话介绍:Chloe是一款原生嵌入CRM的AI销售代理,能自动拨打线索电话、回复邮件、更新记录和跟进,解决小型销售团队人力不足、无法规模化覆盖客户触达的痛点。
Sales Artificial Intelligence CRM
AI销售代理 CRM嵌入 语音外呼 线索跟进 自动化销售 SDR替代 中小企业工具 销售效率 无集成痛点
用户评论摘要:用户反馈正面:称赞其语音质量高、节省SDR成本、显著提升呼叫量(65%-75%)。有人担忧重复联系已有人工跟进的线索,官方回应Chloe会识别并礼貌结束。整体无重大负面问题,但缺乏对复杂销售场景的深度讨论。
AI 锐评

Chloe的定位精准,但并非颠覆性创新,而是对CRM原生能力的一次务实补全。其核心价值不在于AI语音通话的技术突破,而在于“无集成”和“无额外席位费”——这直接击中了中小企业对SDR外包高昂成本和复杂集成栈的厌烦。从公开的340万通电话和300万美元管道数据看,Chloe在高频、标准化、低客单价的外呼场景(如线索初步筛选、会议预约)确实能快速产生ROI,相当于将企业从“48人全员拉通”的不可持续状态拉回至“4人+AI”的可复制增长模型。

然而,需警惕两个潜在天花板:其一,AI在复杂异议处理、深度关系建立上仍无法替代人类,评论中“避免品牌伤害”的担忧依然存在——一旦语调或语境出错,可能伤害早期冷启动的客户印象;其二,数据壁垒——Chloe只能调用Close生态内数据,若销售流程涉及跨平台线索源(如LinkedIn、邮件营销工具),其原生优势可能转为核心局限。本质上,Chloe是Close CRM的“粘性增强器”,而非独立的销售基建。对于已经使用Close的团队,它的性价比极高;对于未深度绑定Close的企业,这更像一台优雅的“内部囚笼”——好用,但换场地成本陡增。

查看原始信息
Chloe by Close
Other AI agents bolt onto your CRM, but Chloe is built into it. Chloe calls leads, replies to emails, enriches contacts, takes notes, and updates Close the second anything changes. No extra seat, no commission, no integration to babysit. Chloe isn't a tool you add. It's like adding 20 salespeople who never sleep, directly in your CRM. Chloe works across your entire sales motion — but her voice is where she really goes to work. Hear Chloe sell now.

I’m Steli — co-founder of Close. Been building CRM for small sales teams for 13 years. Today we’re launching Chloe, our AI voice agent. I can sell you a little bit on Chloe, but she’ll do most of it herself.

Here’s what our Chloe beta taught me: team size no longer has to determine how many conversations you can have.
Chloe has generated over $3M in open pipeline for our customers. Now, it’s your turn.

What makes Chloe different is that she lives natively inside your CRM. She pulls real context on every lead as she calls, qualifies, books, follows up, and updates your CRM all in one loop.

While big companies hire SDR armies, you hire Chloe.

Chloe is available to all U.S. Close customers today, takes minutes to set up, and we gift you AI credits in every plan.
If your #1 goal this year is to scale without scaling headcount, give Chloe a try. I think you’ll be surprised.

Happy to answer anything. Let’s go

:muscle::skin-tone-3:
7
回复
Howdy, PH 🫶🏼 Long-time builder @ Close CRM, first-time poster. Here to give you a look at Chloe, our new AI teammate: the highest-quality voice agent on the market with full, native CRM context and action. A big part of being a PMM is keeping a pulse on real customer feedback (glowing/harsh/anywhere in between). In the 6 years I've been at Close, Chloe is the runaway star of my customer interviews. Close users come to me buzzing, show me the money she's making, the meetings she's booking, the AI hesitance she suavely overcomes, and their excitement to have hot, qualified deals already on their calendars in the AM. Our Beta customers loved the insanely high quality of her voice and the tens of thousands of dollars they're saving on outsourcing SDRs. In the last 30 days, Chloe has placed 340K calls, had 50K live conversations, and has over $3M in open pipeline for our customers. Come join them. If you're a founder, owner, or sales leader whose 2026 goal is scaling with less headcount, Chloe is your next 20 sales reps. Abandon your standalone voice agents with wacky Zapier/Make setups and scripted convos. Chloe can call, qualify, book, follow up, and update your CRM with the context native to your CRM. No bolt-ons. Chloe is now available for all Close customers. Set her up in just a few minutes. We gift free Chloe credits to all our customers included on our plans (nope, we didn't raise subscription fees). Excited to see if you have the same *oh, shiiiiiit!* moments that hundreds of Close Beta testers had last month. Happy closin' 💰 Use PHCHLOE for 30% off for the first 3 months of Close until June 30th at midnight.
5
回复

I work at Close and I'd riot if they took our AI calling agent away from me

I'm not some random user. I do work at Close. I implemented Chloe within our own company, for our own sales and post-sales teams. I have skin in this game. But I have to tell you, it's the real deal.

Our best calling month ever came from mobilizing 38 people across the company. A massive all-hands effort you can't just replicate on any given Tuesday. Our actual sales team is 4 people. Chloe has increased our calling output by 65-75% consistently over our best calling month, without the extra effort, without hiring, without pulling anyone from their day-to-day. I'm not a scientist, but I've run the numbers enough times: more calls = more money.


We started with one agent. We now run 12 that call warm leads within minutes of signup, book onboarding appointments, even collect beta feedback from Chloe users (yes, Chloe called people to ask what they thought of Chloe). We're building agents to catch dropoff moments inside the product. None of it required hiring. All of it freed our team up to focus on what we actually want them doing: building relationships.

The fear with AI calling is always "will it feel robotic, will it damage our brand." What I've watched instead is our human team getting more focused on conversations that actually require a human. We call our warm leads faster, more consistently, and without burning out our team. It's not complicated. It just works.

1
回复

@liz_stephany1 Chloe 🤝 Close's real sales team. Always our own best power users!

0
回复

It's a nice concept. What happens when Chloe reaches out to a lead already talking with a human rep?

1
回复

@dhiraj_patel5 The great thing about a native voice AI agent in a CRM like ours is that there are no silos between everyones customer communication :) But if that ever happens, it would be the exact same scenario as a sales rep calling a prospect that is already chatting with another account executive. The prospect would tell this to Chloe, and she would acknowledge it and make a note in the CRM and wish the prospect a lovely day :)

3
回复