Product Hunt 每日热榜 2026-07-22

PH热榜 | 2026-07-22

#1
Humalike x Hermes
Social intelligence plugin for Hermes Agent
451
一句话介绍:Humalike x Hermes 是一款为Hermes代理注入社交智能的插件,通过一条命令即可让AI在Slack、Telegram和WhatsApp等群聊中自主判断何时发言、模仿群组语气并记住成员背景,解决AI代理在社交场景中“有问必答、像机器人一样尬聊”的痛点。
API Developer Tools Artificial Intelligence
社交智能AI AI行为层 群聊代理 对话节奏控制 个性化AI AI插件 Hermes代理 群组动态适应 AI记忆上下文 AI静默决策
用户评论摘要:用户高度认可“何时发言”的核心价值,关注行为层与模型解耦的架构优势。主要疑问集中于:多用户记忆边界如何管理(防止隐私泄露)?静默决策的具体机制和延迟成本?是否支持按群组自定义话痨程度?有用户建议暴露轮换/打断处理API用于语音场景。
AI 锐评

Humalike x Hermes的野心不在于让AI更聪明,而在于让它更“像人”。这恰恰是当前AI应用最被低估、也最难以攻克的鸿沟——当模型智商已够用,行为情商才是决定用户是否愿意与AI长期共存的胜负手。产品精准切中代理在群聊中“话痨、抢话、没有节奏”的社交无能症,用“决策是否发言”这一轻量级门控取代“全量应答”,从架构上就比传统全知全能的Agent更务实、更省钱。

但务必要警惕其宣传中的软肋:多用户记忆边界管理目前“没有绝对解决方案”。在家庭、团队等混合敏感信息的群聊中,一旦AI把私聊里的吐槽放到工作群复读,信任将瞬间崩塌。这或许是比“何时发言”更难处理的地狱级难题——社交智能的核心不仅是时机与语气,更是场合与权限。另外,SOC 2认证虽在推进,但用户记忆和消息如何路由、是否完全可控,需要更明确的技术文档背书。

总体而言,这是AI从“工具”走向“伙伴”迈出的正确一步,但别急着把它变成你不该说的真话的“局内人”。真正的社交智能,既要知道什么时候说话,更要知道什么话不该说。

查看原始信息
Humalike x Hermes
One command gives your Hermes agent social intelligence. It decides when to speak, adapts to your group's tone and remembers who said what. Works in group chats on Slack, Telegram and WhatsApp.

Hey PH 👋 Martí here, co-founder of Humalike.

What is Humalike? The behavioral infrastructure for humanlike AI agents. The social skills your agents have been missing.

Today we're shipping the fastest way to feel what that means: a plugin that makes your Hermes agent fit in 1-1's / groups.

The problem
We run Hermes agents in our own Slack and Telegram. Brilliant at tasks. Painful to use in groups and treat is as a companion. It answered every single message, talked over people, spammed 10 lines when one was enough. Everyone knew it was a bot instantly. That's not a model problem, it's a behavior problem.

- What the plugin does: Decides, when to jump in and when to stay quiet
- Paces replies like a human: typing speed, pauses, 1-3 short messages instead of a wall
- Learns how your group talks and matches the tone
- Remembers who people are and what matters to them

One command to install. Works in group chats on Slack, Telegram & WhatsApp.

Where it shines
👨‍💻 Coworker: steps up when it can actually help
🍻 Group chats: no longer the awkward one in the room
👨‍👩‍👦 Family chat: reacts to the puppy photo like everyone else
🤝 Friend: knows you well enough to say no

What we'd love from you

Install it and tell us the when you had the "aha" moment (if it didn't, that's the feedback that matters most).
We'll be here all day reading everything!

Backed by the first investors in ElevenLabs, Revolut & more.
Still the same tiny 🇪🇸×🇵🇱 team, & still not sleeping much :))

20
回复

@mcarmonas Congrats Martí 👏

The “when to speak and how to speak” layer is underrated.

We’re doing something similar for content — teaching the AI your real voice so it doesn’t sound like every other tool.

Best of luck today!

0
回复

@mcarmonas congrats Martí - wish you and the team all the success in the world.

4
回复

@mcarmonas Congrats on the launch. In groups where people have mixed expectations, how does Humalike decide which social role to adopt, and can users teach or correct that behavior over time?

1
回复

I've been thinking about this a lot lately , AI keeps getting smarter, but not necessarily better at interacting with people . Love that Humalike is focusing on the behavioral layer. Feels like a missing piece for truly useful AI agents.

8
回复

@james_anderson77 :))))) tysm James! Happy to chat a bit more in our Discord community https://discord.gg/7bZFjm9aHH

0
回复

Congrats on the launch! The behavior layer being separate from the model is interesting. Everyone tunes what the agent says, few teams touch when it should shut up. Are you guys biasing it toward silence by default? Socially that feels right, but the failure mode I'd worry about is probably the false negative, the message it should've jumped on and skipped instead

5
回复

@artstavenka1 Thanks for the question! 🫶

2
回复

@artstavenka1Hey Art! These are some sharp insights you've got here. I'd say it's way harder to make agents biased toward silence, false positives are much harder to get rid of than false negatives, and they make the agent super annoying.

I'm here if you have anymore questions :))

1
回复

Hey!! Maks here, Founding Engineer at Humalike.

All you need to give your agent social intelligence is one command, that's it. First credits are on us :))

It works in group chats on Slack, Telegram and WhatsApp.

Feel free to ask any questions. I'm here to help :))

4
回复

Hey everyone, I’m Mateusz, founding researcher at Humalike.

This launch is an easy way to see our work in action.

We help the Hermes agent:

  • know when to speak,

  • fit naturally into a group,

  • behave less like a bot.

Excited to see how Hermes users experience it.

4
回复

Hey! Ignacio here, Founding Product Engineer at Humalike.

We encourage you to integrate humalike into your Hermes agent and watch it's performance improve immediately in social scenarios. Trust me, you won't want to go back to your old agent behaviour. ;)

P.S. Enjoy your free credits on sign-up!

4
回复

The "decides when to speak" gating is the actually hard part here — an agent that stays quiet 90% of the time is worth more than one that answers every message, but that means the plugin holds group context the base Hermes agent doesn t. Where does the "who said what" memory live: inside my own Hermes runtime, or a hosted Humalike service the group messages get routed through? And is the speak/stay-silent decision a model call on every message (latency + cost per turn) or a lightweight local classifier?

3
回复

@noctis06 Hey Noctis, these parts live on Humalike side. Both memory and classifier are on our side for few reasons:
1. ease of integration
2. easy way to ship improvements
3. the Hermes plugin uses Humalike APIs that are transferable from product to product. That means if you ever wanted to create your own agent you can easily use Humalike APIs and get similar behavior as Hermes x Humalike

Cost per turn is low from what we tested, when you sign up you get 10.000 free credits which should be enough for months of normal usage for hermes agent.

Let me know if you are concerned about something else specifically:))

1
回复

Congrats on the launch @marticarmona! Decoupling the behavioral layer from the LLM core is definitely the right architectural move for conversational agents.

Because the plugin remembers 'who people are and what matters to them' in shared group settings (Slack, Telegram, WhatsApp), how does Humalike manage multi-user memory boundaries? If User A shares context in a 1-on-1 thread with the agent, how does the underlying knowledge graph prevent that memory from surfacing in a shared team or family group chat where User B is present?

3
回复

@franz_brionesFor now, the underlying memory doesn't prevent it. We want the behavioral layer to be able to correctly assess what information should be shared and what shouldn't :)) It's, of course, a very hard problem and we don't claim to have an absolute solution.

2
回复

@franz_briones Tysm for the support Franz!

0
回复

A suggestion from someone who works with voice agents a lot: it would be great to see a turn-taking or interruption handling module exposed through the API. Right now most stacks handle barge-in in a pretty mechanical way, so a behavioral layer that knows when to pause, yield, or gently hold the floor would be a really meaningful addition.

3
回复

@feride178675 Yep, exactly! That's what we mean by behavioral layer, and turn-taking API is acutally one of the 7 APIs we have available, although it's not tailor-made for voice. If you have use-case for voice where you need humanlike turn-taking, let's connect and chat:))

1
回复
0
回复

Love your product, i have a question how you optimize the token usage ?

3
回复

@kartikmalik 🫶🫶

0
回复

@kartikmalik Thanks Kartik! The optimization is already in how the turn-taking component works: a tiny "should I reply?" check runs on every message, and the full pipeline fires only when the agent actually has something to say!

2
回复

@kartikmalik Appreciate it!

0
回复

The turn-taking API is the part that got me. Knowing when to stay silent is honestly the harder problem than knowing what to say. Curious how you handle group dynamics where multiple people are typing at once, does the model factor in who's been most active? Congrats on the launch, excited to see where this goes!

3
回复

@vanmurten_eth Alexey! You're right, this problem is super hard. We handle it like a real human would, which means that as a new message comes in, the agent looks at the whole conversation and decides whether it should reply to any message from the conversation, taking the new message into account as context.

I am happy to answer any other questions you might have :))

2
回复

@vanmurten_eth Thanks for the supp Alexey!

0
回复
0
回复

Congrats @mcarmonas & team. What was the main challenge you wanted to solve when building Humalike?

3
回复

@hamza_afzal_butt tysm! Social skills & proactiveness!

1
回复

Hello PH, I'm Mateusz co-founder and CTO @ Humalike. I will be active here to answer any question! 🙋‍♂️

We've been dogfooding Hermes x Humalike plugin for 2 weeks now. Here are few use-cases where I see the most potential:

  • Gangprompting

  • Personal assistant (like Poke)

  • Agent for personal group - friend or family

It's possible thanks to our internal research team combined with extreme speed of our engineering. We would love to hear your feedback!

SOC 2 in progress 🔒

3
回复

@mateusz_jacniacki Had the same thought recently , our Ai agents can solve complex tasks but still struggle with basic human interctions . Love seeing someone tackle this problem head on . Best of luck with the laucnh.

1
回复

Does Humalike adapt to different group dynamics over time, or does each workspace start from a predefined behavioral profile?

2
回复

@tarqiya_forgah Indeed it does adapt! We continuously update its understanding of the group through components like Norms, Social Memory, Theory of Mind, and Social Signals.

That lets the agent pick up jokes, understand what matters to different people and align with the local behavior.

3
回复

Love the concept. One question: can developers customize how talkative the agent should be depending on the workspace or group?

2
回复

@nayan_joshi Thanks! Yes! Talkativeness is part of the agent's personality, and each workspace/group runs its own, so the same agent can be all business in one room and playful in another.

And even without touching anything, thanks to social learning API, it reads the room and matches each group's rhythm on its own: quiet team, quieter agent.

1
回复
0
回复

so you're giving soul to hermes??

2
回复

@jiteshghanchi Hahah, actually good that you ask, we should clarify.
We add whole new layer for turn-taking: decide when to speak vs stay stilent, that works well in group chats. We change how hermes acts when he is interrupted. We change the way he responds (short natural burst messages instead of long report-like messages). This is engineering work, not just a prompt.

We add Social Memory -> another layer of memory focused on people, their preferences, humor and opinions.

We add social learning that fires every few messages -> Hermes picks up new jokes, style and norms without you having to touch soul[.]md at all. It happens automatically many times during one chat.

Annnddd yes, we also improve your Soul based on our experience building 4 different humanlike agent products:))

You should give it a try, the jump you can make with Soul alone is 1.5x, with everything we got it's 10x

1
回复

@jiteshghanchi 👀👀👀

0
回复

Does the memory work across different group chats or is it per-group? Curious how much conversation history it needs before it starts picking up tone shifts.

2
回复

@talhakhalidmtk One memory for all of your groups, platforms (discord, telegram, slack, whatsapp...) and threads. Give it a try and push it to the limit

1
回复

Congrats on the launch. Social intelligence gets useful only when the signal is explainable. When Hermes surfaces an insight about a person or community, can users inspect the evidence behind it and correct the signal if the system gets it wrong?

2
回复

@yaroslav_stelmakh Hermes is the definition of auto-improving agent, it has memory and learns new skills. Humalike plugin also uses this information to adjust behavior, and even makes it stronger thanks to Social Learning API used under the hood.

0
回复
0
回复

Finally someone tackling the awkwardness problem head on. Tried the API and the proactivity cues feel way more natural than the stiff reminders my agent used to send.

2
回复

@savazverjtjn :)))) That's great to hear!

0
回复

the 'decides when to speak' bit is the whole game imo. most agents in a group chat have zero restraint and just talk over everyone

2
回复

@alex_watson2110 100%.. it's unusable by default. Give our plugin a try :))

0
回复

@alex_watson2110 Thanks for the comment Alex! I'm here if you have anymore questions

0
回复

I appreciate that this tackles a problem many developers notice after deployment. Technical accuracy alone rarely creates comfortable interactions. Will there be tools for testing personality consistency before releasing agents?

2
回复

@john_michael31 Hey, yes! We have an audit page where you can paste transcript of any agent (powered by Humalike or not) and you will have detailed analysis of how it performs.

We also have CLI for it if you want to run this analysis on recurring basis e.g. once per week!

1
回复

I want to know whether the hardest part was making the agent more human, or making it less eager to participate.

1
回复

@reda_roqai_chaoui Hahaha that's a good one! It depends a lot...
We built a AI Community Manager in the past that, if it was properly guided (through channels) and with a lot of constraints (behavior), the hardest part was how to make it human (infinite issues). Now, if you are building something with a lot of constraints and it's against product nature (having constraints is bad), making it less eager to participate is so hard.

1
回复

congrats on the launch. the "who said what" memory question that hasn't come up yet: since it runs across Slack, Telegram and WhatsApp, if the same person messages the agent under different handles on two of those platforms, does Humalike link that back to one identity, or does each platform get its own isolated memory of that person? seems like it'd change a lot for someone who's in the same group across multiple apps.

1
回复

@galdayan Hey Gal, good question! Today: "no", guessing two accounts are the same human is risky (hard to know who is who across platforms). But, if you build with our APIs and "tell us" (mark) ids/person, memory unifies across platforms :))

0
回复

"social intelligence" is the part I'd want a concrete definition of before buying in - is there an actual benchmark or eval suite behind that claim, or is it mostly vibes-based (agent seems less awkward in a demo)? proactiveness is even harder to score objectively since an agent that's too proactive just becomes annoying. curious what metric you're actually optimizing against internally.

1
回复

@omri_ben_shoham1 Hey Omri! Social intelligence is a big statement and we are aware of it. We could argue there is no "good or bad" when judging behavior, so it's really hard to measure. As of today, we are mesuring it "vibes-based", and also working on benchmarks such as https://arxiv.org/abs/2606.14600. We (& all devs/researchers) have a long way to go on delivering proper social intelligence, but we love the problem and we are building/doing research in that direction!

0
回复

decides when to speak is the whole game for me. i've killed more than one bot in a group chat because it replied to everything and drove people nuts. staying quiet at the right moment is way harder than having something to say. how does it learn a specific group's threshold, or do you set it?

1
回复

@terminal_candy Hey Peter!! Totally... we had the same problem. We realized how big of a problem this was when we were building an AI Community Manager in the past. Making an AI moderate, do support and have genuine conversations with a group of people is really really hard.
About groups norms, we have an API called Norms that exists just to solve that problem, + combining it with other APIs such as social intelligence and theory of mind, we are able to make agents adapt really well so far!

0
回复

Splitting those two apart is fair, and the interruption handling is the bit I want to poke at. When we did this we cancelled the in-flight generation whenever a new message landed, and the annoying part wasn't the cancel, it was deciding whether to restart from the new context or drop the turn, because on a busy channel restarting meant the agent never finished a sentence. Does yours re-run the decision after an interruption, or stay quiet for that turn?

1
回复

@dipankar_sarkar We've tried "everything" and so far the best way to approach this is re-running the decision after an interruption. There are ways to change this and it's valid if the community or chat is really active. Happy to go deep into it and share our experience :))

0
回复

Congrats team! 🚀 Love that you’re treating social skills as infrastructure rather than a prompt-engineering problem. Quick technical question: how does the Social Signals API detect things like typing pauses and deleted reactions? Does it need platform-level event hooks like Discord or Slack webhooks, or can it infer these from message streams alone? Curious how portable that is across stacks.

1
回复
1
回复

What I'd want to know is how long the decision takes. In a live group the window where a message is still useful is a couple of seconds wide, and every should-I-speak check we tried added another model call, so by the time it came back yes the thread had moved on two messages and the reply read as odd. Is the gating a small fast classifier sitting in front of the main model, or does the same model produce both the decision and the message?

1
回复

@dipankar_sarkar Hey, I think these are two separate things:
1. actual latency for making a decision
2. "group has moved on" problem

We solve 2 really well by handling interruptions and making sure agent doesn't respond to old messages. About 1, yes, it adds latency, and it's not perfect. In our testing it's not a deal breaker in chat, LLMs are fast enough (compared to humans that acutally have to type each message). I encourage you to give it a try and see how you like it, we would love your feedback!

1
回复
When will you guys support discord
1
回复

@dake_zhang1 Dake! You can build a Discord Bot with Humalike APIs in 10min. I have an old Discord community were I sold software (with thousands of users) and it runs completely in autopilot with a bot that's powered by Humalike :))

0
回复

Can this be used with other agent frameworks like Eve or OpenClaw?

1
回复

@gilad_tsehori Hey, humalike APIs are easy to build with. You can connect any kind of agent to Humalike. We don't have out-of-the-box integrations with Eve or OpenClaw though, let us know your use and let's chat!

0
回复

@gilad_tsehori Not yet, Hermes is first integration for now. If people find value on Hermes x Humalike, really soon it will be on Eve or OpenClaw

0
回复
#2
Migma AI
AI runs your email marketing. Better with every send.
409
一句话介绍:Migma AI是一个AI驱动的全自动邮件营销平台,从策略规划、内容生成到兼容性渲染、发送及优化,一站式解决邮件营销中跨客户端渲染不一致、个性化和合规性等痛点。
Email Analytics Marketing
AI邮件营销 智能营销 邮件自动化 邮件兼容性 Zinn标记语言 个性化营销 本地化 营销自动化 SaaS B2B
用户评论摘要:用户认可其“减少上下文切换”和端到端自动化。核心疑问集中在:Zinn引擎是否仅靠保守标记规避渲染问题,还是真实测试?小编辑会消耗配额吗?能否导出HTML至其他ESP?对合规和实时预览有强需求,赞赏非复制生成而是一体化工作流。有用户提到“约1/12生成需重试”。
AI 锐评

Migma AI的野心不在于“又一个AI写邮件工具”,而在于试图重构邮件营销的生产关系。它的核心价值命题非常清晰:用AI Agent接管从“理解业务目标”到“生成兼容各客户端的营销邮件”再到“基于数据闭环优化”的全链条。这切中了邮件营销的两个实质性痛点:一是“渲染矩阵”问题——无数工具生成的邮件在Outlook或Gmail上“翻车”,Zinn作为自研标记语言是对行业积弊的正面突进;二是“端到端效率”——Klaviyo、HubSpot的做法是给人类操作员加上AI助手,Migma则试图成为那个操作员本身,显著降低人力成本。

然而,“All-in”的另一面是“All-or-nothing”。评论中用户对于Zinn渲染机制、配额消耗策略、API集成细节的密集追问,恰恰暴露了这套自动化逻辑最脆弱的一环:信任。营销人员仍需保留“最后一公里”的否决权和控制力。Migma声称“人类批准”,但若生成频率与配额管理冲突(如小编辑消耗配额),或自动化工作流处理异常场景的韧性不足(如1/12的重试率),都会从“效率神器”滑向“需要兜底的麻烦”。此外,其能否吸引已深度绑定Klaviyo/Mailchimp的存量用户迁移,取决于技术层(Zinn的兼容性是否能超越主流库?)和成本层($49/20万封的长期定价是否具有竞争力)。

一句话总结:Migma在技术层面(Zinn)和产品哲学层面(Agent-first)都足够激进,但颠覆一个依赖稳定和信任的行业,光靠“快”是不够的,它还需要证明自己的“稳”。下一次公关危机,可能就是某个大客户CEO发现自己的促销邮件在Outlook 2003上成了一堆乱码。

查看原始信息
Migma AI
Migma creates full campaigns before you ask. It understands your goals and events, creates every email, personalizes and localizes it, and renders consistently across inboxes. It handles sending, preferences and compliance, then learns what drives revenue.

Hi Product Hunt,

Over the last year, Liam and I have been obsessed with one idea:

Your email platform should understand what your business is trying to achieve, carry the work all the way to the inbox, then learn what works.

That is what we built Migma to do.

Give Migma your brand, audience and goal.

It plans the campaign.

Creates the complete sequence.

Makes the right versions for different audiences and languages.

Then sends it.

ChatGPT, Claude and other AI tools can create emails that look good in a browser, but break when they reach real inboxes. They don’t understand compatibility, accessibility or compliance in the way email requires.

Since our last launch, we learned that compatibility had to be the foundation.

React Email, MJML and other libraries were still failing to produce emails that looked consistent everywhere.

So we built Zinn, our own markup language.

It generates emails quickly and efficiently while keeping them compatible across email clients, even Outlook 2003.

We didn’t build it just to create more emails.

If you’re paying for a platform, the emails should be worth sending and built to convert.

That’s our mission, and why we watch every metric.

Migma personalizes and localizes each version. Your audience in Spain receives the email in Spanish. Customers in Canada can receive it in French or English based on their location and preferences.

Migma tracks which emails drive clicks, purchases and revenue, then uses those results to improve future campaigns.

Your first campaign should never be as good as your last.

Thanks to our partnership with Cloudflare, we’re also able to offer sending infrastructure you can trust.

Migma handles domain setup, warmup, compliance, unsubscribes and customer preferences.

You can work directly in Migma or from your AI agent, Slack or Telegram.

All of this is also available through our API for marketing and transactional email. We’re already partnering with other platforms to bring Zinn into their products.

Since our last launch, we’ve raised a six-figure pre-seed round and started growing from a team of two brothers into a bigger team.

Today, we’re offering Premium at 50% off for the first 200 customers forever.

You can create and send up to 200,000 emails per month for $49, and keep that price for as long as you stay subscribed.

We wouldn’t be here without your early validation, feedback and support.

Thank you,

Adam and Liam

35
回复

@adam_lab Amazing !!

I absolutely love it. So excited to be part of this journey.

Migma has just revolutionized email marketing.

5
回复

@adam_lab This is really huge!! Been watching Adam and Liam build this for months and the Zinn markup piece alone is a game changer, finally an email tool that doesn't break the second it hits Outlook. Compatibility + personalization + actual localization in one place? That's the SaaS email marketing has been missing.

5
回复

@adam_lab The future of email marketing is actually here!! Seeing it on PH today feels surreal. Thank you all so much for the early support... genuinely means everything to us right now.

2
回复

Can we export raw zinn markup or compiled production HTML straight to our own external ESP via API? huge congrats🙌 on the launch @liam007

5
回复

@priya_kushwaha1 Yep, compiled HTML, absolutely. Raw Zinn markup, not at the moment.

You get clean, email-ready HTML with merge variables already in the native syntax for Klaviyo, Mailchimp, HubSpot, Brevo, or Omnisend, so you can drop it straight in without having to rewire personalization.

You can push it through the native connectors or pull it via MCP or a webhook.

And the best part is that the HTML is built to render properly across Gmail, Outlook, Apple Mail, dark mode, mobile, and even older Outlook versions, so you know it’ll look right before you send.

Would genuinely love to hear your feedback once you’ve tried it 🙌

3
回复

I like products that reduce context switching. This feels like a platform designed around that idea. Great work!

4
回复

@1mirul Thanks so much! Cutting context switching was a big part of the design, glad it comes through.

0
回复

Migma feels strongest where most AI email tools fall apart: not the copy generation, but the render and compliance layer. Zinn is the interesting bet here because email HTML is still weirdly unforgiving.

Curious how much testing happens before send. Is Migma validating against real inbox/client previews, or mainly compiling into conservative markup that avoids the usual Outlook, Gmail, dark mode, and clipping issues?

4
回复

@aditya_harish_2002 You’re spot on, email HTML is still surprisingly unforgiving.

Zinn handles the compilation side, but Migma also checks the campaign before send: real inbox and device previews, dark mode behavior, Outlook and Gmail quirks, clipping risk, broken links, and compliance essentials like unsubscribe and sender details.

So it’s not just “generate and hope.” You can review the actual output, catch issues early, and only send once everything looks right.

Would love for you to try it with one of your tougher templates and tell us where it breaks 🙌

1
回复

Really cool product - congratulation on the launch. In my experience, the hard part in email isn't just the generation, it's the render matrix. Outlook's Word engine, Gmail clipping at 102KB, dark mode inverting your PNGs. Curious whether the "proprietary engine" means real client testing or just conservative table markup.

4
回复

@aidan_codefox That’s exactly what our Zinn engine is built to solve.

It goes beyond conservative table markup. It handles tricky cases like dark mode color inversion, especially in Gmail on iOS and Android, and you can also run real tests across real devices in Migma before sending to make sure everything renders correctly.

Really appreciate you bringing this up, the render matrix is one of the hardest parts of email, and it’s something we’ve spent a lot of time on.

Give it a try and put it through its paces. I’d genuinely love to hear what you think 🙌

1
回复

The idea of carrying the whole workflow from campaign planning to the actual inbox is probably the strongest part of Migma for me. It feels much more useful than another tool that only generates copy. How much control do marketers keep before a campaign is sent, and can they require approval at specific stages?

4
回复

@nico_mandera You keep full control. Migma gets everything ready, but a human approves it before it goes out.

Once approved, you can send it right away or schedule it, and Migma will handle the send automatically at the time you choose.

You can also review, comment, edit, or pause anything along the way. Migma just takes care of the busywork in between 🙏

3
回复

Brand consistency is always hard at scale. Looks like Migma is tackling that problem in a practical way.

3
回复

@monir_ Appreciate that 🙏
It's actually the first layer we solved early on, so marketers spend time driving revenue instead of fixing things they shouldn't have to worry about.

Try Migma for free today.

0
回复

klaviyo and hubspot slapped ai email gen on this year too tbh. what's actually different here besides zinn?

3
回复

@levi_mitchell1 Fair question.
Zinn's part of it, but the bigger piece is that Klaviyo and HubSpot bolted AI onto tools built for a human to run. Migma was built agent first, so it plugs straight into your own project through MCP and can handle your campaigns end to end, not just draft copy for you to finish but the entire workflow.

Come try a full cycle and see the difference.

1
回复

ok so per a review, every little edit makes a new email version and burns your daily quota. fast iteration vs hitting the cap fast, real tension there. any plan to batch small edits so they don't each count separately?

3
回复

@owen_parker4 Good question, small edits are grouped, so every tiny change doesn’t burn a separate credit.

You can also use the visual editor for free, which makes it easy to tweak and iterate without eating into your quota.

Really appreciate you flagging this, and would love to hear how it feels once you’ve tried it 🙌

1
回复
Congrats on the launch @liam007! Being able to trigger and manage campaigns directly from Slack or Telegram sounds super smooth. Can we review and approve full campaign previews right inside those channels before they go live?
2
回复

@hannesh Yes exactly!

From Slack you can ask Migma what to send next, check the dashboard for results, or hand it a Figma frame and get back a ready email.

Come give it a try

0
回复

Direct emission is the ambitious version of this, so I'm curious what catches a bad token. When we tried it we ended up with a validate-and-retry loop, and roughly 1 in 12 generations needed a second pass, which is fine for a background job and painful when someone is sitting there waiting. Are you constraining the decode with a grammar, or letting it write freely and repairing after?

2
回复

@dipankar_sarkar The model writes Zinn directly instead of HTML, that's what keeps it fast and consistent across clients.

Come try it see how its super fast

0
回复

LETS GOOOOO!

2
回复

@yahia_bakour3 yoo, please watch your APIs, we may burn your servers.

1
回复

@yahia_bakour3 Yahiaaaa 🤗

0
回复

Super well thought out blend of email and AI. Really excited about this one. Congrats!

2
回复

@greggblanchard Migma <3 SendView thanks a lot for your support.

0
回复

@greggblanchard Thanks so much! Means a lot coming from someone who gets both sides of it.

Come take it for a spin.

0
回复

Upvoted, good luck guys! 🇩🇰 🤝 🇸🇪

2
回复

@philip_sorensen Thanks Philip! Let’s keep the big ones uncomfortable 😄

0
回复

One thing that would really help me is a side-by-side preview of how the email looks across different clients like Gmail, Outlook, and Apple Mail before I hit approve. It would save me from those surprise layout breaks that always seem to pop up after sending.

2
回复

@yiitqx9x Our engine already renders every email to match across clients properly, so the base is solid before anything else kicks in.
On top of that, Preflight lets you preview it on real devices and clients before you approve, so you catch anything before it ships.

Try it here for free: https://migma.ai/email-html-spam-checker

1
回复

The auto-generated campaigns idea is genuinely cool, but I'd love to see a side-by-side preview showing how each email renders across Gmail, Outlook and Apple Mail before I hit approve. Would save me from those surprise layout breaks on Outlook especially.

2
回复

@nihaln20560 Love this, and good news, we already have it 🙌 Preflight shows client previews (Gmail, Outlook, Apple Mail, dark mode, and more) before you approve, so you catch any Outlook weirdness before it goes out, not after.

Come check it out for free.

1
回复

Email marketing that actually learns from results is what most tools claim but dont deliver. How long does it take for the AI to start making noticeble improvements after first few sends?

2
回复

@abdurrahman_fakhrul Totally fair pushback.

Every send feeds data back in immediately, opens, clicks, what converted, so it's not waiting on a big batch to update. We don't have an exact "X sends" number pinned down, but people usually start noticing sharper segmentation and copy within their first few campaigns as it locks into what your audience responds to.

2
回复

Congrats on the launch! Migma’s focus on the full email workflow, from planning to rendering and compliance, feels much more practical than just another AI copy generator. I’m especially curious about how the system learns from approvals, edits, and skipped sends over time. How quickly does it adapt to a brand’s tone and sending judgment?

2
回复

@glebarios Thanks so much, this means a lot 🙏

Brand tone comes from what you feed it (your assets, guidelines, past campaigns) and it keeps refining that as you approve, edit, or skip sends, so the more you use it the closer it tracks how you actually write and what you're okay shipping.

Performance also feeds straight back in: opens, clicks, what worked, what didn't, all shaping the next campaign automatically.

Come try a few campaigns and you'll feel it get sharper send over send 📬

1
回复

How does Migma handle email deliverability and domain reputation when sending campaigns from a connected custom domain?

2
回复

@shahriardgm  Great question 🙌

When you connect your own domain, Migma handles setup and warmup for you, so it ramps up gradually and lands in the inbox instead of spam, meeting all the checks that actually help conversion along the way.

If you don't have a domain yet, you can send straight from Migma with a shared address, no setup needed. Or you can buy a new domain right inside Migma at an affordable price.

Come connect yours and try a send, would love to hear how it lands for you 📬

1
回复

the "learns what drives revenue" part is the interesting bit to me - how does it attribute revenue back to a specific email when someone gets touched by 3-4 emails in a sequence before converting? that's usually where email attribution gets messy even for humans doing it manually.

2
回复

@omri_ben_shoham1 You're right, when someone gets 3-4 emails before buying, giving one email the credit is guesswork. So we don't force it onto a single email.

We track which emails a person actually clicked, and spread the credit across those, leaning toward the ones they engaged with most. Opens we mostly ignore since Apple Mail fakes a lot of them.

The part we really trust is the pattern across lots of sends, not one person's path. If the day-3 email with one subject line keeps showing up before people buy, and another version doesn't, that tells us something real even when any single journey is messy.

That's the math no one has time to do by hand.

Best way to judge it is to see it live. Would love for you to try it.

1
回复

Congrats on the launch - really cool idea and looks to be a very smooth product. I have to ask, if clients have all their contacts and operations inside another CRM (think Zoho or GoHighLevel) that currently send their campaigns, how does it work? My clients would be keen to keep their core CRM systems running without having to duplicate data into a specialized email tool.

2
回复

@danielboardman Thanks, really appreciate that 🙌

Your clients can keep Zoho, GoHighLevel, or their current CRM as the core system. Contacts are imported or synced into Migma through native integrations, connected ESPs, or API/webhooks, depending on the setup.

So there’s no need to move the rest of their operations. Migma handles the campaign workflow, while the CRM stays the main source of truth.

Would love for you to try it with one of your client setups and tell us how the workflow feels.

0
回复

Looks like something we could definitely use something like this in the near future - do you have a strategy how to you handle sending reputation/content quality?

2
回复

@foobar_beer Absolutely, that’s a big part of how we think about the product.

Migma helps on three fronts: content and compliance checks before send, deliverability and warm-up guidance, and reputation monitoring through your connected ESP.

The goal is to help you avoid the common issues that hurt sender reputation, not just generate the email and leave the rest to chance.

Would love to hear what your current setup looks like when you get a chance to try it 🙌

2
回复

The line that stands out to me is that your first campaign should never be as good as your last. That learning loop is the real moat here, even more than the rendering work everyone is asking about.

My angle is the awkward one for it: I run email to a small, high-value B2B list in healthcare, a few hundred serious buyers rather than a consumer list. There a tone-deaf automated send costs a relationship I cannot re-earn, and the revenue signal comes back slow and sparse.

So, Adam, does the learning loop need consumer volume to work, or can it improve on thin data, where the win is one right email a month rather than a tuned high-frequency cadence?

2
回复

@clemente_lopez1 No. Volume makes the loop faster, but it is not required.

For a small, high-value list, the useful signals start before opens or purchases: what you approve, edit, reject, who you choose to send to, and when you decide not to send.

Migma should become more conservative in that setting, not more autonomous. It prepares the campaign, learns your judgment, and you keep approval before anything goes out.

The goal is not more sends.

It is one email a month that sounds like you, lands at the right moment, and does not burn a relationship.

0
回复

@clemente_lopez1 You've nailed the real constraint here. High-value B2B lists are unforgiving as one bad send to a spam trap or invalid address tanks your sender reputation instantly, and suddenly Migma's learning loop breaks because mailbox providers are already filtering you.

Before Migma learns what works, you need the list itself to be trusted. That's exactly where we built Mailthentic. We validate every address, flag risky patterns (typos, role-based, honeypots), and remove the deadweight so Migma's AI gets clean signal to learn from.

For your healthcare B2B scenario: clean list first → Migma learns on engaged buyers → revenue compounds. Skip the cleanup and even perfect campaigns get junked.

Happy to chat about integrating list verification into your workflow if you're interested.

0
回复

Congrats on the launch. The preflight checker that catches rendering issues across email clients and fixes them with one click is such a good detail, that cross-client mess is usually the most painful part of building emails. Curious how you're handling something like Outlook's rendering quirks under the hood, that engine has famously been a nightmare for HTML email for years.

1
回复

the render consistently across inboxes line is the part i'd test first. outlook has wrecked more of my campaigns than any bad subject line ever did. does it hold up in the older outlook desktop clients or mostly the modern web ones?

1
回复

@terminal_candy Outlook is exactly the client we built hardest for 🙌 Zinn covers the old desktop versions too, not just the modern web ones, down to Outlook 2003.


Come put it through its paces.

0
回复

Big congrats team. Love that it factors in the calendar, not just brand and audience. Timing is usually the part founders get wrong, so having that baked into the drafts is smart. Excited to see this one grow.

1
回复

@ben_kahan Really appreciate that 🙏
Timing is one of those things that's easy to get wrong, so glad it's landing right.

Come take it for a spin

0
回复

Does your platform heat up emails?

1
回复

@carlos_mendez16 Yes! Connect a new domain and Migma warms it up automatically, so you ramp up sending without hurting your reputation.


Come try it for free!

0
回复

The Zinn part is what I'd want to know more about. Every time we've asked a model to emit a bespoke format it drifts back to whatever it saw most in training, so ours kept quietly producing plain HTML where our own schema was required, and prompting never fixed it, a constrained schema did. Is the model writing Zinn directly, or does it emit a structured plan that a deterministic compiler turns into Zinn?

1
回复

@dipankar_sarkar Great question! The model writes Zinn directly, not plain HTML. That's a big part of why generation is so fast and stays consistent across clients.

0
回复

@dipankar_sarkar It took us years of engineering to reach this point.

0
回复

I love this – congrats on the launch. What signal does it actually learn from: opens, clicks, replies, or revenue?

1
回复

@derrickshowers Thanks so much! It's not just one signal, we look at a mix of engagement data, opens, clicks, heatmaps, and how people actually interact with the email, not just a single number.

Come give it a try for free.

0
回复

"Creates before you ask" is the real pitch — most email tools wait for you to know what to say.

1
回复

@ringo_td5 Love how you put that 🙏 That's exactly the shift we're going for, Migma doesn't wait around for a brief.

Come see what it drafts for you

0
回复
#3
Buzzy
Your creative AI co-director
389
一句话介绍:Buzzy将AI视频创作从生成15秒短片的“玩具”升级为制作20分钟广告片、电影长片的专业导演工作室,通过无限画布和多智能体协作,解决视频创作中角色一致性、场景精准控制、叙事连续性难以兼顾的行业痛点。
Movies Marketing Artificial Intelligence
AI视频生成 无限画布 智能体协作 角色一致性 商业广告制作 电影级长片 多模型集成 视频编辑 创意导演工具 工作流模板
用户评论摘要:用户主要关注角色/品牌视觉在多场景、多模型切换时能否保持一致性;能否通过画布而非对话控制叙事节奏、情感与灯光;需要持久的配音和音频对齐功能;对企业用户而言,安全合规与团队协作是刚需。多数评论对“非抽奖式”的确定性控制感到兴奋。
AI 锐评

Buzzy的“无限画布+多智能体”架构精准切中了当前AI视频行业的两大痛点:一是“生成即失控”,传统工具随机性强,用户像赌徒一样拉下拉杆,等待模型赏赐好结果,而Buzzy通过画布给予用户逐帧导演的权力,把随机性压缩到可控区间;二是“碎片化”,多数工具只能产出15秒“GIF片段”,无法支撑真正的叙事长片,Buzzy通过可复用的角色、灯光、场景参考及跨模型一致性维持,提供了从创意到成片的完整管线。

但其面临的挑战同样明显。第一,70+模型的集成意味着巨大的工程维护成本和推理成本,且不同模型间的“漂移”虽能容忍,但全自动路由策略是否足够智能仍需经受长片制作检验;第二,当前评论中缺少对生成速度、渲染时间、用户实际创作20分钟成片的成本(金钱与时间)的具体反馈,产品是否真能达到“商业可用”级别存疑;第三,团队协作、安全合规等企业级功能的缺失,使其主要停留在个人创作者工具层面,难以渗透到品牌广告公司的生产流水线。

本质上看,Buzzy不是在“生成”视频,而是在“制作”视频。它更像一个AI混编的后期合成软件,而非简单的文生视频工具。对于那些厌倦了“AI抽卡”、希望精确控制每一帧画面的专业创作者而言,它是有价值的进化方向。但若想真正替代传统影视制作流程,Buzzy还需要拿出更多长周期、多角色、复杂叙事的标杆案例,并解决协作与成本问题。

查看原始信息
Buzzy
Buzzy is the first truly agentic infinite canvas for video creation. It helps you spark ideas, build moodboards, and precisely photoshop your video in one click — while creating unlimited-length blockbuster films and commercial ads, keeping unlimited subjects consistent, controlling lighting and camera angles, and precisely editing video like photoshop in one workspace.
Hey everyone! 👋 Ella here, maker and founder of Buzzy. We are already facing a critical transition point where AI videos are starting to move into real production pipelines instead of being just a funny toy. That's why we build Buzzy. What it is: Buzzy is the first truly agentic infinite canvas for that revolutionizes the video generation in filmmaking and TVC ads landscape. What makes it different: It's no longer a 15-second short GIF, but 20-minute blockbuster films or studio-quality ads!😲 Creators get full control of every scene using Buzzy's infinite canvas and more than 50 tools and more than 70 image and video models, and the whole workflow is enhanced and accelerated by the smartest agent on the canvas. The seamless integration of canvas and multi-model agent empowers creators, marketers, filmmakers, and studios to craft narratives with unprecedented control. Key features: - Create unlimited-length blockbuster films and studio-quality ads - keep unlimited subjects consistent 🤯 - Steal thousands of professional workflows from filmmakers and brand advertisers - switch view-angle, control lighting seamlessly - Precisely edit videos like Photoshop - Agentic Canvas with all image and video models available 🎁 SPECIAL OFFER 🎁 To celebrate Buzzy's launch, we're offering a discount for the Product Hunt community: - Use HUNTED2026 to enjoy 20% discount for your first monthly subscription - The yearly subscription is already at 60% discount The discount code is valid for the next 48 hours. I can’t wait to see what you create with it. Let’s bring your ideas to blockbuster films together. 👇 USEFUL LINKS 👇 - Create your first film on Buzzy: https://www.buzzy.now/ - Watch amazing films made by Buzzy: https://www.buzzy.now/agent/long... - Try create your own ads: https://www.buzzy.now/agent/6a38... - YouTube tutorial: https://youtu.be/6oK8vnlfhXU?si=... Cheers, The Buzzy Team
16
回复

@shiying_zhang1 That's awesome - being able to have entities/identities being persistent across shots is very cool, are you able to use those identities across different projects? (thinking you can then run a consistent 'character' for all brand media).

4
回复

@shiying_zhang1 how Buzzy handles creative consistency over longer projects. When users are building a multi-scene video or campaign asset, does the AI maintain context around brand guidelines, characters, visual style, and storytelling direction across iterations? Also, are there plans to introduce team collaboration features for creative teams and marketing departments?

0
回复

@shiying_zhang1 Curious, how do you handle governance for AI-generated applications? For example, can teams enforce security, compliance, and architecture standards across every app built on Buzzy?

0
回复

The canvas instead of a chat timeline clicks! Wondering re 70+ models...If I build a character in one scene, then switch the underlying video model for the next shot because it handles motion better, does the character hold or drift? Great product overall.

3
回复

@artstavenka1 Yes — you can switch the underlying model from shot to shot while keeping the same shared character and references. Some variation can still depend on the model, but Buzzy gives you a consistent project context and lets you refine any drift without rebuilding everything from scratch.

1
回复

Curious how much control the tool gives on tone. Story tools depend heavily on whether you can control pacing and emotion.

2
回复

@aix_pan Exactly. Buzzy lets you shape the story scene by scene, revise individual shots, and adjust pacing and mood without starting over.

0
回复

Stopped making anything narrative in AI because it was easier to just draw storyboards. This is the first thing that made me reconsider.

2
回复

Is this closer to a creative studio you sit inside, or a chatbot you send prompts to?

2
回复

@boyso Closer to a creative studio you work inside. You organize your story, characters, references, and scenes on the canvas, then use conversation to revise specific parts—not just send prompts and wait.

0
回复

The mood in the trailer sample holds all the way through. That's rarer than it sounds.

2
回复

@li_huang1 Really appreciate that

0
回复

Supporting 70+ image and video models sounds like a huge engineering challenge. How do you decide which model to route a specific task to? Is it automatic based on the edit, or does the user have fine-grained control?

1
回复

Love seeing AI video move away from short disconnected GIFs toward actual production-grade film pipelines @shiying_zhang1!

For creators using Buzzy to produce 20-minute films or commercial ads, how does the platform handle lip-syncing and dialogue audio track alignment across persistent characters? Do character entities support native voice embedding alongside visual references to keep vocal tone consistent across shots?

1
回复

@franz_briones Hi Franz, this is a very good question. For vocal consistency, we do have audio node on our canvas, where user can generate their specific character voice using the audio node. Then for each video clip, they can reference that audio node so seedance 2 will keep the reference to that voice and keep the voice consistent across all video clips.

We don't use lip sync a lot, since the results seems not quite good but using the seedance 2 directly by referencing that audio node will really help

0
回复

Love the idea!!! I've explored some tools like this in the past. How quickly does AI start understanding your style?

1
回复

@niallcleaver Great question! You can establish your style through references, characters, lighting, and visual direction, then reuse them across scenes. The more clearly you define those elements, the faster the workflow becomes consistent.

0
回复

Thinking about it for onboarding videos at work. Same character walking new hires through 5 departments.

1
回复

@mingyouagi That’s a great use case.

0
回复

Something about the pacing of the demo felt intentional in a way most AI reels don't.

1
回复

Noticed the samples include a few 8-10 shot scenes, not just 3-shot cuts. Curious how deep the consistency really goes.

1
回复

@andrewstone11 Great catch! Buzzy keeps characters, locations, and references persistent across scenes, so consistency can extend well beyond short 3-shot clips. We’re continuing to push it further for longer sequences.

1
回复

Every AI video tool right now feels like a slot machine. Pull, pray, repeat. Curious if this feels different.

1
回复

Reads like something between a video generator and a story tool. Which side does it lean more?

1
回复

@fayann It leans more toward a directing workspace than a pure video generator. You can generate footage, but the real focus is organizing and refining stories, characters, scenes, lighting, and shots with continuity.

0
回复

Whoever cut the demo reel understands rhythm. That's a signal about the underlying tool.

1
回复

The storyboard is doing more work here than the video model. Once you have characters, objects, and locations as first-class entities that persist across shots, consistency stops being a prompting trick and becomes a data problem. Curious whether edits propagate back down to already-rendered shots. Super cool, and excited to use it.

1
回复

@aidan_codefox Exactly — Buzzy treats characters, objects, and locations as persistent entities, while shot-level edits can stay local. We’re also exploring better controls for propagating changes across a sequence. Great question!

0
回复

Downloaded the demo clip and showed my partner. Their reaction: "wait, you made that?" I didn't, but I want to.

1
回复

The agent workflow demo where it suggests the next shot, didn't know I wanted that until I saw it.

1
回复

Sat through the trailer sample twice. Wasn't planning to. Highest compliment I can give a landing page.

1
回复

Not my usual corner of PH but stopped by to say nice launch. The samples do the talking.

1
回复

Making anything longer than 10 seconds coherent has been the wall for me. Curious if the co-director framing actually helps.

1
回复

Watching the samples I realized this is what I've been trying to make manually with 4 different tools stitched together. Finally, one place.

1
回复

The co-director framing is doing more work than people realize. A tool that treats you like the director, not the prompter.

1
回复

@kaining_luo You nailed it. The goal is to keep the creative decisions with you while Buzzy handles the execution—not turn filmmaking into prompt engineering.

0
回复

The "first shot great, everything after falls apart" curve is the pattern I keep hitting. Genuinely rooting for someone to break it.

1
回复

I've been circling the AI video space for about a year now, mostly lurking in discords and testing every new thing that drops. What kept bugging me is that every tool felt like it was built for the prompt, not for the story. You'd get a gorgeous 4-second clip and then have zero idea how to make it talk to the next clip. Reading through your page, the co-director framing clicked instantly — it's the difference between handing someone a paintbrush vs handing them a storyboard. Genuinely excited to try this over the weekend.

0
回复

For creators posting daily, could this replace the intro they film every morning? Same look, different subject.

0
回复

Pitch phase is where most of our budget disappears. Any tool that helps us show a real-looking concept before shooting is worth a serious look.

0
回复

Overall experience is great, but complex scenes can still have some small details that need fixing.

0
回复

The biggest benefit is the creative freedom. I can explore ideas without worrying about production costs. :))

0
回复

The samples read like the team watches films, not AI reels. That's rarer than it sounds.

0
回复
#4
Kastra
Runtime authorization for Claude, Cursor, Codex and OpenClaw
337
一句话介绍:Kastra 是一个为AI智能体(如Claude、Cursor、Codex等)设计的运行时授权层,在工具调用、命令执行、数据访问前以亚毫秒级延迟强制执行安全策略,解决企业使用自主AI时“无法在动作执行前阻止不当行为”的核心痛点。
Developer Tools Artificial Intelligence GitHub Security
AI安全 运行时授权 Agent治理 策略引擎 提示注入防护 敏感数据防护 预执行控制 跨框架支持 开发工具 企业级
用户评论摘要:用户普遍认为预执行控制理念正确,但质疑其抗对抗性攻击能力(如绕过策略的诱导)。关注策略引擎如何处理嵌套工具调用中下游连锁动作的授权。关心运行时不可达时的“容错模式/故障关闭”默认行为,认为单点故障风险需谨慎。开发者也表示需要动态上下文感知策略及调试优化工具。
AI 锐评

Kastra抓住了当前AI Agent在生产环境中“裸奔”的死穴。绝大多数团队还在依赖“提示工程”这种概率性手段来约束Agent,本质上是在跟一个不可控的黑盒“商量”。Kastra的逻辑很清晰:既然Agent本身不可信,那就不要在“让它变乖”上浪费精力,直接在动作执行路径上架一道铁闸,用确定性规则否决掉它任何越界行为。

这个方向是对的,但在光环之下有两点必须警惕。其一,所谓的“亚毫秒级延迟”在单层策略判断上轻而易举,但面对用户评论中提到的“嵌套工具链”和“复杂上下文”时,策略引擎的评估深度和延迟如何平衡,是工程上的巨大考验。目前仅回应“每个下游动作独立评估”,但并未解释复杂策略树下的实际性能衰减。其二,将授权从“监督”变为“执行”,本身就是一种傲慢——它预设了管理员能预判所有风险场景,并写出百无一漏的规则。当Agent的行为组合超出预期,或企业团队因规则复杂产生误配,导致大量误杀或漏判时,Kastra很可能从“安全救星”变成“开发阻碍”。

不过,对于一个刚上线的工具,Kastra至少解决了从“零控制”到“有控制”的质变。它最大的价值不是技术指标,而是提供了“信任但验证”的工程范式。当所有人都在吹嘘Agent的自主能力时,Kastra冷静地说:先把门锁好,再谈让管家自由活动。对于正在把AI嵌入核心业务流程的团队,这种“煞风景”的产品思路,恰恰是最务实的保险。但长久来看,它必须学会与概率性的风险评估共存,而非仅仅依赖冰冷的开关。

查看原始信息
Kastra
Kastra is the runtime authorization layer for AI agents. It decides what agents can and cannot do before actions execute, enforcing policies with sub-1 ms latency across tools, prompts, inputs, and outputs. Use one control plane to govern agents and policies across Claude Code, Cursor, Codex, OpenClaw, the Anthropic SDK, the OpenAI SDK, and more. Prevent unauthorized tool use, prompt injection, and exposure of sensitive data before they become incidents. Trust the rules, not the agents.

Cross-framework support is huge. Claude Code, Cursor, Codex, one policy layer, that's rare.

48
回复

@sulemna_ola Our neutral enforcement engine across models and agentic systems enables us to enforce the same rules throughout the stack. We want to enhance how you use the tools you love, not replace them.

2
回复

@sulemna_ola Yes, and we are adding support for more tools every day! Thank you!

1
回复

What I like here is the "trust the rules, not the agents" framing. It flips the usual assumption that agents will behave if you prompt them nicely enough. They won't always.

48
回复

@nancy_philip Exactly, we have seen our users want to customize the rules that adjust to their agentic coding style. We allow them to do so; at scale, this helps build trust, and trust allows them to delegate more work to their coding agents.

2
回复

@nancy_philip Exactly, only a deterministic, neutral layer can enforce this.

1
回复

Sensitive data exposure prevention baked into the control plane, not bolted on afterward. That's the correct order of operations.

41
回复

@peter_victor That's exactly the philosophy behind Kastra. We believe authorization should sit directly in the execution path, where every tool call, shell command, API request, and file operation can be evaluated before it executes rather than remediated afterward.

2
回复

I've watched teams try to solve this with prompt engineering alone, telling the agent what it's allowed to touch and hoping it listens. That's not authorization, that's a suggestion. What Kastra seems to be doing is treating agent permissions the way we already treat API permissions, with enforced rules instead of good intentions. Overdue shift.

40
回复

@ramish_saje Yes, the prompt engineering rules/policies are often probabilistic approaches that can't be guaranteed to work 100% of the time. There is a margin of error that some engineers and enterprises can't take the risk of "trusting". That is why we built our engine fully deterministic, so we can build these rules for every custom angle but make sure the highest-risk actions are avoided before they happen. What exists in the market is mostly post-action observability and monitoring, which is insufficient for autonomous AI.

2
回复

Genuinely curious how the policy engine handles nested tool calls, where one authorized action triggers a second unauthorized one downstream.

39
回复

@steven_granata Great question. We evaluate every execution independently, not just the initial request.

If an authorized action triggers additional tool calls, each downstream action is intercepted and evaluated against the same policy engine before it executes. Authorization doesn't "carry over" simply because the parent action was allowed.

3
回复

I've been burned by an agent calling a tool it shouldn't have. A pre-execution control plane like this feels overdue, honestly.

28
回复

@grayson_carter3 Yes, this is a very common problem. We've heard similar stories from a lot of developers over the past few months. As agents become more autonomous, trusting them isn't enough; you also need deterministic guardrails around what they're allowed to do. That's the gap we set out to solve with Kastra.

1
回复

Hi PH. I'm Carlos. Fernando and I built Kastra.

Here's what kept happening. We'd deploy AI agents in real environments: a coding agent, a support agent, and an infra agent. It had credentials and access, and there was nothing structural stopping it from running any action it decided to take, even in production.

Most teams have authentication, logging, monitoring, and evals. Almost none have a layer that decides "this specific action cannot execute" before it executes. The controls that exist look at what the agent already did. We wanted something that decides on the input and output to cover most of the execution level risks. This required a new infra layer for autonomous AI around the authorization angle.

That's Kastra. Every agent action is evaluated against policy and gets an allow or deny decision in under a millisecond. When something needs a human review, that approval now completes in about a second, across our desktop app, web console, and macOS notifications. We reduce that number down every week, and it changes how the product feels. Sub 1ms latency means human-in-the-loop authorization is practical, not theoretical. Kastra is neutral and deterministic across the stack and works with the tools you already use.

We integrate via 1 click to Claude Code, Codex, OpenClaw, and Cursor, and other regulated enterprise use cases right now. The runtime and policy pack library are open source. The enterprise control plane is commercial. The product is free to try, if you are wondering if you need Kastra run these 2 commands below on your coding agent to scan for potential risks your agent has already done that should have policies built. Everything stays locally on your machine, so you can do this safely.

Run "brew install kastra-labs/tap/kastra-edge" and then once connected to your account, you can run "kastra-edge scan".

Happy to go into the policy engine architecture, the latency story, how we think about fail-open vs. fail-closed, probabilistic vs deterministic rules, or anything else. We'll be in the comments all day.

Carlos & Fernando

14
回复

@carlosjimenez1 Really like that Kastra focuses on governing AI agents before they take action rather than reacting afterward. Having a single policy layer across different AI tools feels like a practical way to improve security and control. Good luck with the launch!

0
回复

@carlosjimenez1 how Kastra handles dynamic permissions , can policies adapt based on an agent’s context, user intent, or risk level in real time? Also, are there plans to provide analytics around blocked actions and security insights to help teams improve their agent workflows over time?

0
回复

@carlosjimenez1 How does Kastra balance strict policy enforcement with developer experience? For example, how do teams debug or fine-tune denied actions without slowing down agent development?

0
回复

I like that this doesn't ask me to trust the agent's judgment at all, it just checks the action against a rule before anything happens. That's the right mental model. Most of the "AI safety" tooling I've evaluated this year still leans on the model behaving well, which never sat right with me.

14
回复

@thomas_jack3 Very accurate! We learned this by testing the product with some of the top enterprises pushing AI products into production. We realized autonomous AI products required a new infra layer and we built Kastra to solve the trust problem in a scalable way.

0
回复

I've watched a few "AI firewall" products overpromise and underdeliver, so I'm cautiously optimistic rather than sold. What would actually convince me is seeing how this performs against a red-teamed agent trying to route around the rules, not just against a well-behaved one following the happy path.

11
回复
@sheikh_umair1 That’s a fair point. This has been one of our biggest design priorities. We’ve already tested Kastra against adversarial scenarios because the organizations evaluating and deploying it are building sensitive AI workloads. The product has to be robust enough to earn that trust, and we continue to harden it through real world deployments and testing. Feel free to try it out and give us your feedback!
1
回复

What I like here is the scope, one control plane across Claude Code, Cursor, Codex and OpenClaw. I've juggled separate configs for each tool before, and it's a mess nobody talks about enough.

5
回复
@edward_moore5 Thanks! That’s exactly one of the problems we kept hearing from developers. Teams don’t want to manage a different set of policies for every AI tool. The goal with Kastra is to define your rules once and have them enforced consistently across all the agents your team uses.
0
回复

you mentioned fail-open vs fail-closed as something you're happy to dig into, so - what's the actual default when the Kastra runtime itself is unreachable or crashes mid-decision? fail-closed sounds obviously "safer" on paper but in practice that just turns your authorization layer into a single point of failure that can halt production agents. is that a global setting, or does it vary per policy depending on how risky the action is?

4
回复

@galdayan This is a great point. We learned from enterprise conversations that policies needed to be tested per environment for high-frequency authorizations in order to be trusted by large companies that could not have a single point of failure.

So we enabled these fail-open and fail-closed features so every integration can be tested and evaluated before launching to production to make sure the risk of the integration was minimal. Think of it as sandboxing policies in a custom environment: one mode logs every decision without blocking, and the other one logs and blocks them. This can be switched on the dashboard at any time once the customer is ready. The interception and enforcement engine works just the same with every authorization.

1
回复
@galdayan Thank you for the support. Exactly that is the goal, reducing friction and risk across the stack.
1
回复
Really interesting approach to AI security. Instead of monitoring after something happens, preventing risky actions before execution feels like the right direction.
3
回复

@1mirul exactly! The primary goal is to deterministically stop things before they happen. The monitoring and logs are a natural byproduct, since we're already intercepting every request anyway.

0
回复

@1mirul Thanks! That's exactly how we think about it. Monitoring is valuable, but once an action has already happened, you're often in incident response mode. We wanted to give developers a deterministic decision point before execution, so they can safely delegate more work to AI without giving up control.

0
回复

How does kastra prevent prompt injection attacks where an agent attempts to bypass tool parameters by obfuscating CLI commands? huge congrats for shipping 🙌@carlosjimenez1

3
回复
@vikramp7470 Thanks! Great question. We don’t rely on the agent resisting prompt injection. Even if an injected prompt changes the agent’s behavior, every tool call and its arguments are still evaluated against deterministic policies before execution. If the resulting action violates policy, it’s denied or held for approval regardless of what prompted it.
1
回复

Congrats on the launch!

2
回复
@peter_tribelhorn Thank you Peter! Appreciate the support 🙌
0
回复

congrats on hitting product hunt @carlosjimenez1 keeping dev velocity high while putting hard guardrails around shell commands is pure leverage.

2
回复
@priya_kushwaha1 Thank you so much Priya! Appreciate the support
0
回复

Huge congrats on the launch! :)

2
回复

@lisa_shmulyan thank you for the support Lisa! 🙏

1
回复

@lisa_shmulyan Thank you for the support, Lisa!

0
回复

Congrats on the launch. Runtime authorization for agents is a strong direction because the risk usually appears after the model decides to act. How do teams define the boundary between low-risk actions that can run automatically and sensitive actions that need explicit approval or audit?

2
回复
@yaroslav_stelmakh Thanks a lot! That’s exactly where policies come in. Teams define which actions should be automatically Allowed, Denied, or Held for human approval based on factors like the tool, environment, data sensitivity, repository, or any custom condition. The goal is to automate the low risk actions while putting guardrails around the actions that actually matter.
0
回复

Finally something that treats agent permissions like a real problem. The sub-1ms latency claim seems legit - I didn't notice any lag in tool calls when I tested it with Claude Code.

2
回复
@baharyakupbcr2 I am glad you were able to test it! Let us know what your Recon scan found out in the “Recommendations” section!
1
回复

@baharyakupbcr2 Awesome Bahar, glad you had a good experience! Keeping the AI flow fast was a concern for us since day one. Every decision was made around adding this security layer without hurting performance.

0
回复
Congratulations on your launch! This feels like the missing piece for regulated AI adoption. Enterprise‑ready - control, neutral, deterministic, and fast.
2
回复
@odeth_negapatan1 Exactly, this is our goal. We want to become the default neutral authorization layer a top enterprise deploying large scale product and also an indie developer can use. Appreciate the support!
0
回复

Thank you@odeth_negapatan1, that's exactly the gap we set out to close!

0
回复

The deterministic claim holds right up until the tool is generic.

Denying delete_file or write_prod is the easy case, because the intent sits in the tool name and you can rule on it in under a millisecond. But a coding agent needs a shell. Once bash or run_command is permitted, the dangerous action stops being a tool call and becomes an argument string, and the engine is no longer authorizing an action, it is parsing arbitrary shell to predict one. That is the part that cannot be made deterministic.

It is also where the real incidents will come from, because nobody takes the terminal away from the agent. They deny it the scary sounding tools it was never going to reach for anyway.

How does a policy express that boundary? Can you constrain inside a permitted generic executor, or does allowing bash effectively allow everything downstream of it?

2
回复
@abdullah_javaid3 Great question. We don’t treat bash as a single trusted action. We inspect the command and its arguments before execution, so policies can constrain what is allowed inside a generic executor, not just whether the tool itself is allowed. The goal is to authorize the actual operation being performed, not simply the wrapper used to execute it. Thanks for the thoughtful question.
1
回复

Carlos, the deterministic engine is the part that matters for my world. I sell AI into healthcare, and "the agent usually behaves" is not a line you can put in front of an auditor or a BAA, since a probabilistic guardrail is a suggestion while a deny decision in the execution path is a real control.

Where I would push: preventing exposure of sensitive data assumes you can recognize it first. Is a policy written against the tool or endpoint, or can it key off the content itself, say a rule that blocks any output carrying patient identifiers no matter which tool produced it?

And does every allow or deny land in an audit trail I can hand to a compliance review? In regulated work, proof that the control fired is worth as much as the control.

2
回复
@clemente_lopez1 Really appreciate the thoughtful questions. Yes, policies can evaluate both the action and the content, so you’re not limited to matching on a specific tool or endpoint. You can define rules like blocking outputs containing patient identifiers regardless of which tool generated them. And yes, every Allow, Deny, and Hold decision is recorded with the policy that matched and the supporting context, so you have an audit trail you can use for compliance and investigations. We completely agree with your last point. In regulated environments, proving that the control executed is just as important as the control itself. That’s a core design principle for us. We are working with healthcare enterprises with the same use cases.
1
回复

Runtime authorization is the missing layer for coding agents. Prompt rules and post-action logs are useful, but they do not help much once an agent has already touched prod or leaked data.

The part I’d want to understand is policy authoring. For a team using Claude Code and Cursor heavily, do policies start from templates, scanned risky actions, or does each org end up writing a lot of custom rules by hand?

2
回复
@aditya_harish_2002 Great question! We support all three. Teams can start with prebuilt policy templates, scan their existing agent activity to automatically generate recommendations, or write custom rules in plain English. Most customers end up combining all three, then refining their policies as they learn how their agents work.
1
回复
@alieksia The plain english rule building feature is enabled in our Pro and Teams plan, but you can still use the Recon scan and our policy packs templates in the free plan that covers some of the highest risks. The policies you build on your own, you can export them and they are opensource.
1
回复
@aditya_harish_2002 Yes, the policies you build with the plain english policy building feature are your own. You can export them and they are opensource when using the Free, Pro and Teams plan. The policy packs and custom rules we build for enterprises we limit exporting as they hold our IP on these rules for complex regulated environments.
0
回复

Governing agents before they act instead of auditing after the fact is the shift I've been hoping for.

2
回复
@charlos_brat Appreciate that! That’s exactly the shift we believe needs to happen. Observability tells you what already happened. Authorization decides what is allowed to happen before the action executes.
0
回复

@charlos_brat Exactly! And the nicest part is it isn't either/or. Every decision Kastra makes before an action runs is written to a tamper-evident, per-tenant hash-chained decision log, so the "audit after the fact" is just the other face of governing before it. And because it's chained, you can prove nothing was altered or dropped. There's a chain-verification API an auditor runs, not a log file you take on faith.

0
回复

The line I keep going back to is "approval now completes in about a second".

We run an approve-before-send step on the support side, and latency was never what broke it. Attention was. While the queue is short people actually read what they are approving. Once it gets long they stop reading and start clicking, and you have a human in the loop on paper but not in practice. Making approval fast is good, but it also makes it cheap, and cheap approvals get rubber-stamped.

I have never solved this properly, so genuinely asking: does the console show time-to-decision, not just allow/deny counts? A median of 0.4s across 200 approvals a day, and the same allow rate across 20, are two very different situations and only one of them is human-in-the-loop.

One smaller thing, and it may be out of scope for an authorization layer: the approver sees the action and the policy that matched. What usually decides whether an action should run is what the agent read three steps earlier. Do you carry any of that into the approval prompt?

2
回复
@jernej_jan_kocica Great point, and it’s something we’ve thought about a lot. We don’t think every action should require approval. If everything goes through a human, people eventually stop reading and start clicking, just like you described. Our approach is that approvals should be the exception, not the default. Policies can Allow, Deny, or Hold. Hold pauses execution and asks for human approval only for the scenarios you define, while everything else is handled automatically. Also yes, we already include the execution context that led to the Hold decision, not just the action itself, so the approver has the information needed to make an informed decision. Regarding the time to decision this is one of our strongest angles that support billions of interceptions per environment with sub 1ms latency. We did a lot of infra work to get to this scale.
1
回复

I have been testing kastra with my team and even though I’m not very technical it’s very simple to use and it has helped us a lot while building our product.

2
回复

@paula_jimenez4 Kastra is great for new builders learning how to code; it teaches them the risks and helps them understand the agentic coding workloads better by being able to evaluate all the actions, risks, and decisions at runtime.

0
回复

I use Kastra everyday to discover the risks while coding and it’s very useful.

Highly recommend 👍🏻

2
回复

@carlos_campos93 This is great, thank you. We're happy we can bring value to you!

0
回复

Love the recon scan to know all the unexpected/dangerous things Claude Code has done in the past. Some things I already knew but it caught some others I didn't know about it. And with automatic policy generation I don't need to worry about it messing up again!

2
回复

@churds Exactly, Richard, this gives you a baseline of policies that adjust to your coding behavior first. Then, once you have this baseline, you can build new daily policies on top of them to continue reducing risks you learn along the way. Some Recon scans from our customers show many important high-risk actions they were not aware of; some thought they were protected with their existing frameworks, and the reality was different.

0
回复

@churds Happy Recon could help, Richard!

0
回复

The integrations with tools like Cursor and Claude Code caught my attention. Looking forward to testing this in a real workflow.

1
回复

@monir_ We hope you enjoy it! Thank you for the support! 🙏

0
回复

the part that keeps me from letting my coding agents run fully unattended is exactly this. they will happily do something you didn't want and you only find out after the fact. a layer that decides no before the action runs is the piece i keep wishing i had. how granular can the policies get per tool?

1
回复

@terminal_candy very granular. Kastra sits in front of the action (a pre-tool hook for Claude Code / Codex / Cursor, or the LLM proxy), so every tool call is evaluated before it executes. A rule can target:

  • A specific tool: exact name, a list of tools, or a regex (e.g. only Bash, or every mcp__* tool).

  • The tool's arguments: substring, multi-substring, or regex conditions over the tool input (e.g. block Bash when the command contains rm -rf or git push --force; block Write when the file path matches .env). On the proxy path, arguments are flattened per field (output.tool_call.arguments.<key>), so you can match one named argument in isolation.

  • The context around the call: working directory, git repo/branch/dirty state, session, OS, agent permission mode, environment (dev/prod), model, principal/agent identity, data classification.

  • Content: built-in scanners (secrets, PII) can run over the tool input or model output with confidence/count thresholds.

And "no" isn't the only answer. Each rule chooses its effect: hard deny, hold (pause the agent and page a human to approve, with a timeout that fails open or closed, your choice), monitor (allow but flag), redact, rate-limit, or spend-cap.

0
回复

Congrats on the launch. The "trust the rules, not the agents" framing is the right instinct — we've landed on the same philosophy internally. One thing I'm curious about: when a policy check fails, is deny the only outcome, or can a policy route the action to a human approval queue instead? In practice I've found agents need three outcomes, not two — allow, deny, and "a human decides this one." Sub-1ms is impressive for the allow path; wondering what the escalation path looks like.

1
回复
@nitish_garg4 Hey Nitish, for these scenarios users can create hold logic rules that can control these actions and wait for a human approval within a limited timeframe of response. This works for most scenarios, with webhooks we have enabled in the app we can extend this timeframe.
0
回复
#5
box
Simple computers for agent w/ full VMs
266
一句话介绍:Box通过命令行2秒内提供带桌面和SSH权限的完整Ubuntu虚拟机,以低至0.036美元/小时的超低成本,解决AI代理在开发、测试和运行时因计算资源供给慢、成本高而受限的痛点。
Developer Tools Artificial Intelligence
AI代理云电脑 虚拟机即服务 Agent沙箱 GPU云实例 低成本云主机 自动化开发环境 DevOps工具 云上IDE 弹性计算 开发者工具
用户评论摘要:用户普遍赞赏其低成本、全虚拟机架构和快照/分支功能,认为优于容器方案。主要关切点包括:物理隔离边界(同一宿主机不同盒子的内核/管理程序隔离性)、出口流量管控(防止Agent泄露凭据),以及安全审计流程是否能跳过内部采购环节。
AI 锐评

Box抓住了AI代理市场一个被忽视但极其致命的瓶颈——计算资源的供给效率和成本。当行业疯狂追逐模型能力迭代时,忽视了“等两个月才能拿到集群”的荒诞现实。从产品层面看,Box做了正确的取舍:坚持全虚拟机而非容器,虽然看似“笨重”,但彻底解决了容器沙箱在环境依赖、系统级权限和持久状态上的致命短板,这对于需要安装复杂依赖、运行嵌套进程或跨任务保持工作流的“真Agent”至关重要。其成本优势并非源于恶性补贴,而是通过自研快照系统、极简架构和欧洲数据中心本身较低的基础设施成本实现的,商业模式可持续。

然而,安全与治理的追问直指核心:快照回滚只能恢复机器状态,无法撤销Agent已执行的外部操作(如推送代码、发送邮件)。这意味着,对于需要接入外部系统的生产级Agent,Box目前实际上将信任和出口风控的责任完全甩给了用户。若不能提供基于网络策略的精细化出口管控和更严格的内存/内核隔离,它更适合开发测试、沙箱评估等场景,而非直接面向高安全敏感的生产环境。此外,“通过命令行在2秒内获得管理员权限的VM”本身就是双刃剑,虽降低了使用门槛,也同时大幅降低了恶意利用的阻力。Box要想从“好用的工具”跃升为“可靠的基础设施”,必须在安全边界上给出更明确的解决方案,否则其极致的效率优势反而可能成为风险放大器。

查看原始信息
box
Box offers the simplest, cheapest cloud computers for agents, thought for builders of agentic platforms & software factories. Run 'box new' in your terminal, in 2s get a beefy ubuntu VM, with admin rights, a desktop, ssh access. At $0.036/hr it is 10x less expensive than the likes of E2B, Daytona, Modal, so you can run more agents, or run them longer. Run up to 1000 concurrent boxes fully self-served, or ask us for more, with same-day support from the founders.

Hello Product Hunt,

2 years ago, I worked at the European Space Agency, scaling their black holes & binary stars research efforts.

At that time, GPT 4o already helped us find many optimizations that made the code more scalable. But to test them, our agents needed a cluster of cloud compute, which we had to wait 2 months to get. AI may work fast but compute provisioning is slowing it down.

After leaving, I spent 10 months building cloud coding agents, which needed cloud computers too, so I tried many self-serve options for "computers for agents" and "AI sandboxes".

But there was no fit, they were optimized for bursty, short-lived scenarios on containers rather than full computers where anything just works...

... all while being super expensive, high multiples of what the same compute is worth in the average datacenter in Europe.

So 2 months ago, we built and launched box on X, our on-demand computers for agents, super cost-efficient and powerful.

Now our computers for agents have already run for 10 years in just 2 months, growing steadily every week.

Today, they already power many agentic platforms, software factories, cybersecurity audits platforms, with real customers in production.

Try box and let us know how it goes :) https://box.ascii.dev

12
回复

@anic_dev congrats on the launch!

5
回复

@anic_dev  Congrats on the launch. You guys built a great product. Happy user here!!

1
回复

@anic_dev Curious about how Box handles agent memory and environment persistence ,can agents maintain context, learn workflows over time, and reuse previous setups across different tasks? Also, are there plans to add collaboration features where multiple agents can work together on complex workflows?

0
回复
The months to provision a cluster detail is the real story here that's the actual bottleneck most "AI moves fast" narratives skip over. Compute governance and procurement cycles don't care how fast your model iterates. Curious how you're hitting $0.036/hr while staying full VM rather than container is that just leaner infra overhead, or are you oversubscribing hardware more aggressively than E2B/Modal do? Also: what happens if an agent inside the VM does something destructive is there a snapshot/rollback layer, or is isolation the whole safety model?
4
回复

@thys_beesman we hit such low price because we designed the whole system for cost efficiency, using no dependencies and minimizing the operational costs. we have our own snapshotting system (beautiful tech built in-house), allowing you to easily stop/resume/fork your machine.

1
回复

@thys_beesman Compute in European datacenters is just that cheap, its just hard to build around, we still make v good margins.

0
回复

Hey guys!

im Kirill, the CTO of box.

Today, we're seeing 7 consecutive weeks of big usage growth, and compared to six months ago, our MRR is 130× higher. It's been an incredible ride, but we're still early, and we're shipping every single day.

I'd genuinely love your feedback: good or bad. If you're building coding agents, personal assistants, or any agentic product, you can find box super useful!

3
回复

@luaroncrew The goat

0
回复

Interesting. I just implemented this use case and landed on E2B. I'm running a platform that does agent evaluations of dev tools and the agents need sandbox environments to build in to validate the dev tool. How do you differ from E2B? I just built this feature a couple days ago so its quite new!

2
回复

@tessak22 box.ascii.dev/compare explains well how we're positioned
tldr: if all you need is 4vCPU 8GB ram sandboxes we're 9-10x cheaper with comparable feature set

2
回复

Full VM + snapshot/fork is the right shape for agents that need real state, not just a clean container every run.

The self-serve scale is the part I’d want to stress test. If someone spins up hundreds of boxes for long-running coding agents, do fork/resume times stay predictable, or does it vary a lot once the fleet is hot?

2
回复

@aditya_harish_2002 no it's stable, even at the thousands box per user scale

1
回复

Good luck with the launch guys, amazing product!

2
回复

@meta_nkt thank you Nikita!!

0
回复

Using box to offload large tasks from strawberry.computer phone prototype - works pretty well for me!

1
回复
0
回复

Cool stuff. Where's best to DM you?

1
回复

@ogandreakiro X anic_dev

1
回复

The safety half of Brandon's question never got answered, and it is the more interesting half.

Snapshot and fork is recovery, not prevention. Restoring the machine undoes what the agent did to the machine. It does not undo what already left it. A pushed commit, a sent email, a dropped table on a managed database, a paid API call, none of those roll back when you restore the box.

Isolation bounds the blast radius to the VM. It does not bound it to the world, because the thing that makes an agent useful is the credentials you put inside the box with it.

So which is it today: is egress the customer's problem to govern above box, or is there something at the boundary? That answer changes who box is safe to hand to.

1
回复

Full VMs with admin + ssh instead of ephemeral containers is the right tradeoff for agents that install system deps or spawn nested processes — that is exactly where container sandboxes fall over. Two things I would verify before running many concurrent: what is the isolation boundary between boxes on the same host (separate kernels/hypervisor vs shared), and can I lock down egress per box or do agents get open outbound by default? Also, does box state persist across runs or is every "box new" a clean image?

1
回复

Brandon's "procurement cycles don't care how fast your model iterates" line hit hard — that's WinBidIQ's whole world, just for federal contracts instead of compute. Skipping procurement is the pitch here, but does any customer ever need a security review anyway before spinning up VMs with admin rights?

1
回复

ran box new and had an ubuntu vm up in a couple seconds, which genuinely surprised me. the price point makes it easy to justify spinning up extra agents without watching the meter.

1
回复

Congrats on the launch @anic_dev

I’ll give box a spin! btw the branding is goated 📦

1
回复

Full VM per agent is smart approach. Isolation at that level removes a lot of headaches with shared state. Whats the cold start time like when spinning up a new box?

1
回复

"box new → beefy Ubuntu VM in 2s with admin rights" is a great DX. The thing I'd want to see front and center for agent VMs is teardown and isolation — an agent with admin rights and ssh is one bad loop away from doing real damage or leaking creds across boxes. How ephemeral is each one, and what stops an agent in box A from reaching box B? That's the part that decides whether I'd let something autonomous loose in it. Following.

1
回复

I’ve been using box for a few weeks now. By far the best provider and it’s not even close. Great prices, super fast and the founder proactively reached out when they saw I had an issue. My preferred sandbox provider! 📦📦

1
回复

Congrats on the launch!

As agents move into enterprise workflows, how are you thinking about secure access to private networks, internal systems, identity, and policy, not just the execution environment itself?

1
回复

Cool to see pricing finally getting competitive for agent VMs. One thing that would really help builders like me: a built-in snapshot or image registry so I can prebake a box with my project's tooling and spin up ready-to-go environments instead of installing dependencies every time. Would save a ton of time and make the 2s boot claim even more useful.

1
回复

@mnevvertumazji you already can! see the templates section of docs.ascii.dev

0
回复

Congrats on the launch! Been using box and ascii for a few months now and couldn't be happier! Also support is really good, the founders are great. Love the CLI btw

1
回复

This looks really nice for spinning up cheap VMs quickly. One thing that would be a huge help: a built-in way to snapshot a box's state and resume it later, since agent workflows often need to pause and pick up where they left off rather than start from a fresh ubuntu every time.

1
回复

@hediyedilli it is totally already built in!

0
回复

too simple to use and spin up - been working on integrating this in my platform.

1
回复

@arth_tyagi Hope it's going well! Lmk if I can help in any way :)

0
回复

I'm a box user, been using it for (I think) a few weeks now. It's insanely fast, cheap (actually ridiculous), and good. Try them out! And congrats on the launch guys

1
回复

@kartikskabadi Thank you so much Kartik! It's a pleasure to help you

0
回复

Very cool product. Having deployed at scale, most agent sandbox pain is re-provisioning, not compute. Curious how fast fork actually is under load. The persistent-VM + fork-from-snapshot combo is the part that actually changes how you build agent loops.

1
回复

@aidan_codefox So many hard problems to solve in that space for so many different usecase.

We catter best to long running agent whose data is precious and never to be lost

1
回复

the isolation questions above are the ones I'd want answered before trusting this with anything real, but the one nobody's asked yet is billing-side: at $0.036/hr an agent that hangs or loops instead of exiting cleanly is cheap per hour but not necessarily cheap per incident if nobody's watching it. is there an auto-kill on idle/runaway processes, a hard spend cap per box, or is the ephemeral-and-cheap pricing itself the safety net and you're expected to monitor it yourself?

0
回复

$0.036/hr for a full VM with admin rights is a wild undercut of E2B/Daytona - what's actually different on the infra side that gets you there, bare metal instead of nested virtualization, oversubscribed hosts, or just thinner margins for now? asking because a price that much lower than everyone else usually means either a real efficiency win or something that gets quietly walked back once you're past the free/intro tier.

0
回复

box new giving you a real ubuntu box in 2s is the part i like. the thing i actually want this for is handing an agent a throwaway machine so it can go wild without touching my real setup. does the vm tear itself down after, or do i manage cleanup myself?

0
回复

The 2s-to-VM claim caught my eye — that's the number that matters for agent workloads. Question: can you snapshot a warm VM and fork it? The pattern I keep hitting is N agents that all need the same base environment (deps installed, repo cloned) — cold-booting each one wastes most of the cost advantage. If 'box new --from-snapshot' exists, that's the killer feature. Congrats on the launch.

0
回复

Full VMs instead of a locked-down sandbox feels like the right call for agents that actually need to do things. A lot of agent failures come from environments too restricted to be useful. The tension you're managing is the interesting part: a real computer per agent is powerful, but it's also the thing that keeps security folks up at night. How do you keep a confused or rogue agent contained without clipping the wings that make box worth using?

0
回复

The ESA-to-cheap-VMs origin is a fun one. The number I keep circling is the $0.036/hr — is that flat always-on, or does a box idle between agent turns bill differently? I run a bursty API on Fly and the thing that makes the economics work there is machines auto-stopping to near-zero between requests. An agent box is the opposite shape though: long-lived, holding state, mostly waiting on a model call. At 1000 concurrent that idle-but-alive time is most of the bill. Do you lean on suspend/resume for that, or is the bet that flat-and-cheap beats clever-and-metered for this workload?

0
回复

honestly the pricing is what caught my eye, super competitive. one thing that would help me though is a built in way to snapshot a box and fork it, like a template system, so i can spin up a preconfigured env instead of setting up every vm from scratch each time. would save a ton of time

0
回复
#6
ACME.BOT
No-slop AI SEO agent that interviews you first
202
一句话介绍:ACME.BOT是一款通过“先采访用户”来提取其专业见解、再生成符合Google EEAT标准SEO博客文章的AI代理工具,帮助中小企业在移动端低门槛地创建有深度、非千篇一律的高质量内容。
Marketing SEO Artificial Intelligence
AI SEO代理 内容创作 专家见解提取 人机协作 博客自动化 知识库 EEAT优化 移动优先 深度研究 反AI内容
用户评论摘要:用户高度认可“采访优先”避免了内容同质化,但也提出若干关键问题:采访能否帮用户自己形成观点?深度研究的来源是否可审查?审批关卡能否自定义阶段?产品反馈强调,通用AI草稿常因缺乏一手数据而致使创始人放弃发布,希望工具能真正保留“只有我才有的数据”和真实观点。
AI 锐评

ACME.BOT的核心价值在于其“反Slop”的人机协作机制,而非又一个AI代笔工具。一句“Competent and generic at once”精准戳中了当前AI内容创作的死穴——工具化输出只能生产信息密度极低的“行货”。而ACME.BOT通过“先采访你”这个反直觉的流程设计,将创始人真实的一手经验、反直觉观点、内部数据作为内容锚点,试图解决三个根本问题:1)内容“祛魅”问题——用户不信任缺乏个人经验的AI内容;2)工作流摩擦问题——用手机端短问答替代桌面端长篇编辑,降低创始人的参与门槛;3)内容复用与积累问题——通过持续演化的知识库让每一次采访都比上一次更深,形成竞争壁垒。但是,其产品设计也存在潜在脆弱性:在严格监管的B2B医疗等垂直行业,采访提取的“深度”能否真正抵御合规风险,是否能持续输出有价值的深度内容而不沦为下一个“采访界面包装的AI生成器”,还有待用户长期检验。另外,能否处理那些真没有明确观点或专家经验的创始人,在这一场景下产品价值可能会迅速退化。总的来说,ACME.BOT揭示了一个更本质的产品哲学:在AI可以轻易生产“大量内容”的时代,真正稀缺的是“非你不可的内容采集与转化能力”。它值得关注,但不可忽视其依赖使用者自身专业积累的前提条件。

查看原始信息
ACME.BOT
ACME.BOT turns your expertise into no slop SEO articles. Answer a few questions on a topic (from your phone) and your first-hand experience shapes your blog post. AI handles the boring stuff. Like keyword research, illustrations, references, internal linking, and publishing cadence. Create helpful non-commodity content that also satisfies Google's EEAT requirements.

Thank you @rohanrecommends for hunting us here today!

I left my job as a search infrastructure engineer at Google to start building AI products as an indie maker almost a decade ago, with my first launch being on this very platform.

A big pain point I faced while running my SaaS was the ability to reliably run a business blog. My options were

  1. Hire content automation team / Hire freelancers  - extremely expensive money-wise and time-wise. Not for most small-medium sized companies.

  2. Duct tape an AI automation - time consuming and fragile / hard to get right and be consistent

  3. “Slop Cannon” blog automation tools - dishonest to your audience / prone to search engine penalties

There had to be a way to reliably create content that takes the best of me (my opinions / counter-intutive thesis / expertise) and turn into posts that is genuinely useful to my audience - but it seemed like it just did not exist.

So, we built ACME.BOT - first for us. To be able to run our own blog with our own AI SEO agent.

Some key features:

  • 💬 Chat driven — accessible via web chat / Telegram even on your phone!

  • 🎙️ Interview mode — conducts an interview on the topic to understand your perspective and build the basis for the post

  • 🔍 Deep Research — finds supporting facts and research backing your point of view

  • 🗓️ Keyword research & content calendar

  • 📊 Diagrams — visually represent concepts

  • 🎨 Brand Tone — learns your voice from reference material

  • ✍️ Humanizer — reduces the chance of being flagged as AI writing

  • 🚧 Approval Gates — makes sure human checks are in place when needed

  • 🚀 Auto-publish to Shopify and WordPress

  • 🧠 Knowledge Base — everything you do on the platform feeds an ever-improving repository of your business context

At the heart of ACME.BOT is Human-In-The-Loop mechanic that drives the whole blog writing process collaboratively with low friction. So you step in only where you are genuniely irreplacable. Let ACME.BOT handle the rest.

Our initial set of users loved our features (like interviews, diagrams, deep research, run from phone).

Excited to be launching it here today!

As a part of our PH launch week we're offering free trials (no CC required) to everyone from PH. Sign up for the free trial here.

6
回复

@abhishek_iyer The interview-first step is the interesting bet. Most AI SEO tools I've tried fail because they never learn what the business actually sells, so they optimize for traffic nobody converts. What does the interview capture that a good brief wouldn't, and how do you keep it from drifting back to generic keyword output after a few runs?

2
回复

@rohanrecommends  @abhishek_iyer Congrats on the launch! Really like that this is built around actual first-hand expertise instead of just cranking out generic AI content, that EEAT problem is real and most tools just ignore it.

Couple questions:

With Interview mode, do I need to already have a clear opinion going in, or does it help me figure out what I actually think as we go? For Deep Research, do I get to see/check the sources before they end up in the draft, or does it just fold them in automatically? Can I choose where the Approval Gates kick in (like after the outline vs. right before publish), or is that fixed? Good luck with the launch!

0
回复

Other co-founder here 👋

@abhishek_iyer covered what it does, so I'll just say why I wanted it to exist.

What worked for my previous product was blog posts targeting the specific things people search right before they buy. It worked better than ads and outbound for us.

We tried to scale that with writers, then with AI. Failed both times.

Took me ages to see why.

The posts that worked had things nobody else could write. Numbers from our own work, opinions from years of doing it. That never made it into a brief. So every draft needed hours of editing before I'd put my name on it, and that time never existed. 

I generated a lot. I published almost none. 

The interview flow is our answer to that. It does the research, reads your docs, then asks you the few things it couldn't know and your answers anchor the piece.

Basically the tool I wanted when I was doing this by hand.

The feedback I most want: if you've tried AI writing tools and quietly stopped publishing the drafts, tell me why. That's the failure mode we are builting against.

4
回复

@abhishek_iyer  @dyuti We did not abandon ours, we changed what the model was allowed to do. Ours used to handle the mechanical parts too - quotes and dashes, stage tracking, SEO fields - and the quality was uneven. Once that moved into scripts and it was left with only the writing, the drafts got better and cost less. The generic tone was not the model. It was how much non-writing work we had given it.

1
回复

finaly we can ship blogs with our opinions!!

2
回复

@jiteshghanchi In a sea of AI Slop blogs - turns out human opinion is the most important missing piece!

1
回复

Love that the questions come to your phone instead of forcing you into a desktop session, that small choice removes the friction of actually sitting down to write.

2
回复

@yeterezlb Yes - part of building this tool for ourselves (being the lazy people we are) meant that we wanted to remove all friction to posting - that includes opening a website / desktop session. The feeling of writing content pieces on the go is truly magical :)

1
回复

what if I don't have a strong opinion yet on the topic, does it help me form one.

2
回复

@chloe_bennett Iteresting use case - like a good interviewer, it is designed to ask probing questions on a topic and will genuinely force you to think on a topic.

By default it asks only three questions which might not be enough to form a strong opinion on the topic. You can, however ask it to do a longer interview and it will oblige.

1
回复

Love the angle on using real expertise over generic AI slop, that's genuinely refreshing. One thing that would seal it for me though, can you add a way to pull in my own raw voice notes or transcripts and have the bot match my writing style? Right now I'm worried the output might still sound like AI even with my answers.

1
回复

@sedat127428 Couple of thingsr

  1. Your raw notes can verbatim uploaded in the Knowledge Base it will be used / quoted as a part of research and forming perspectives.

  2. For the raw voice of the narration, at the moment we're only have the option to learn voice from a reference URL that is already published. Based on your feedback will consider adding adding an upload option for the brand voice reference as well.

1
回复

Honestly the questionnaire flow is way faster than I expected, answered like 6 prompts on my phone during a coffee break and it spit out a solid first draft. The internal linking is the part that surprised me, it actually found relevant stuff on my old posts.

1
回复

@alparslansril Awesome - glad you found it useful :)

1
回复

The interview-first approach is actually clever. Most SEO tools just take a brief and run with it, ends up generic. Does the interview data get reused for future content or each piece starts fresh?

1
回复

@abdurrahman_fakhrul Yes. All interviews go verbatim in a knowledge base. So not only does the content get reused.

Even the interview builds on top of the previous interviews so that you're not asked similar set of questions again and again. Instead you get to go deeper into the topic with every interview.

2
回复

ACME.BOT

The “interviews you first” approach is interesting. SEO agents often jump straight into execution, but the quality of the initial diagnosis probably determines whether the recommendations are actually useful. Curious to see how you handle cases where the founder’s stated SEO problem isn’t the real bottleneck.

1
回复

@aryan787544 Great question, and worth clarifying a little about the purpose of the interview. It's less about being an SEO diagnosis of your content, and more of a way to pull out your first-hand material, the opinion and numbers that make a post worth reading.

Research handles the diagnostic part separately. Keyword research and content calendars are based on search data from the audience, not from what you assume they search.

So if your stated problem is not the real bottleneck, the data usually surfaces that on its own.

1
回复

the interview-first angle makes sense as a way to avoid the generic AI-blog-post feel. how long is a typical interview session though - is this a quick 5 question voice thing on your phone, or does getting genuinely useful first-hand detail out of someone take more like a 20 minute back and forth?

1
回复

@omri_ben_shoham1 The default is set to a short one, but it offers to go on if youre up for it - so it really does not have an upper limit on how long / deep the interview goes.

What genuinely surprised me is that how even a small 3 question back and forth can shape how an article turns out.

Another thing I'd like to point out is since all interviews go into the Knowledge Base there is continuity in the interviews - so over time the small interviews also add up to get really meaty.

1
回复

Dyuti, you asked why people quietly stop publishing AI drafts, so here is mine: every draft read as competent and generic at once. No number I had earned, no opinion I would defend in a room, so putting my name on it felt like a small lie, and fixing that by hand cost more than writing from scratch.

The interview flow is the right fix, because the missing input was never the words, it was the first-hand material only I had.

One question, from running SEO in a narrow B2B healthcare niche: the posts that convert for me have almost no search volume and rank on depth, not keywords. Does the research chase volume, or can it work the low and zero-volume long-tail questions real buyers actually type?

1
回复

@clemente_lopez1 

On volume: So search volume is a lagging, averaged signal. In a narrow B2B niche the queries that convert often show as zero in every tool because the tools round down anything under ~10/month. So no, the research doesn't rank topics by volume alone. The interview + your knowledge base surface the questions your actual buyers ask (which you know and no keyword tool does), and the keyword research is there to shape the piece around how those questions get phrased

Honestly, healthcare B2B specifically is a good stress test for us - esp. if you try it on one of those zero-volume posts, Would love to hear whether the interview pulled out the depth that would make the post useful and make it rank!

PS: "Competent and generic at once" is going on a wall somewhere

🙂🙂
1
回复

The interview-first approach is the real unlock here. Most AI SEO tools start from keywords and end up sounding like the same recycled SERP summary. Pulling in the founder’s actual opinions, examples, and counterintuitive takes before drafting feels much closer to how useful blog posts get made.

The GitHub/Markdown connector is also interesting for custom blogs. Curious how much control users get over the final publishing flow: can teams keep approval gates and edits inside their existing repo workflow before anything goes live?

1
回复

@aditya_harish_2002  Thanks - that's exactly the bet we made.

Keyword-first tools converge on the same SERP summary because they all start from the same inputs. The interview is the only input your competitors can't have.

On the GitHub connector: yes - we have a review gate that waits for your feedback (as configured). Also drafts land in your repo as a pull request, so your existing review flow is the approval gate. Edit in the PR, request changes, merge when ready. Nothing goes live until the merge. ACME.BOT's own approval gates still apply upstream (before the PR is even opened) if you want a check earlier in the pipeline

1
回复

Love the interview mode and that the article is almost one shot keeping the needs of the brand (and that specific article) in mind.

1
回复

@varun_dhamija Glad you liked the interview mode. The interviews compound over time in your Knowledge Base. So as you keep using ACME.BOT to write more blogs the context will keep building.

1
回复

I always wanted more control over what I post. This looks interesting.

In the process of trying it out for my blog.

Ps - how can I make it work for my custom built blog?

1
回复

@mitalee_rao Hi Mitalee - What custom stack are you on? If your custom blog is on Github / with markdown, we have a GitHub connector. You can try it from the connectors page.

1
回复

The interview-first step is the interesting bit — most SEO tools skip straight to generating commodity content. Once I answer the questions for a topic, does it persist that expertise as reusable context for future posts, or is each article a fresh interview? And for the publishing cadence, which CMSs can it actually push to versus just draft?

0
回复

@noctis06 

re. Storing interviews: Yes, all your interviews go inside a knowledge base. So any future documents can reference it and cite it.

re. Publishing: Right now we support Wordpress, Shopify and GitHub.

0
回复

One thing I'm curious about is how you stop the content from sounding repetitive after a few months. Does the knowledge base help with that?

0
回复

@reda_roqai_chaoui Yes, the Knowledge Base helps - every time we do an interview we end up going one level deeper. So the knowledge from the interviews builds up being a rich and deep repository of your opinions and takes over time. Based on that the later posts will end up going deeper into the takes as opposed to the initial ones.

0
回复

the feature list has a "Humanizer - reduces the chance of being flagged as AI writing" line right next to the EEAT/no-slop pitch, and those two things sit a bit awkwardly together. if the content is genuinely built from a real interview with someone's first-hand expertise, why does the output still need to be disguised from AI detectors? that reads like the underlying draft is more AI-generated-then-dressed-up than the interview-first framing suggests. is the humanizer mostly smoothing awkward phrasing, or actually working to defeat detection tools specifically?

0
回复

@galdayan The motivation is to not disguise - it's actually to just iron out the awkward phrasing etc. The fact that it ends up beating AI detection is a side effect, not the explicit goal.

0
回复

Is it like SEObot?

0
回复

@carlos_mendez16 Not sure which one you mean - since there are a handful of products that go by the same name :)

What ACME.BOT is - proactive AI agent that sets your content schedule, pings you for your inputs when its writing, waits for your reviews / approvals and publishes. All from your phone. You end up publishing your own unique take ie. non-commodity content - not average of wikipedia

0
回复
#7
Gemini 3.6 Flash Family
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
188
一句话介绍:Gemini 3.6 Flash系列模型专为大规模AI代理构建而设计,通过优化效率、降低延迟和提升可靠性,解决了代理开发中响应速度慢、成本高、工具调用不稳定等核心痛点。
SaaS Artificial Intelligence Development
AI模型 大语言模型 代理开发 低延迟 高可靠性 效率优化 模型家族 安全护栏 成本控制
用户评论摘要:用户肯定Flash系列在代理任务中的延迟降低,但尖锐指出缺乏硬性数据(延迟、可靠性基准),认为发布更像营销而非工程。开发者强烈呼吁提供统一仪表盘,以实时对比各变体的成本、延迟和推理质量,并希望明确安全性与低拒绝率的平衡策略。
AI 锐评

此次发布看似是一次“全家桶”式迭代,实则暴露了Google在AI Agent赛道上的焦虑与短板。Gemini 3.6 Flash系列将“效率、延迟、可靠性”作为核心卖点,切中了开发者构建多步骤代理时API响应不稳定、工具调用断裂的切肤之痛。然而,正如评论区多位资深开发者所指出的,整个发布充斥着“形容词”而非“数字”。没有公开的延迟百分位图、没有压力测试下的工具调用一致性报告、没有跨变体的生产成本对比表,所谓“为规模而生”的宣言就显得苍白无力。更致命的是,Flash-Lite与Flash-Cyber的定位模糊,用户需要自行从散乱文档中拼凑差异,这暴露了产品矩阵缺乏统一规划。这种“先发布、再细化”的作风,在竞争白热化的模型市场中是致命的信任消耗。真正的价值不在于模型本身,而在于Google能否兑现承诺:提供一套可观测、可调试、可预测的工程化工具链。否则,再快的Flash也只是停留在演示Demo中的“闪光”,无法在复杂的生产系统中点亮稳定的Agent之路。

查看原始信息
Gemini 3.6 Flash Family
We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber. Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale.

Congrats@sundar_pichaion the launch. Exciting to see 3.5 Flash Cyber. Since it's gated to governments/trusted partners for now; is wider access on the roadmap, or is this meant to stay a permanent limited-access tool given the dual-use risk?

3
回复
How the reasoning of 3.6 compare to previous 3.5 flash models for task like high level coding, edits, fixes and reviews?
0
回复

Finally gave the new Flash a spin on a quick agent task and the latency drop was honestly noticeable, made the whole loop feel snappier than I expected.

0
回复

Adding first-class observability dashboards built into the model API would be huge. Real-time token usage, latency breakdowns, and failure pattern detection per agent run would make it way easier to debug and optimize at scale.

0
回复

The positioning around “efficiency, latency, and reliability for agents at scale” feels incomplete. Right now every claim is qualitative no hard numbers, no latency charts, no reliability benchmarks, no failure‑mode transparency.

Developers building multi‑step agents don’t just need fast models; they need predictable tool‑use behaviour across long chains. Without published consistency metrics, uptime guarantees, or real‑world agent stress tests, it’s impossible to evaluate whether 3.6 Flash actually solves the reliability gaps that break production agents.

Also: Flash‑Lite and Flash‑Cyber are introduced as if they’re cleanly differentiated, but the docs force us to dig through scattered pages to compare cost, reasoning quality, and safety behaviour. A unified dashboard isn’t a “nice to have” it’s essential.

Right now this launch reads more like marketing than engineering. If Google wants developers to commit workloads, we need numbers, not adjectives.

0
回复

Been testing flash models quite heavily lately. The lite variant is interesting for high volume use cases where cost matters more than raw performance. How does 3.6 compare to 3.5 on reasoning tasks?

0
回复

Love the focus on efficiency for agent workflows. One thing that would help me as a developer is a built-in token usage dashboard per model variant, so I can compare cost and latency between Flash, Flash-Lite, and Flash Cyber in real time during testing without needing separate tooling. Would save a lot of guesswork when picking the right model for each part of a pipeline.

0
回复

A unified dashboard to compare latency, cost, and quality across the three Flash variants side by side would be huge. Right now I have to dig through docs and run my own benchmarks to figure out whether Flash Lite or Flash Cyber fits a given workload, and a built-in comparison view with sample prompts would save a ton of evaluation time before committing to a model.

0
回复

The latency/reliability positioning is the part that matters most for agents. Cheap tokens are nice, but agents fall apart fast when tool calls get slow or flaky.

Curious how you’re measuring reliability here: is it mostly uptime/API stability, or are you also tracking things like consistent tool-use behavior across long multi-step agent runs?

0
回复

BTW, I'm also curious how you balance "stronger safety guardrails" with "fewer refusals" cos those feel like they'd pull in opposite directions. How do you know when you've got that balance right?

0
回复
#8
MonoCloud for Startups
One identity layer for your customers, APIs, and agents
176
一句话介绍:MonoCloud为初创企业提供一个统一身份层,不仅管理客户登录,还授权和审计API、AI代理的访问权限,解决多工具拼凑带来的安全与合规痛点。
SaaS Developer Tools Security
统一身份平台 客户身份认证 API安全 AI代理权限 细粒度授权Cedar 证书绑定mTLS SSO 审计溯源 M2M通信 免费一年
用户评论摘要:用户关注AI代理的问责性、基于Cedar的授权下放、mTLS撤销延迟、SPIFFE跨环境兼容、定价连续性及代理链审计。团队回应强调支持引用令牌即时撤销、下放权限不超用户上限、Cedar策略强制执行。
AI 锐评

MonoCloud切中的是一个尴尬的中间地带:小工具止于登录,大企业方案臃肿贵胄。它声称用Cedar做细粒度授权、mTLS做信任锚点、SPIFFE做身份证明,从一开始就把AI代理列为“一等公民”,试图填上大多数平台对“事后授权与审计”的漠视。这确实是有眼光的选择——当代理开始自主行动,谁负责和如何撤销成了致命问题,MonoCloud至少给出了完整的技术解答。

但问题也在此:它把“身份”做成了准操作系统级的平台,而非可插拔的模块。开发团队要考虑的是,是平白增添一个全权信任的中间层,还是忍受五套工具的拼凑风险?框架够“意见”少,但绑定深——Cedar策略必须与MonoCloud的令牌发行挂钩,才能完整发挥。这意味着一开始选它,后续逃离将成本高昂。

另一个隐忧是:免费一年养熟了用户,但收费后能否续命?留言区已有人问“第二年涨价多猛”,尽管团队答得在情在理,但定价敏感期的初创企业终究是走是留,取决于产品能否真正成为“不可替代的基建”——而非另一个“免费上瘾,付费戒毒”的SAAS陷阱。此外,当前仅有10个设计伙伴验证,尚缺大规模压测下的稳定性与性能数据,评论区已有人关注单点故障。整体看,方向正确,但产品尚在早期,值得跟进,不必盲信。

查看原始信息
MonoCloud for Startups
MonoCloud is one identity layer for your customers, your APIs, and your agents. Most tools stop at a login box. We go past login into authorization and accountability: decide exactly what every user, service, and AI agent can access, prove what it did, and revoke it in an instant. Fine-grained Cedar authorization, passkeys and SSO, API protection, M2M, and mTLS with certificate-bound trust, all on one platform. Startups get the full platform free for one year.

Hey Product Hunt! 👋 Riya here. I do product marketing at MonoCloud, and I am one of the makers.

Most of my week is spent talking to founders about auth, and I keep seeing the same arc. In the beginning nobody wants to think about it, so they drop in a login box, ship, and move on.

Six months later, those same people are drowning in it.

A customer wants SSO.

A security review shows up with forty questions.

Their services need to talk to each other safely. And their AI agents are now acting on behalf of users with no real control over what they can reach.

The reason it keeps happening is that most identity tools were built for one slice of the problem.

The easy ones stop at the login screen.

The enterprise ones are powerful but slow and expensive to adopt. So teams end up owning the hard middle themselves, which is the exact part they have no time to build.

MonoCloud covers all of it on one platform, your customers, your APIs, and your agents, from your first user to enterprise scale. Not just who signs in, but what every user, service, and agent is allowed to do afterwards, with the ability to prove it and pull it back in an instant.

All of it is live today, not a demo or a waitlist. Ten design partners across SaaS and AI companies are already building on it.

👉 If you are a startup, you get the full platform free for a year, every feature, no gates. Apply here: MonoCloud for Startups - Free for One Year 👈

And if identity is central to what you are building and you want to shape where this goes, we are opening a few more design partner spots. Leave a comment or reach me on LinkedIn (link in profile).

I will be here all day and would genuinely love to know how you are handling identity for your AI agents right now, and where it falls apart first. That is the conversation I care about most.

7
回复

Hey Product Hunt 👋 Vishal here, founder of MonoCloud.
Every company is racing to put AI agents into production. Almost none of them can answer one simple question. When an agent does something it should not, who is accountable?

The log will say an agent did it. But an agent has no intent. Someone gave it that access, and today most teams cannot say who, cannot prove it, and cannot pull it back fast. That gap is the biggest unsolved problem in identity right now, and it is why I built MonoCloud.

I have built authentication into products four times over my career. Every time, the login box took a weekend and everything that mattered took months. Passkeys alone were over 800 engineering hours before they were production ready. I learned the hard way that identity is not a feature you bolt on. It is infrastructure, and it decides whether you can be trusted.

MonoCloud is that infrastructure. One platform for your customers, your APIs, and your agents, where you decide exactly what each can access, prove what it did, and revoke it in an instant.

Two convictions shaped it, and I still stand by both:

🔹 Build on open standards, not shortcuts. Authorization runs on Cedar, workload identity on SPIFFE, and trust on mTLS with certificate-bound tokens. Anyone can demo a policy dashboard. Far fewer hold up in a real security review.

🔹 Treat agents as first-class identities from day one. Every agent gets a scoped identity, acts on behalf of a user only within the limits you set, and leaves an audit trail you can revoke in one click.

All of it is live today. Ten design partners across SaaS and AI companies are building on it now.

👉 Startups get the full platform free for one year, every premium feature, no gates. Apply here: MonoCloud for Startups - Free for One Year 👈

If identity is core to what you are building, we are opening a few more design partner spots. Leave a comment or reach me on LinkedIn.

I will be here all day. The one thing I want to hear is this. How are you handling identity for your agents today, and what happens the first time one does something it should not?

6
回复

@maganuk The accountability question is the real one. "Accountability, judgment, and ownership still belong to people" isn't just a nice line, it's the actual gap most teams are ignoring while they rush agents into production.

Here's the layer I'd want to understand for client delivery work specifically: when we build an agent for a client, who's the principal it's acting on behalf of, the client's end user, or us as the vendor who built and deployed it? That distinction matters the moment something goes wrong and someone has to own it.

0
回复

Hey Product Hunt! 👋 Shivangi here, building MonoCloud with our founder Vishal @monocloud.

Auth always looks finished the day login works, but it never is. A customer asks for SSO. Your services need machine-to-machine keys. RBAC roles stop being able to express what you actually need. And now AI agents are acting on behalf of your users, and nobody can say what they are allowed to touch. Each piece gets bolted on separately, and a year later "auth" is five half-built tools held together with glue.

We built MonoCloud so you do not do that. It is one identity layer for your customers, your APIs, and your agents. Login is table stakes. The part that matters is what happens after: you decide exactly what every user, service, and agent can access, prove what it did, and revoke it in an instant, before an agent is ever issued a token.

Most tools make you pick a side. The login-box platforms stop at the sign-in screen. The enterprise suites are powerful but heavy and built for large companies only. MonoCloud is the one platform that covers customer identity, API protection, and agent identity together, from your very first user all the way to enterprise scale.

What you can do with it:


🔐 Every way your customers sign in: passwordless, passkeys, OTP, social, and SSO, ready on day one
🧩 Fine-grained authorization with Cedar, applied at the token, not just roles that stop being granular the moment your app grows
🤖 Identity for AI agents and workloads, with scoped access, on-behalf-of delegation, M2M, and SPIFFE, so an agent can only ever do what you allowed
📜 Certificate-bound trust and mTLS with instant revocation, plus full audit logs, so you can prove who did what and cut off access the second you need to
🛡️ Step-up authentication that forces fresh verification before high-risk actions like payments, transfers, or admin changes, plus attack and brute-force protection built in

We are building with 10 design partners across SaaS and AI companies right now, and their feedback is shaping what ships next. If identity is core to what you are building and you want to shape where this goes, we are opening a few more design partner spots. Leave a comment here, or reach me on LinkedIn, and I will follow up.

👉 For the Product Hunt community, startups get the full MonoCloud platform free for one year, every premium feature, no gates. Apply here: https://tally.so/r/ZjKMKB 👈

We will be here all day. The one thing I would really love your feedback on: how are you handling identity for AI agents right now, and where does it break first? That is the piece we are building hardest, and honest answers help us more than any upvote.

I am happy to get into anything, the product or the roadmap.

3
回复
@roguetink One platform for users APIs and AI agents sounds super convenient but how do you prevent single point of failure issues as workloads scale heavily?
0
回复

Congrats on the launch identity is one of those problems that only looks simple until you have to build it yourself Really Building authentication and identity is never as straightforward as it seems Excited to see products taking a fresh approach here Congratulations to the team and hope you have a fantastic launch

3
回复

@suryansh_tiwari2 thank you so much. it always looks like a solved problem until you're actually in it 😄 really appreciate the support!

1
回复

The agent-on-behalf-of-user piece is what most auth stacks punt on, so going straight at it with Cedar is the right call. Concretely: when an AI agent acts for a user, is its access always a downscoped subset of that user's own permissions (so it can never exceed what the human could do), and can you express that delegation in a single Cedar policy? And if you revoke mid-task, does an already-issued agent token stop on the next call or only at expiry?

3
回复

@hazy0 Good question, this is exactly what our API Access Policies are built for. Cedar runs at token issuance (the IssueAccessToken action), so it gates what the agent can get: which API, which scopes, which token settings.

 
Downscoping: yes, as a ceiling you author, default-deny, and forbid overrides permit. The agent is the principal, the user rides along as context.user, and requested scopes come in as context.requested_api_scopes, so a single permit can require that a real user is present, that the user is in an eligible group, and that requested scopes never exceed the ceiling you set. 

 

Mid-task revocation depends on token type. Reference (opaque) tokens are checked against server state on every call, so revocation is immediate, the next call fails. JWTs are self-contained and live until exp. For agents that need instant revocation: set AccessTokenType = Reference, keep lifetimes short, and optionally BindTokensToSession so ending the user's session kills the agent's tokens too. Policies are re-evaluated on every issuance and refresh either way, so pulling the agent's permit stops new tokens at the next refresh — only already-minted JWTs survive to expiry.

 

Happy to walk through your exact agent setup on a call.

1
回复

Interesting that mTLS with certificate-bound tokens is a first-class thing here rather than an enterprise add-on. Most CIAM tools treat that as an afterthought. What's the story for revocation latency — how fast does a revoked cert actually stop working in practice?

2
回复

@maurya_abhiranjan Revocation is checked live at validation time, before a client can get or use another token, not left until the token expires. You've got live OCSP for status, CRLs, and a local deny list to block a cert outright. How fast it actually takes effect comes down to the cache windows you control on the trust store, OCSP status is cached five minutes by default and CRLs fifteen, and you can tighten those toward near-instant. The full set of mTLS validation and revocation settings is laid out here if you want it, monocloud.com/features/mtls

0
回复

Amazing, the product looks so good, the tiers are so generous, so excited to try it out 🥳

2
回复
@itsharshag Thank you. Would love for you to try it and share feedback with us.
0
回复

Is the SPIFFE workload identity tied to a specific runtime, or does it work across mixed environments (some services on Kubernetes, some on bare VMs)? Trying to figure out how much of my setup I'd have to standardize.

1
回复

@suyash_kr Not tied to a runtime, it's SPIFFE-based, so anything with a valid SVID works, whether it's a Kubernetes pod or a workload on a VM. If they share a trust domain one trust store covers both, and if they're separate domains it's one trust store each. The setup walkthrough is here if it helps, monocloud.com/docs/guides/workload-identity

0
回复

First of all a great congrats to @riya_pariyar @roguetink @maganuk
I am just curious that you've said you're building the auth layer for agentic products "from the ground up" instead of bolting agent support onto human auth. Concretely, what does a token look like when an agent calls another agent three steps downstream? How do you actually scope and audit that so the blast radius doesn't become "every permission that token ever carried"?

1
回复

@arduvey29 Thanks! A token never inherits permissions, it's decided fresh at every issuance under default-deny, and you can forbid a sensitive API from ever sharing a token with others, so a token can't quietly widen to reach something it shouldn't. Tokens issued over mTLS, including X.509 SVIDs, are certificate-bound, so a copy is useless without the matching cert. If you want to see how that's written, there's a full guide with example policies at monocloud.com/docs/guides/api-access-policy

0
回复

I like the emphasis on authorization over authentication. How opinionated is MonoCloud about application architecture? Could a team adopt just the Cedar authorization layer while keeping their existing auth provider, or is the biggest value unlocked by using the full identity stack?

1
回复

@tarqiya_forgah Not very opinionated, it's standard OAuth and OIDC with drop-in SDKs, so you adopt it piece by piece rather than rewriting your app. On the Cedar part, to be precise, the authorization policies are evaluated at the moment MonoCloud issues the access token, so the authz layer and token issuance go together. The full value shows up when MonoCloud is in the token path for the APIs you want to govern. If you're thinking of running only the policy layer on top of a different auth provider, I'd love to get into your exact setup.

0
回复

the one year free offer for startups is honestly such a smart move, basically lets you build the whole stack without cobbling together five different auth tools.

1
回复
@fakkelek20444 That's the exact sentence we want people saying a year from now. Cobbling five tools together is where most identity bugs are born, the seams between tools are where things fall through. Thanks Fatma.
1
回复

How does pricing work after year one? Curious if startups that grow during the free period end up sticking or if the jump to paid feels steep.

1
回复
@talhakhalidmtk Fair question and honestly the one we thought hardest about. After the year, startups move to the plans. The goal is that by then you're paying because the product earned it, not because migration is painful. If the jump ever feels steep for a team, we'd rather talk than lose them, so reach out before the year ends and we'll figure it out.
0
回复

Congrats on the launch. The customer/API/agent identity angle is timely because agent access can blur normal user boundaries. How do teams inspect which identity performed an action, which permissions were used, and whether the action came from a human or an agent?

1
回复

@yaroslav_stelmakh Good question, and it's a real gap most setups have. Every token request is audited with the acting principal, the trust store and certificate context for mTLS, and the policy evaluation itself, so you can see who acted and whether it was allowed. Human versus agent is built into the model rather than guessed, clients carry a type of either application or agent, and machine-to-machine requests carry no user in the request, so an agent action looks structurally different from a human one. The permissions side is the Cedar decision that gets logged with the request. The client-type and context details are here if you want them, monocloud.com/docs/guides/api-access-policy

0
回复

honestly the one free year for startups is kind of a big deal, but what really stands out is the mTLS with certificate-bound trust. most platforms treat that as an enterprise add-on and bury it in pricing. shipping it on day one for early teams shows you actually thought about the security stack instead of bolting it on later.

1
回复

@kervangamz80440 Thanks, that was a deliberate choice. Certificate-bound tokens make a stolen token useless without the key, and that shouldn't be something you only get once you're big enough for an enterprise tier. The free year for startups includes it, so early teams can build on cert-bound trust from the start instead of retrofitting it later.

0
回复

Identity layer that covers customers, APIs AND agents in one place is actually pretty rare. Most solutions handles only one or two of those well. Free for a year is generous, what happens at the end?

1
回复
@abdurrahman_fakhrul Thanks Abdurrahman, and you're right that most tools pick one lane. Customer identity, API auth, and agent identity share the same underlying primitives, so splitting them across three tools always felt like the wrong default to us. On the free year: after it ends, startups move to our paid plans. We kept the jump deliberate and gradual because a team that grew with us for a year churning over sticker shock would mean we failed at pricing, not them. If it ever feels steep, reach out before the year ends and we'll figure it out.
0
回复

Congrats on the launch - it's an impressive offering!
I'm curious how MonoCloud tracks the agent acting on behalf of the user - is the agent treated as a separate identity?
And can I use Cedar to manage access to application-level entities? How does that work in practice?

I'm currently using Clerk, whose separation between development and production environments makes it easy to copy the configuration to production when launching. Do you offer a similar workflow or support migrations from other authentication providers?

1
回复

@mateuszkonik Yes, the agent is treated as a separate identity. It doesn't reuse the user's token. The agent obtains its own token to act on behalf of the user, so agent activity is always distinguishable from the user acting directly. Cedar comes in at token issuance: each API (a resource client in MonoCloud) can have multiple Cedar policies attached, and whenever that API is requested  via the authorization or token endpoint, all of its policies are evaluated first. A token is issued only if every policy passes, so access to an API is gated by policy right at issuance. As for migrations, you can bring users over programmatically via our Create User endpoint (MonoCloud - Authentication Platform for Developers | SSO, Passkeys & Next.js - you can use user claims, private data and public data fields to capture all of the users properties from your current provider. Happy to dig into any of this further, and there's a lot more detail in the docs (https://www.monocloud.com/docs).
I'll be happy to walk you through the setup.

1
回复
Honestly, right now it's the unglamorous version: a single Supabase service role key that our enrichment pipeline uses to call the Claude API and write back to the database. It works because it's the only "agent" we run today, but it's already the wrong shape that key can touch every table, not just the columns it needs to enrich. Where it'll break first for us: the moment we add a second automated process (say, something that also handles billing or subscription logic), we've got two different "agents" sharing one all-or-nothing credential. There's no scoping, no per-task revocation, no audit trail beyond generic DB logs. Right now "identity for agents" for us just means "hope the key doesn't leak," which isn't a real answer, it's a gap.
1
回复

@thys_beesman You are right that it is already the wrong shape. A single service role key with full reach is where almost everyone starts, so you are in good company.
The second process is where it really starts to hurt. Once a billing job and an enrichment job are sharing one credential, a leak means you cannot even tell which one did what, and you end up rotating the key on both.

The way we think about it, each process should get its own scoped identity instead of everyone leaning on one key. You give that identity a policy for what it is allowed to do, so the enrichment worker can only enrich and the billing worker can only touch billing. If one of them is ever compromised, you revoke just that identity without taking the others down, and every action is logged against that specific agent instead of getting lost in generic DB logs.

I am curious, are these long-running services or short-lived jobs, and where do they run, somewhere like Kubernetes, Fly, or Railway? That is what will decide whether workload identity with SPIFFE or short-lived machine-to-machine tokens is the cleaner path for you.

3
回复

the instant-revoke plus proof-of-what-it-did combo is the part that actually matters once agents are calling real APIs on someone's behalf. most identity providers treat agents like just another OAuth client and call it done. does the audit trail capture the actual reasoning/prompt context behind an action too, or just the API call itself? that distinction matters a lot when you're trying to explain after the fact why an agent did something unexpected.

0
回复

@omri_ben_shoham1 We audit the identity layer, and it's fairly detailed, the principal that acted, the mTLS and trust store context, and the policy evaluation itself, including whether the decision was allow or deny and why, for example a revoked certificate. What it doesn't capture is the model's prompt or reasoning, and that belongs in your agent or observability layer. You join the two by the agent's identity to explain why it acted.

0
回复

Everything in the thread assumes a human kicked the agent off. A good chunk of what I run is on a cron, hourly, nobody logged in, so there is no session to ride along as context.user, and that is exactly where "on behalf of" gets slippery: the agent is acting for an intent someone wrote down three weeks ago, not for a person who is awake to notice. Does that shape get its own principal, or does it collapse into a machine client with a service identity and you lose the delegation trail?

0
回复

@dipankar_sarkar When there's no human, there's no user in the request. Machine-to-machine grants carry no user context, so the agent authenticates as its own principal, an agent client or a SPIFFE workload, not an anonymous session, and what it's allowed to do is set by policy on that identity. The grant-type and context detail is all written up here if you want it, monocloud.com/docs/guides/api-access-policy

0
回复

Vishal — the step-up-before-high-risk-action line caught me. If an AI agent, not a human at a keyboard, is the one about to trigger something like a payment or a live bid submission, can step-up auth even work — or does that whole model assume there's a human around to re-verify?

0
回复

@medal411 Sharp catch. Classic step-up does assume a human is there to re-verify, so re-prompting a fully autonomous agent for a passkey doesn't make sense. Two things cover the agent case instead. When a human should sign off, you use human-in-the-loop consent, pausing the agent before the sensitive action and requiring a fresh approval. When it's genuinely autonomous with no human, you don't step up at all, you gate the action with policy, a Cedar rule that decides whether that agent, with that scope, is even allowed to trigger the payment, rather than re-verifying someone who isn't there.

0
回复

This is the unglamorous part that makes agent work production-safe. Small teams often start with one all-powerful service key because it ships fastest, then the second or third automated process turns that shortcut into real operational risk. Per-agent identity plus fast revocation is the boring layer that lets you delegate without pretending the risk disappeared.

0
回复

@krekeltronics Exactly this. The one all-powerful key is where almost everyone starts because it ships fastest, and it stays invisible right up until the second automated process has to share it. Scoped identity per agent plus revocation you can actually trigger is what makes delegating feel safe instead of just hoping. Well put.

0
回复

the agent-as-a-separate-identity model is the right architecture, and Cedar for token-issuance policy is a solid choice over roles that stop scaling. the question I didn't see answered yet is the business side of the free year: what does year two actually look like pricing-wise for a startup that's grown into real usage by then, is it a predictable seat/request-based scale-up or more of a cliff where the free tier suddenly becomes an enterprise sales conversation? that's usually the thing that decides whether teams migrate off something core like identity later out of pricing fear rather than product fit.

0
回复

@galdayan Fair question, and the right one to ask about anything as sticky as identity. There's no cliff. When the free year ends you move onto our standard published plans, not into an enterprise sales conversation, so it's a predictable, self-serve scale-up rather than a renegotiation. We kept the step up gradual on purpose, because a team that grew on us for a year and then churned over sticker shock would mean we got pricing wrong. If the jump ever looks steep for where you are, reach out before the year's up and we'll sort it out. The whole point of the free year is to take pricing off the table as a reason to ever leave something this core.

0
回复

Lessssgooooo MonoCloud

0
回复
#9
Trovio For Brands
Communities love your brand - Our AI team finds them for you
170
一句话介绍:Trovio For Brands 是一个AI驱动的社区发现与创作者营销平台,帮助品牌绕过传统的搜索关键词和头部KOL,精准定位产品实际用户所在的小众信任社区,并匹配高转化率的小型创作者,实现从触达到转化的全链路追踪。
Marketing Influencer marketing Social media marketing
AI创作者匹配 社区营销 小众KOL 转化追踪 品牌合作 精准营销 微影响力 信任经济 内容电商 数据分析
用户评论摘要:用户高度认可“社区匹配”逻辑,提出担忧:如何识别真正欢迎品牌的社区而非抵触者?如何检测虚假粉丝?能否支持端到端转化追踪?如何保证推荐质量并支持品牌手动调整?对小型创作者的转化效果表示好奇,并询问收费模式(非抽成,订阅制)及平台方如何动态更新社区图谱。
AI 锐评

Trovio的价值不在于又一个“AI匹配工具”,而在于它精准戳破了营销行业一个浮华的泡沫:粉丝数并不等于购买力。在流量红利见顶、用户对硬广免疫的当下,它承认了一个朴素的商业真理——人们只买他们信任的人的推荐。

它的聪明之处在于重构了“发现”逻辑。传统模式是在“创作者分类”里找网红,本质是“人找人”;而Trovio跳到了“在社区里找信任源”,变成了“价值网找人”。这意味着一个卖蛋白粉的品牌,不必再去死磕健身博主,而是能找到每天站10小时的小护士,因为这个社区的人(疲惫的上班族)才是真正的核心用户。这种跨越表面标签的“跨品类匹配”,才是其AI在商业洞察上最犀利的地方。

但冷静来看,产品目前仍存在巨大挑战。首先,社区生态的动态平衡极难维持:一个社区昨天欢迎品牌,今天可能因为一条生硬广告就集体反感,仅靠“API数据和内容信号”来判断社区红线的精度存疑。其次,它声称“不做抽成、只收订阅费”是一把双刃剑——这有利于吸引中小品牌和素人创作者,但也意味着平台的增长引擎完全依赖于能否持续证明“小创作者+精准社区”的转化率远高于传统模式。一旦某次匹配失误导致转化崩塌,品牌方的信任成本极高。

Trovio目前更像一个“高效的筛选漏斗”,而非“自动印钞机”。它解决了发现效率,但尚未完全回答品牌方最深的恐惧:如何确保每一笔钱都花在“能带来真实销售的信任”上,而不是另一场精美的数据包装。

查看原始信息
Trovio For Brands
A nurse on hour ten of her shift sells your protein bar better than any athlete. Gone are the days when mega creators moved the needle on conversion. Today people buy from who they trust: small, close-knit trusted communities. Trovio finds those communities, and their influencers who actually use and need your product. - Influencers matched on community in seconds - Post a brief, set terms, and we handle the rest - Full analytics and conversion tracking Trust converts. Reach doesn't.

Hey Product Hunt!

I'm Andrew, co-founder of Trovio. Last time we were here, we launched Pitch Kit to help creators land brand deals without an agent. Today we're launching the other side of the marketplace: the Trovio Brand Dashboard.

Here's what surprised us building the creator side. The best creator for your brand often isn't in the category you'd expect. They match your customer. We keep seeing it: a streamer who sits eight hours a day is a genuinely strong fit for comfortable activewear. A budgeting app's most convincing voice turns out to be a debt-payoff creator, not a finance one. Not because category creators are bad, but because the audience that buys often lives somewhere else.

Search struggles to find those people. A search bar retrieves creators by the label they filed themselves under, and your customers don't live in one label. They live in communities: comfy-clothes people, no-time-to-eat people, always-in-transit people. One community pulls from a dozen niches, so a lot of the best matches never show up when you type a keyword.

So we built the Brand Dashboard around communities instead:

  • Tell us about your brand and we surface the communities your customers are already in, the influencers they trust, and the reason each one fits

  • Post a brief, set terms, and manage every conversation, deliverable, and deadline in one place

  • Track what actually matters: conversions, not just reach

We'd especially love your take on:

  • How community matches feel vs. your current discovery process

  • The match that surprised you most, and whether the "why" made sense at a glance

  • What metrics you'd want in the dashboard that aren't there

Ask us anything, and thanks for checking out Trovio!

10
回复

@andrew_lukas We've paid people to do exactly this by hand, and the hard part was never finding communities, it was reading which ones would actually welcome a brand versus quietly resent it. How do you tell an engaged community from one that turns on outside brands the moment you show up?

4
回复

The filename being renamed is normal on many sites. Check the Network tab in DevTools during the upload—the response often includes the new file URL or ID. If it doesn't, the files are likely stored privately and not directly accessible.

0
回复

@andrew_lukas big fan of yours! let's go gener8tor companies!

1
回复

How do you handle end-to-end conversion tracking via custom affiliate links or promo codes? congrats @andrew_lukas on hitting product hunt today🙌

5
回复

@priya_kushwaha1 Short links, mostly. We don't really need promo codes to get there. The link carries the attribution, so you get click through to conversion on the brand's side without asking a creator to remember a code or a customer to type one in.

The other half is that our creators are actually on Trovio with their accounts connected, so we can pull the post-side numbers straight from the platform instead of relying on screenshots or self-reported stats. So you get both ends... what the post did, and what it drove.

Thanks again for checking it out!

2
回复

How does a brand know a creator's audience is even real? Fake followers and bot engagement are everywhere.

4
回复

@tyler_bush Great question, and honestly it's one of the biggest problems in the space. A few layers to how we handle it. First, creators on Trovio connect their accounts directly, so everything we work from comes straight from the platform APIs and refreshes nightly. No screenshots, no self-reported media kits. Second, our matching leans on engagement quality and consistency over time rather than follower counts. Bought followers are pretty loud in the data: big reach numbers with engagement that never follows.

There's also a quieter filter built into the model itself. Every creator on Trovio is a paying subscriber actually running their business here, and bot farms don't pay a monthly subscription to manage a fake audience. And the final backstop is that the dashboard tracks conversions, not just reach. Bots don't buy things. If an audience isn't real, it shows up exactly where you're looking.

2
回复

Sweet! Congrats! I’ve not seen someone be able to truly do attribution on creator marketing. How do you do attribution?

4
回复

@olsoj029 Yeah, honestly that's the thing that bugged me most about this space. Creator attribution is usually either a vanity metric (views, engagement, "estimated reach") or a pile of UTM params that nobody sets up consistently and that get stripped the second someone copies a link.

We do it at the link layer instead. Every creator on a campaign gets their own short link... clean enough to say out loud or drop in a bio, with the attribution baked into the link itself rather than a query string trailing off the end.

Once you have that, the more interesting part opens up. You can see which creators are actually driving results, not just impressions, and that becomes real feedback for the brand... who to re-sign, who to go deeper with, who's worth building an actual relationship, etc.

And because creators tend to cluster into communities, once we know someone's performing in a given community we can point the brand at other creators in that same community and then measure the whole group against each other. That's the part I'm most excited about!!

Thanks for checking us out!

1
回复
This is awesome! Do smaller creators actually convert for brands, or do you still need recognizable names? I feel like I only see big creators mostly in ads still.
4
回复

@steven_snider Thank you!! Totally fair question, and honestly it's why ads still look the way they do: big names are easier to find and easier to justify internally. But the conversion math points the other way. Based on industry benchmarks (HubSpot's Influencer Marketing Report + Influencer Marketing Hub's 2025 benchmarks), $50K on one macro creator gets you about 2.6M views and roughly 1,150 conversions. That same $50K spread across 10 smaller creators gets fewer views (about 1.9M) but roughly 4,700 conversions. Fewer eyeballs, about 4x the buyers, because people actually act on recommendations from creators they trust.

The reason brands haven't done this already is logistics. Finding, vetting, and managing 10 creators used to be way more work than booking one name. That's the part Trovio takes off your plate.

2
回复

This is cool! I could see us using this. How does the money work, does Trovio take a cut of the deals?

4
回复

@sean_adams_pantano1 Thanks Sean!! Great question. We don't take a cut, and honestly that's the core of the whole model. The terms you set in your brief are exactly what the creator gets, no markup on either side. This allows us to run on a subscription model vs. only chasing 6 figure deals.

1
回复

Is it possible to switch the creators I see if I don't like the initial recommendations?

3
回复

@shane_mueller Absolutely - you get ultimate say on who you decide to work with and the conversations you think match best to your brand!

1
回复

Hey Andrew, as we know communities are changing often how do you stay on top of the changes being made to make sure the communities people want to reach are active and accept brands to potentially collaborate with

3
回复

@niallcleaver Great question. Since we have all of our creators connected to Trovio - and not just a scraped list, we actually keep up with each and every post and their engagement so we can see over time as niches, communities, and attitudes change and then change in real time with them, so the brands are never getting stale matches.

1
回复

Finally tried posting a brief last night and was impressed how quickly it matched me with a niche fitness mom group instead of random big accounts. The analytics actually show real conversions which is rare.

3
回复

@doaniikrinu Thanks for checking us out!!

0
回复

@andrew_lukas How does Trovio learn about my brand and the communities that might fit us?

3
回复

@jake_whitman Good question! A bit of both. You give us the starting point, basically what you sell and who it's for, and honestly even just your site gets us going. Then we do the research from there. The other side of Trovio is thousands of creators with their real accounts connected, so we've built a map of communities from their actual content and who engages with it. We place your brand into that map and surface where your customers are already hanging out, the creators those communities trust, and why each one fits.

The fun part is the matches you wouldn't have searched for. Your customers usually live in way more communities than your category suggests, and that's exactly what keyword search misses. Every match comes with the why in plain language, so you can gut-check it instead of trusting a black box.

0
回复

The community-first matching is the sharpest part here. Most creator discovery still feels like searching categories and follower counts, when the real buying intent is probably sitting inside smaller, weirder communities.

Curious how you explain the match quality to brands at a glance. Is the "why this creator fits" mostly based on audience/community signals, past conversion patterns, or the creator's own content and product usage?

3
回复

@aditya_harish_2002 Really thoughtful question, and the honest ranked answer is: mostly community and content signals today, with conversion patterns as the layer that compounds over time.

Concretely, each "why this fits" is built from where the creator's audience actually lives (the communities they're part of and who actually engages with them) plus the creator's own content: what they genuinely make and talk about, which is how we infer whether they'd actually use and need what you sell rather than whether they tagged themselves in your category. We deliberately show those signals in plain language instead of one opaque score, because a surprising match without a legible why just reads as broken matching.

On past conversion patterns: that's the part that gets stronger as deals run through the platform, since every brief has conversion tracking attached. I'd rather be straight that we're early there than pretend we have years of conversion history. The nice property is it compounds, every campaign that runs makes the next match rationale sharper.

0
回复

How does it figure out which communities a brand's customers are in? What's that based on?

3
回复

@betsy_seus Good question. There are two ingredients. On the creator side, thousands of creators have connected their actual accounts, so we build the community map from their real content and who engages with it, not from labels people picked in a dropdown. That's what surfaces communities that don't match any single niche, like comfy-clothes people pulling from gamers, nurses, and new parents all at once.

On the brand side, you tell us about your product and who buys it, and we place your brand into that same map. So the matching happens in one shared space (communities), rather than trying to translate a keyword into a creator category. And every match shows the reasoning in plain language, so you can sanity check it rather than trusting a black box

0
回复

Congrats on the launch. The community matching angle is the useful part here because it moves the brief away from follower counts and toward actual trust. For brands that sell across several customer segments, do you let them define different community profiles inside one campaign, or is the matching model better when each brief stays focused on one clear buyer group?

3
回复

@wesc Thank you! This is an awesome question.

The matching itself already runs at the community level, and when you tell us about your brand we usually surface several distinct communities, because almost no brand maps to just one buyer group. So multi-segment discovery is native to how it works. For briefs, my honest take today is that one clear buyer group per brief performs better. Terms, creative direction, and what a good deliverable looks like usually differ by segment anyway, so a few focused briefs beat one broad one that averages across audiences. But letting brands define named community profiles and track them side by side inside one campaign is a really good idea. Noting it for the roadmap.

1
回复

Interesting! How much of this is AI deciding who represents a brand?

3
回复

@chase_mine Good question! The honest answer is AI does the finding, you do the deciding. Where the AI earns its keep is the discovery problem: mapping which communities your customers actually live in and which creators those communities trust. That's a many-to-many matching problem that a search bar (or a person with a spreadsheet) genuinely can't do at scale.

But nobody represents your brand without you choosing them. Every match comes with the reason it fits in plain language, and you review the matches, decide who to invite, set the terms in your brief, and approve the final deal. So the AI narrows the field from millions of creators to the handful worth your time, and the actual call is always yours. If a match ever feels wrong and the "why" doesn't make sense at a glance, that's something we'd genuinely want to hear about.

1
回复

Does this surface opportunities with those who have never monetized their content before? Curious to learn more about how it surfaces previously unknown communities for brands.

2
回复

@ihabib Thanks for the Question. We helped 3 creators just last week get their first ever deals. Honestly that's kind of the heart of the whole thing. The creator side of Trovio is built for the 99% who don't have representation, and a lot of them have never done a brand deal before. They have real audiences and real trust, they've just never had a way in, because the agency model only pencils out for the top 1% and everyone else gets ignored. Those are exactly the people who show up in your matches, and they're often the highest-trust voices you can get, precisely because their audience isn't used to seeing them do ads.

On surfacing unknown communities: the map is built from creators' actual content and who genuinely engages with it, not from categories anyone picked. Communities are many-to-many, one community pulls from a dozen niches, so when we place your brand in that map, you see communities you'd never have thought to search for. That's the whole point really. If you already knew to search for it, you didn't need us. Every match comes with the why in plain language so the surprising ones make sense at a glance.

0
回复

The "match on community, not category" thesis is the right one — for a launch the highest-converting voice is usually someone already trusted in an adjacent niche, not the on-topic mega-account. I d use this to seed a Web3 product launch where the real communities are Discord/Telegram-native rather than IG/TikTok. One concrete thing: does the matching reach those closed community spaces, or is it sourcing purely from public social graphs — and how do you screen out communities with inflated/bot membership before I put a brief and budget behind them?

2
回复

@hazy0 Really appreciate this one, and you clearly get the thesis. Honest answer on the platform question: today the matching is built on creators who connect their accounts directly, with Instagram and TikTok first-class and YouTube rolling out now. Discord and Telegram native communities aren't in the graph yet, but our whole approach is meeting creators where they are, so any platform where real communities live is on our radar. Web3 is honestly a great example of trust living in closed spaces rather than public feeds.

The connect-their-account part is also the answer to your bot question. We don't scrape public profiles and guess. Creators link their real accounts, so we see true engagement straight from the platform, not follower counts, and accounts flooded with bots are pretty loud in that data: big reach, engagement that never follows. That verification layer is exactly what we'd bring to any new platform we add, closed communities included, because putting a brief and budget behind an unverified community is the thing we're trying to completely eliminate.

0
回复

Love that you matched on actual community fit rather than just follower count — feels like the most honest take on influencer marketing I've seen in a while.

2
回复

@cemil1216871 Thanks so much!! We really appreciate it!

0
回复

Amazing update! Congrats! One QQ; Do you match brands with one creator at a time or can they spread their budget across multiple?

2
回复

@sarah_showers Thank you!! And definitely multiple. Honestly spreading the budget is kind of the whole thesis. When you post a brief we surface a set of matched creators, and you decide how many to work with, whether that's one or ten. Terms are per creator, so you can mix bigger and smaller partnerships inside the same budget, and everything runs from one place so ten conversations doesn't mean ten times the work.

The reason we push this way: by industry benchmark math, the same spend spread across a group of well matched smaller creators drives about 4x the conversions of putting it all on one big name. Fewer total views, way more buyers. A portfolio also derisks things, you learn which communities convert for you and put the next dollar there instead of betting everything on one bet working out.

0
回复

Finding communities that genuinely engage with your brand vs just mentioning it is hard problem to solve. Curious what data sources Trovio pulls from, is it mainly Reddit and Discord type platforms?

2
回复

@abdurrahman_fakhrul Not Reddit or Discord, actually. The communities we care about live around the creators themselves, so mainly TikTok, Instagram, and YouTube today, expanding to more platforms to meet our creators where they're at.

The bigger thing is how we match. A niche describes the creator ("fitness influencer" is a label someone filed themselves under). A community describes the audience. Those are different axes, which is why searching and filtering by creator tag feels productive but gets brands nowhere.

So say you make a protein bar. The obvious move is to go search athlete or nutrition influencers. What we'll surface instead is something like nurses... a nurse on hour ten of her shift sells your protein bar better than any athlete does. Same product, totally different community, and it converts because it's her actual life and not a sponsorship.

We can do that because of the creator side. We're already working with creators to build these communities, or they showed up with one already built, so we're not trying to infer a community from a pile of brand mentions. We already know what it is and who's in it, then we match that to what a brand is looking for.

0
回复

Would love a way to see the actual conversation thread or content context where an influencer mentioned a brand. Right now the analytics show reach and conversions but not the organic conversation that drove it, which makes it harder to brief creators on what messaging is actually landing in those communities.

2
回复

@ceyda299740 Yeah great feedback! When an influencer posts we do show and track the post, and the threads that emerge to show sentiment on how the message is landing beyond just the conversion and reach numbers. We can be better about saying that upfront.

1
回复

How would a brand know what's fair to offer a nano creator?

2
回复

@stefan_stumpfl1 Good question, and there's honestly no universal rate card, which is part of the problem. My advice is to price the work, not the follower count. Start from what you're actually asking for: how many deliverables, how much production effort, and especially usage rights. Rerunning a creator's content as ads is worth real money and it's the thing most brands underprice. For smaller creators, a fair offer is usually less about hitting a magic number and more about being clear up front: the deliverables, the terms, and the budget, instead of making them guess and negotiate blind.

This is also something we help with directly. Trovio guides you toward a fair range for what you're asking, and matching already takes budget alignment into account, so the creators you see are ones where your range and their expectations actually overlap. No markup either, so whatever you offer is exactly what the creator sees. From there creators can accept or counter, and honestly a clearly scoped offer to a handful of well matched creators will calibrate you faster than any benchmark article.

1
回复

Do product-exchange deals still work or do creators all want cash now?

2
回复

@yasir_khan Honest take: gifting absolutely still works, it just works for different reasons than it used to. Cash has become the norm for bigger asks, but for a lot of creators a product they genuinely want is a totally fair trade for a low-lift post, especially early in a relationship with a brand. We see almost 50% of our first deals on a product exchange.

The key is fit, which is really a targeting problem. A gifting offer to a random creator in your category reads as "work for free." The same offer to someone who actually lives in the community your product serves reads as a perk, because it's something they'd probably buy anyway. That's honestly one of my favorite side effects of matching on communities: when the fit is real, gifting and hybrid deals (product plus a smaller cash component) land really well, and some of the most authentic content comes from exactly those deals. On Trovio you set the terms in your brief, so product-exchange, hybrid, and cash briefs are all totally doable. My advice: product-exchange is a great way to start relationships and test fit, then put cash behind the creators and deliverables that prove out.

1
回复

I really like this! But, there are a hundred influencer tools out there, what's actually different about this one?

2
回复

Thanks,@laura_butler! Fair challenge, and honestly it's the right question to ask. Most influencer tools are some version of the same thing: a searchable database of creator profiles. You type keywords, filter by follower count, and get back creators sorted by the label they filed themselves under. The tools differ in polish, but the underlying model is retrieval by category.

Three things we do differently. First, we match on communities instead of categories. A niche describes the creator, but a community describes the audience, and one community pulls from a dozen niches. That's a many-to-many problem a search bar structurally can't solve, and it's why the best match for your brand often isn't in the category you'd search. Second, the creators on Trovio aren't scraped profiles. They're actually here, running their business on our creator side with their real accounts connected, so the data is first-party and the conversations happen in one place. Third, the incentives: we take no cut and add no markup, so a $500 brief is worth doing right, not just a six figure one. My honest take is that the difference isn't a feature, it's that the whole thing is built around who your customer trusts instead of who has the most followers.

1
回复

Really like the insight behind this. The best creator for a brand usually isn't the biggest creator. Finding trusted niche communities instead of chasing vanity metrics is where the real ROI comes from. Excited to see where this goes

1
回复

@suryansh_tiwari2 thanks! Appreciate the support. We think this direction unlocks a ton for creator and brand relationships.

0
回复

This is really interesting! So I understand you match brands with creators but who handles the exchanging of money and signing on the dotted line to get the work done? Excited to learn more.

1
回复

@vanessa_brost Thank you!! Great question. Both happen inside Trovio. When you post a brief you set the terms, and a creator accepting is the agreement, so the scope, deliverables, and price everyone signed up for are documented in the platform rather than living in some email thread. Payment runs through Trovio too. You're not chasing invoices or wiring money to someone you met last week, and the creator isn't chasing you to get paid, which honestly is half the reason smaller deals never used to happen.

And since we take no cut and add no markup, the number in the brief is exactly what the creator receives. We make money on the subscription, not the transaction. Would love to show you more, and happy to answer anything else!

1
回复

the community-not-category angle is genuinely different from the usual influencer platform pitch. one thing I didn't see covered: these community voices (a nurse, a debt-payoff creator) are often not professional influencers with a media kit and an agent who already know the FTC disclosure rules cold. does Trovio handle the #ad/#sponsored disclosure piece for them, or is that left to the creator to get right on their own? that feels like the kind of compliance gap that only shows up after a campaign, not during matching.

0
回复

Congrats on the launch!

Are there any industries where this community-based approach works especially well?

I’d also be curious whether the strongest matches tend to come from TikTok, YouTube, Instagram, or Twitch.

0
回复

Very impressed by how you are leveraging agents to solve a real gap. This was relatable to me:

“beginner-friendly strength training. In Maya's comments, women share things like “I finally walked into the free-weight section without wanting to turn around.” That's the specific conversation a women's apparel brand wants to be part of — the kind they'd never find by searching for “fitness creators.” Her agent reads those signals, matches her with the brand, and drafts the pitch in her voice. Sent the same afternoon.”

0
回复
#10
Light Flip
Nostalgia for a time before smartphones
147
一句话介绍:Light Flip 是一款极简主义的翻盖功能手机,旨在帮助用户摆脱智能手机的过度依赖,回归有意识的数字生活,在保持通话、短信等核心通信功能的同时,通过刻意限制应用和功能来防止无意识刷屏。
Hardware Cell Phone
极简手机 翻盖手机 数字排毒 功能手机 反智能手机 专注工具 防沉迷 去网络化 电子墨水屏 怀旧科技
用户评论摘要:用户普遍认同“数字排毒”价值,但核心疑虑集中在新品与原有Light Phone的定位差异:是否增添了更多功能(如导航、相机导出、联系人迁移)?用户担心怀旧外形会导致“悄无声息地增加功能”,破坏了原有的极简哲学,并关心非智能手机在过渡期的实用性。
AI 锐评

Light Flip 打出的是一张高级的“心理牌”——不是贩卖硬件,而是贩卖一种摆脱数字焦虑的幻觉。从评论看,这批核心用户并非抵制科技,而是厌恶“失控的便利”。他们真正需要的是一个“有尊严的牢笼”:功能足够满足生存底线(导航、联系人、播客),同时又强制隔离抖音和推特。

Light 的聪明之处在于,它没有像诺基亚那样做一款“更差的旧手机”,而是将反智能本身品牌化、社群化。靠147票的社区热度,印证了这一细分需求确实存在。但隐忧也十分明显:它本质上在赌“多少用户愿意为反主流而付费”。一旦Flip为了扩大市场开始偷偷加回浏览器、微信等功能,它将立刻从一个意识形态产品退化为一个丑陋的、功能残缺的安卓套壳,丢失所有稀缺性。目前最尖锐的质问正是来自早期信徒:你们会不会为了出货量而“漂白”极简主义?答案决定了这款产品是精神图腾,还是失败的情怀生意。

查看原始信息
Light Flip
Light is a radically different technology company. We design beautiful tools that respect and empower our users and our first product is The Light Phone.
Time and time again, we‘ve heard from people – especially the younger generations - who love flip phones, but the current options are poorly made, slow, unsupported, and often pre-installed with the same addictive apps built in. Light is taking everything that already works for tens of thousands of Light Phone users and is bringing that into a heavily desired yet timeless tactile form factor of a flip phone.
2
回复

As someone who has eyed a Light Phone but always worried about the day-one gap — what actually makes the cut on the Flip beyond calls and texts? Is there turn-by-turn navigation, and can I move my contacts and podcasts over without keeping a smartphone in the loop? The flip form factor is the part that would actually get me to leave the smartphone at home.

0
回复

How do you get the photos from the camera to your computer?

0
回复

Been using one as a weekend detox and honestly the e-ink display is way easier on my eyes than I expected. Love that I can leave the smartphone at home and still get calls and directions without doomscrolling.

0
回复

The dumbphone market is growing for good reason. Constant connectivity is draining for a lot of people. Is this a full hardware product or more of a software layer on existing devices?

0
回复

Love this concept — there's a growing appetite for intentionally "dumb" tech that still respects modern needs. What's the biggest tradeoff you had to make to keep it simple without making it frustrating to actually use day-to-day?

0
回复

The flip form factor is genuinely interesting because it changes the default behavior before software even gets involved. Opening a phone to do something feels more intentional than endlessly waking a slab.

Curious how much of the original Light Phone constraint system carries over here. Is Light Flip still deliberately limited to essentials, or does the tactile nostalgia come with a slightly broader app/tool set than the earlier Light Phones?

0
回复

the whole reason people trust Light Phone is that it's deliberately incapable of most things, not just shaped like an old phone. curious where Flip lands on that spectrum - is this the same "as little as possible" philosophy in a flip shell, or does the nostalgia angle mean you've quietly added back more apps/features than the original Light Phone had, since that's usually how "fun retro form factor" products drift over a few versions.

0
回复
#11
Arkor
Fine-tune and Deploy Open-weight Models in TypeScript
140
一句话介绍:Arkor 让开发者无需管理GPU基础设施或编写Python代码,即可在TypeScript项目中通过代码助手快速微调开源模型,并将其部署为API。
Open Source Developer Tools GitHub Development
模型微调 TypeScript 开发者工具 低代码/无代码ML AI部署 开源模型 GPU管理 代码助手集成 数据安全
用户评论摘要:用户关注训练数据泄漏与评估验证问题(如训练/验证集重复、数据泄露);担忧本地数据集意外提交至Git仓库造成安全风险;询问是否支持下载模型权重、自托管训练与推理;肯定免Python、低摩擦的开发者体验及代码可见性优势。
AI 锐评

Arkor试图解决的是一个经典且昂贵的痛点:将AI功能从“调API”升级到“模型微调”,其跨度往往横跨一个完整的ML基础设施团队。它巧妙地通过“TypeScript原生体验 + 代码助手驱动”降低了Web开发者进入的门槛,其核心价值并非在于训练速度或模型效果,而在于将复杂、非透明的ML过程“软件工程化”。

然而,产品目前暴露的短板非常致命。多数评论指向一个核心矛盾:Arkor的卖点是“无需ML专业知识”,但用户提出的训练数据泄露、过拟合、评估指标等恰恰是ML领域的核心坑。将“检查数据质量”和“判断模型好坏”的责任完全交给一个不了解数据的代码助手,无异于让新手驾驶员在不知路况的情况下盲开。Maker在回复中承认尚未解决这些问题,这暗示产品目前更接近一个“高效的模型训练启动器”,而非一个“可靠的模型生产平台”。

其“Managed-first”的策略虽然在初期降低了使用阻力,但也带来了供应商锁定(不能自托管、不能下载权重)的隐忧,这与开源社区“掌控自己模型”的价值观存在冲突。Arkor真正的挑战在于,它需要从“让训练跑起来”快速进化到“让训练跑对且可控”,并妥善处理数据主权问题。如果只是将ML的复杂性从“Python脚本”转移到“神秘的黑盒操作”,那它并未解决根本问题。目前来看,它更适合极速原型验证,距离真正的生产级工作流尚有一段距离。

查看原始信息
Arkor
Start a real LLM training in 10 minutes. Tell Claude Code or Codex what the model is for; they prepare the datasets and create the training project. Click "Run Training" in the local studio. Arkor runs the training, and deploys the trained model as an OpenAI-compatible API. Think Next.js and Vercel for fine-tuning: code you can review, infrastructure you do not have to manage, and a model your app can call. No GPU setup. No ML expertise required. No Python training code to write.

Hi Product Hunt! 👋

I’m Hina, one of the makers of Arkor.

We built Arkor because we kept running into the same gap: adding an AI feature to an application is easy, but adapting a model to a specific task still feels like joining an ML infrastructure team.

The moment you want to fine-tune, you often end up maintaining a separate Python project, converting datasets into unfamiliar formats, provisioning GPUs, and moving between scripts, notebooks, and dashboards.

We wanted model development to feel more like ordinary software development.

With Arkor, the workflow starts inside your existing repository:

  1. Tell Claude Code or Codex what behavior you want from the model.

  2. Your coding agent can find or prepare a dataset, write conversion scripts, create the TypeScript trainer, and add evaluation.

  3. Review the generated code and changes.

  4. The agent runs pnpm dev, and Arkor Studio opens locally at localhost:4000.

  5. Click Run Training, monitor the loss and checkpoints, test the trained adapter, and deploy the result.

For example, you can give your coding agent a prompt like:

Use Arkor to fine-tune a model that rewrites rough drafts as tweets in my style: https://github.com/arkorlab/arkor
Find or prepare a suitable dataset, create the TypeScript training workflow, and tell me when it is ready to review in Arkor Studio.

Arkor is not intended to be a magical prompt-to-model black box. It is a developer-controlled workflow where coding agents can handle much of the setup, while the training code, data transformations, evaluation, and final decisions remain visible and reviewable.

We’re especially interested in feedback on:

  • Whether the coding-agent workflow feels intuitive

  • Which parts of fine-tuning still feel unclear or intimidating

  • What you would need before using Arkor for a production model

  • Which models, datasets, and deployment workflows we should support next

Thanks for checking out Arkor. We’ll be here throughout the launch and would love to hear what you try building. 🙏

5
回复
The dataset-prep step is where I'd be nervous, since that's exactly the kind of thing a coding agent will happily do and happily get subtly wrong a bad train/val split, near duplicate examples leaking across the two, or silently deduping less than it should. None of that shows up as an error, it just shows up later as a model that looks great in Arkor Studio's loss curve and then underperforms in the real world. Does Arkor Studio do anything to catch that automatically flag overlap between train/val sets, or run a sanity check before burning GPU time on the actual training run? That's the piece I'd want visible before trusting the agent-driven prep step for something going to production.
2
回复

@thys_beesman Hey Brandon, I really appreciate your comments!

You’re right, and this is exactly the kind of failure a healthy-looking loss curve will not reveal.

Arkor Studio does not currently detect train/validation leakage, near-duplicate overlap, or insufficient deduplication automatically. Today, those checks have to be defined in the project’s dataset preparation and evaluation code.

For a framework intended to make fine-tuning accessible to developers who may not already know these failure modes, we do not think that is sufficient. Studio and the framework should provide a preflight layer before GPU time is spent, including split-overlap checks, exact and near-duplicate detection, malformed-example checks, and visibility into eval coverage.

We’re on it.
This is one of the clearest pieces of feedback we’ve received today, and it is helping us prioritize the right responsibility for Studio.

Would you expect these checks to block the run by default, or surface warnings and let the developer decide?

0
回复

the eval/leakage question below is the big one, but there's a second thing nagging me: this workflow lives inside your existing repo, with the coding agent preparing the dataset alongside the code. does that mean the actual training examples end up committed to git history by default? if someone points this at a real support-ticket or user-message dataset to fine-tune on, that's real customer data now living in version control unless you go out of your way to gitignore it - feels like an easy footgun for anyone who isn't already thinking about data handling.

1
回复

@galdayan Hey Gal, thank you for commenting!

Great point. The intended default is not to store raw training data in the repo. Arkor is designed to reference datasets hosted on Hugging Face, but we also support local JSONL files.

You’re right that local files create an easy footgun. Without the right safeguards, real support tickets or user messages could accidentally end up in Git history.

For a framework aimed at developers who may be new to fine-tuning, we should not leave that entirely to the user.
Adding dataset paths to .gitignore by default, warning when local training files are tracked, and surfacing data-handling checks before training are guardrails we should provide.

This is a genuinely helpful point.

Would a default ignore plus a warning for tracked dataset files address the main risk you’re thinking of, or would you expect stricter handling of local datasets?

0
回复

Pushing a fine-tune through without touching a single Python script felt weirdly good. The handoff between Claude Code preparing the data and Arkor just clicking "Run Training" was exactly the kind of low-friction flow I keep wishing existed.

1
回复

@beyzawzzc Hi Beyza!
Thanks so much, that’s exactly the kind of experience we were hoping to create.

I’m curious what you’re building. Are you already using an LLM API inside your product?
If so, what task are you asking the model to handle today, and where does the current model fall short?

0
回复

Looks super cool guys! Curious to know though, how much control do users have when choosing which llm they can fine tune?

1
回复

@lucapiekarski Thanks, Luca!

We’re starting with Gemma 4 for the first release, mainly so we can make the full workflow reliable before expanding the model catalog.

We’re already hearing requests for smaller models for low-latency or local use, plus voice and vision models, so those are definitely areas we’re exploring.

What are you hoping to build, and which model would you want to fine-tune?

0
回复

can you download the resulting weights?

1
回复

@school_4_ants Thanks for asking, Hunter!

Not yet in the current release.

We launched managed-first so people could go from training to inference without setting up GPUs or deployment infrastructure.

That said, avoiding vendor lock-in is important to us, and downloadable training artifacts are on our roadmap.

For your use case, would downloading the LoRA adapter be enough, or would you also need the merged model weights?

And would Hugging Face / safetensors be your preferred format?

0
回复

The 'tell Claude Code what the model is for and it preps the dataset' handoff is the part people usually stall on, so nailing that is smart. When I hit Run Training in the local studio, is the GPU compute actually local to my machine, or does it provision remote GPUs behind that UI — and does the deployed OpenAI-compatible endpoint live self-hosted or on Arkor's side? Also curious how Claude Code/Codex actually drives the dataset prep: an MCP server, a CLI it shells out to, or just file conventions in the project?

0
回复

@noctis06 Great questions.

Today, Arkor Studio is a local control surface, but the actual training runs on Arkor-managed remote GPUs. The deployed OpenAI-compatible endpoint is also hosted by Arkor.

We started fully managed so people could try the entire workflow without provisioning GPUs or setting up serving infrastructure.

That said, we do not want Arkor to become a vendor lock-in. Self-hosting and more portable deployment options are on the roadmap.

Claude Code or Codex works the same way it normally does: you ask it to use Arkor to build a model for a specific task, and it edits the files in your local project, prepares the dataset workflow, writes the training script, and runs the project commands. There is no special MCP-only workflow required.

For your use case, would you prefer bringing your own GPU for training, self-hosting inference, or both?

0
回复

Fine-tuning open-weight models in TypeScript is kind of unusual approach, most tooling for this is Python heavy. Makes it much more accesible for JS devs though. What models are supported so far?

0
回复

@abdurrahman_fakhrul Hi Abdurrahman, thanks for commenting!

You’re right that the underlying training stack still relies heavily on the Python ecosystem.

Arkor wraps that into a TypeScript framework and managed runtime, so JS and TypeScript developers can build the workflow in their own project and go from training to serving without maintaining a separate Python stack or managing GPUs themselves.

For the initial release, we’re intentionally starting with Gemma 4 so we can make the full experience reliable end-to-end.

We’re already hearing requests for smaller models for lower latency and local inference, and expanding model support is on the roadmap.



What would you want to fine-tune, and for what kind of use case? A specific model family or size would be especially helpful as we decide what to support next.

0
回复

curious about the "no ML expertise required" claim specifically - when Claude Code or Codex picks the hyperparameters and dataset splits, how do you know the fine-tune actually improved anything vs just memorized the training set? that eval step usually needs someone who knows what they're looking at, not just infra automation.

0
回复

@omri_ben_shoham1 Hi Omri! That’s a fair challenge.

Studio lets you inspect training and validation loss, but loss alone cannot tell you whether the model generalized or simply memorized the training set.

You’re also right that “no ML expertise required” is too broad.

What Arkor removes today is the need to maintain a Python training stack, provision GPUs, and build the training and serving infrastructure yourself. A meaningful held-out eval set and task-specific success criteria still matter.

One reason we started with the Gemma 4 was to keep the initial use case focused.
A common workflow is to use a larger model to generate or label examples for a clearly defined task, then fine-tune a smaller model to reproduce that behavior. We think that can work well for application-level semantic tasks such as classification, extraction, routing, and rewriting.

We also think evaluation needs to become a first-class part of the Arkor framework. Projects should be able to define task-specific evals in code, while Studio surfaces base-vs-adapter comparisons, data leakage and duplicate checks, and signs of overfitting before deployment. We are not fully there yet, but this is an important direction for us.

Thanks for pushing on the claim. We're going to position this more clearly.

What would be the most useful first step for you: automated dataset checks, base-vs-adapter evaluation on a held-out set, or support for custom task-specific metrics?

0
回复

Arkor makes the fine-tuning workflow feel much closer to normal app development, which is exactly where this has been too painful for most dev teams.

The part I’d want most visibility into is the dataset and eval step. If the coding agent prepares the data and trainer, does Arkor Studio surface checks for train/val leakage, near-duplicate examples, or weak eval coverage before the GPU time gets spent?

0
回复

@aditya_harish_2002 Hey Adithya! I really appreciate your comment.

This is a very fair point.
Studio shows training progress and loss today, but it does not yet take responsibility for detecting train/validation leakage, near-duplicates, or weak eval coverage before a run.

Right now, those checks need to be implemented in the project’s data preparation and evaluation code.
But for a framework designed for developers who may be new to fine-tuning, leaving all of that to the user is not good enough.

We believe dataset validation and evaluation should become first-class parts of both the Arkor framework and Studio.
Preflight checks for split overlap, duplicates, malformed examples, label and scenario coverage, plus held-out base-vs-adapter evaluation, are exactly the direction we’re working toward.

We’re on it, and this feedback helps us prioritize it. Thank you.

What would be the minimum preflight report you’d need before trusting an agent-prepared dataset?

1
回复
#12
AgentManager
Never miss a Claude Code session waiting for your input
128
一句话介绍:AgentManager是一款macOS原生浮动窗口工具,专为并行运行多个Claude Code会话的用户设计,自动悬浮显示需输入的会话,一键跳转终端,解决“哪个会话在等我”的效率痛点。
Mac Developer Tools Artificial Intelligence
macOS工具 Claude Code管理 终端会话监控 开发者效率 原生Swift应用 本地隐私安全 浮动窗口 像素艺术UI
用户评论摘要:用户认可本地钩子检测等设计,主要建议包括:需增加一键取消注册钩子功能;希望支持Codex CLI、Aider等更多AI编码工具;请求手动置顶会话以便监控长任务;有无iOS通知和Linux/Windows支持考虑;远程SSH会话的兼容性问题待解决。
AI 锐评

AgentManager的巧妙之处在于,它用极度克制的设计哲学解决了一个高度具体的痛点——多个AI编码会话的“接听”问题。从技术层面看,基于Claude Code本地钩子的JSON状态写入、FSEvents监听、进程树解析终端窗口ID,构建了一套无需网络、无后台守护进程的纯本地状态管理管道。这避免了屏幕抓取或输出日志嗅探的脆弱性,确保了“需要输入”状态的语义准确性。其“毛玻璃浮动窗口仅在需要时显现”的交互范式,以及对通知中心的彻底弃用(代之以系统级浮窗+猫叫声),直击了多会话工作流的认知摩擦核心:不是信息不足,而是信号过载。从用户评论的深度讨论可见,开发者对共享配置文件合并写入的原子性处理、进程PID复用防欺诈、远程SSH会话的严格边界声明,都体现了一种“基础设施思维”——功能不溢出其应有的职权范围。

但这款工具的壁垒同样明显:它完全绑定于macOS原生生态(SwiftUI+AppKit的窗口层、Automation API的跳转实现),这意味着其精致体验无法跨平台。更根本的局限在于,它的价值直接受限于Claude Code的流行度——若用户迁移至其他AI工具(如Cursor、Codex CLI),其状态检测机制便需从头适配,而开发者已明确表示“不猜状态”。此外,其7天免费试用后4.99美元/月的定价,在纯本地工具中偏高,除非它未来能建立跨代理生态的“标准状态层”——成为所有AI编码终端的中立监控面板。目前,它是一款服务于重度Claude Code用户的精致效率配件,而非一个更宏大的AI开发平台基础设施。

查看原始信息
AgentManager
A floating macOS window for every Claude Code session. The moment one needs your input, it surfaces — and hides when all is clear. Jump to the exact terminal in one click. Native Swift, notarized, 7-day free trial.
Hi Product Hunt! 👋 I'm a solo developer from Japan, and like many of you I run several Claude Code sessions in parallel. I kept doing the same dumb thing: kick off a session, switch to something else, and realize ten minutes later that Claude had been sitting there waiting for a single "yes" the whole time. AgentManager is a native macOS app that fixes exactly that: 🐈 A floating window shows every session's state — running, waiting, done ⚡ It surfaces the moment a session needs your input, and hides when all is clear 🖱️ Click a row to jump straight to the right terminal (iTerm2 pane, Terminal tab, VS Code / Cursor window…) 🔔 Cat-meow alerts instead of Notification Center banners (mutable, promise) 🎨 Sessions live as pixel-art cats in a little room — day and night change with your clock. There's a Simple mode if cats aren't your thing. A few things that matter to me: • Native Swift, ~2.5 MB download, notarized by Apple • Your session data (prompts, paths, names) never leaves your Mac • 7-day free trial with no credit card. Then $4.99/mo or $48/yr. I'd love your feedback — especially which terminals or IDEs you'd like better jump support for. I'll be here all day answering questions!
3
回复

@umechanhika  The quiet/no-banner design is a nice choice for multi-agent work. One edge case I would want to test is remote Claude Code inside tmux: if the state file syncs back but the jump target cannot map to the right pane, does AgentManager still show it as needs attention without trying to focus the wrong window?

0
回复

this is exactly the problem with running a few Claude Code sessions in parallel - you either tab through terminals constantly to check status or you get pulled back way too late because you forgot which one was still going. a floating window that only shows up when something actually needs you is the right shape for this, way better than a generic desktop notification that gets lost in the pile. does it distinguish between "needs input" and "just finished, fyi" or is it one signal for both right now?

0
回复

@omri_ben_shoham1 

You've described my pre-AgentManager life exactly 😄 And yes — they're fully distinct states, styled by urgency. "Needs input" is amber: a pulsing glow, a stripe on the row's edge, and the one short meow. "Just finished, fyi" is green and completely silent — the row turns calm, the window surfaces so you notice, but nothing chirps at you. Blue means "still working, don't touch."

One subtle case I made sure of: when a session finishes and then just sits there, Claude Code emits an idle notification — that stays green. A finished session never escalates itself into fake "needs input" just because time passed. The amber signal only ever means a real blocker: a permission prompt, a plan approval, or a question waiting on you.

0
回复

hooks in ~/.claude/settings.json is the robust call — the seam is that it's one shared file. the other claude code watchers on this page register there too, and an uninstall leaves a stale path still firing on every event.

0
回复

@qifengzheng 

You've found the real seam, and I won't pretend otherwise. What AgentManager does about the shared-file half: registration is a merge, not a write — it checks whether its command already exists per event, appends alongside other tools' hooks without touching them, and the edit is a text splice that only modifies the top-level hooks key, so your formatting, comments, and symlinks survive. A .bak is written first, and if the JSON doesn't parse, it refuses to write at all rather than "fix" a file other tools depend on.

One deliberate choice against dangling: the hook script and binary are installed under ~/.claude/agent-manager/, not inside the app bundle — so moving or deleting the .app doesn't leave settings.json pointing into a Trash'd bundle, and the path is written as $HOME/... so it survives machine migrations.

But the uninstall half of your point stands: today there's no one-click unregister, so fully removing AgentManager means cleaning its entries out of settings.json yourself. That's exactly the stale-path scenario you're describing and it deserves a proper "remove hooks" button that does the surgical reverse of registration. It's going on the list — this is the most technically useful comment on the page, thank you.

0
回复

Mostly Claude Code as the daily driver, with Codex CLI as the second seat and Aider for quick one-off edits. Codex is the one I d want next since it has a hooks/notify config, so a small state-writer there would slot straight into your FSEvents watcher without you inventing anything new. Aider s the weak link: no real lifecycle events, so it d be exactly the guess-from-output case you re right to refuse.

0
回复

@noctis06 

This is genuinely useful — you've basically written the integration sketch for me. Codex CLI's notify config is exactly the kind of seam I'd want: a small state-writer emitting the same JSON into the watched directory, and the app wouldn't even know the difference. The homework on my side is mapping its events onto the four-state model without cheating — the value here isn't "a session exists," it's the reliable distinction between blocked on you, working, and done, plus resolving the jump target the same way. If Codex's events carry enough signal for that, it slots in; if they don't, I'd rather wait than ship a lamp that lies.

And agreed on Aider — no lifecycle events means it stays out on principle, not out of laziness. Thanks for taking the time on this thread; Codex just moved to the top of the "next agent" shortlist.

0
回复

Love this idea, especially the jump-to-terminal one click since that always trips me up. One thing that would seal the deal for me is per project status icons in the floating window itself, like a little red dot or spinner right next to a project name so I can tell at a glance which session needs me without scanning titles. Would also be great if those icons respected Do Not Disturb or hid when a session is just running a long thinking step versus truly waiting on me.

0
回复

@dndcltj 

On the "which session needs me" part — the state lamps are already there: every row has a colored light next to the project name, and the waiting one is a pulsing amber glow with a stripe down the row's edge, so it catches your eye before you've read anything. Blue with a slow breathing animation means "working, leave it alone." And that thinking-vs-waiting distinction is structural: state comes from Claude Code's own hook events, so a long thinking step or tool call stays blue — amber only fires on explicit signals like a permission prompt or a question to you. No false alarms mid-stream.

But if what you're after is telling which project is which without reading names — like a distinct icon or color per project — that doesn't exist yet, and it's a genuinely good idea. Right now identity is the project folder name plus stable row order (rows never reorder, so muscle memory does a lot). A per-project visual mark would make that instant. Adding it to the list — and tell me which you actually meant, I'm curious now 😄

DND is a good call too: there are no Notification Center banners at all and the meow is mutable, but auto-quieting during a macOS Focus session would be the polished version. Noted!

0
回复

Finally something that solves the "which Claude session is stuck" problem. The floating window just appearing when input is needed is exactly the kind of low-key native touch I want from a Mac utility.

0
回复

@ozandenizm84853 

Thank you! "Low-key" is the highest compliment for this app — the bar I set was that a good Mac utility should feel like part of the OS: show up exactly when needed, disappear the moment it's not, and never ask for attention it hasn't earned. That's also why there are no Notification Center banners — the window appearing quietly is the notification.

Hope it keeps your stuck sessions from hiding on you 🐈

0
回复

The native Swift choice really shows here — having a floating window that just appears and disappears instead of fighting with another Electron app is such a relief.

0
回复

@ege_karama27779 

Thank you — that's exactly why it had to be native. An app whose job is to appear and vanish dozens of times a day, while living permanently next to your real work, has no business shipping a browser runtime with it. Swift + AppKit keeps the panel instant and the footprint tiny, so it earns its spot as a resident instead of competing with your actual tools for resources.

(Also, pixel-art cats deserve to be rendered natively. It's a matter of principle 🐈)

0
回复

Does it track multiple Claude Code sessions at once, or is it one window per session? Wondering if this gets noisy if you're juggling several agents.

0
回复

@talhakhalidmtk 

One window, all sessions — each gets a row with its own state light, and the list grows as sessions come and go. No per-session windows to manage.

Noise was actually the main design constraint, because the whole point is that you're juggling several agents and the tool shouldn't add to it. So: the window hides itself whenever nothing needs you and only surfaces when a session is waiting or finished. There are no Notification Center banners at all. And the sound alert is a single short meow with a global debounce — if three sessions hit "waiting" at once, you hear one meow, not three. The more agents you run, the more that restraint matters.

0
回复

this looks super useful for keeping tabs on multiple sessions, honestly the auto-surfacing is a nice touch. one thing though, it would be great if you could pin certain sessions to stay visible even when nothing needs input, so you can monitor like a long-running build or deploy without it disappearing on you

0
回复

@sebahat45760404 

Thanks! Good news for the build/deploy case specifically: you may not need to watch it at all. The moment a long-running session finishes, it flips to "done" and the window surfaces on its own — completion is treated as an event worth your attention, same as waiting for input. And while it's running, the menu bar keeps a live per-state count, so a glance tells you it's still going without any window open. ⌥Space summons the full list any time you want a closer look.

That said, "keep it visible, I just want to watch" is a fair and different ask — and you're actually the second person requesting a pin option today, which officially moves it up the list 😄 Auto-hide is the default because staying quiet when nothing needs you is the app's whole personality, but an opt-in pin wouldn't fight that. Thanks for the concrete use case!

0
回复

Claude Code sessions that need input mid-task is real pain, especially when you're juggling multiple things. Does it support notifications on iOS or just Mac desktop alerts?

0
回复

@abdurrahman_fakhrul 

Mac only, and honestly it's likely to stay that way for a while. The app is built around a strict "nothing leaves your Mac" rule — no server, no network, session state is just local files. Pushing notifications to an iPhone would mean routing your session activity through some backend, and that's a privacy tradeoff I don't want to make casually.

It's also a slightly different problem: AgentManager is built for the at-your-desk case — you're juggling sessions and other work on the same machine, and it surfaces the one that needs you. If you're away from the Mac entirely, a blocked session usually isn't something you can act on anyway (the terminal's back at your desk!). That said, if "know before I walk back" is a real need for you, I'd love to hear the scenario — it'd have to be designed carefully, but I'm listening.

0
回复

The local hooks approach is the right tradeoff here. Once an agent has access to repos, terminals, and approval prompts, I want the session monitor to be boring infrastructure: local state, no output scraping, no network dependency, and stable session identity so I can trust the jump target.

0
回复

@krekeltronics 

"Boring infrastructure" is the nicest thing you could call it — that was the design brief, verbatim. A tool that sits between you and approval prompts has no business being clever.

On session identity, since you brought it up: each session is keyed by Claude Code's own session ID, and the hook records the owning claude process's PID plus its start time — so if a PID gets recycled by the OS, a dead session can't masquerade as alive. Sessions whose process vanished without a clean SessionEnd get swept out automatically. And the jump target is captured at hook time, not guessed at click time — e.g. the exact iTerm2 pane GUID is resolved by walking the process tree when the event fires, so the click lands where the session actually lives.

State files are written atomically to a local directory, the app just watches it with FSEvents, and there's no daemon and no network in the loop. If it ever stops being boring, that's a regression 😄

0
回复

the hooks-based detection is the right call, way more robust than scraping terminal output. one case I didn't see covered: what happens with a Claude Code session running on a remote box over SSH rather than locally on the Mac - does the hook binary need to be installed remote-side too, or is this strictly for local sessions right now? I run a mix of local and remote-server sessions and that's usually where these menu-bar tools stop working for me.

0
回复

@galdayan 

Honest answer: strictly local right now. The hook fires on the machine where Claude Code runs, so in an SSH session it would run remote-side and write its state file to the remote filesystem — which the Mac app never sees, since it just watches a local directory. So installing the binary remotely wouldn't help today; you'd get state files stranded on the server.

Could it work? Partially, in principle. Detection is just JSON files in a directory, so a synced/mounted ~/.claude/agent-manager/sessions/ could carry remote state to the Mac — that half is plumbing. The hard half is the jump: locally the hook walks the process tree to find the exact terminal pane, and that chain breaks at the SSH boundary. Mapping a remote session back to the local terminal that's SSH'd into it needs a different mechanism, and I'd rather not ship a jump that lands on the wrong window.

You're right that this is where menu-bar tools usually stop, and a local+remote mix is a real workflow — it's on my radar. Out of curiosity: are your remote sessions inside tmux/screen on the server, or plain SSH in a local terminal tab? That changes which approach would actually be reliable.

0
回复

That's clever. Any plan to support Linux or Windows once Claude Code agents run there?

0
回复

@dhiraj_patel5 

Thanks! Honest answer: no near-term plans. AgentManager is native Swift (SwiftUI + AppKit) on purpose, and the parts that make it feel good are exactly the parts that don't port — the always-on-top panel, FSEvents watching, and the terminal-jump layer built on macOS Automation/Accessibility APIs. A Linux or Windows version wouldn't be a port, it'd be a second app written from scratch.

The one honest caveat: the concept is portable. Detection is just Claude Code hooks writing local JSON, nothing Mac-specific about that half. But as a solo dev I'd rather make one platform genuinely excellent than three platforms mediocre — so for now, Mac gets all of it. If demand elsewhere gets loud enough, that math can change 😄

0
回复

Honestly this looks super useful since I constantly lose track of which Claude session is actually waiting on me. One thing though, would be great if it could play a subtle sound or send a native macOS notification when a session needs input, kind of like how Messages does it. Sometimes the floating window might be hidden behind another app and I would totally miss it otherwise.

0
回复

@araczeliha97691 

Good news — the sound already exists, and it's the app's signature move: a subtle cat meow (~0.35s) the moment a session starts waiting 🐱 It's mutable if meows aren't your office vibe, and if several sessions hit "waiting" at once it's debounced so you get one meow, not a chorus.

On the buried-window worry: the floating window is always-on-top, and it re-surfaces in front the moment any session needs you — so it can't actually get lost behind other apps. Between the meow, the auto-surface, and the live counts in the menu bar, I deliberately skipped Notification Center: banners pile up, get swiped away, and end up as one more inbox to ignore. The window appearing is the notification, and it disappears once you've handled things.

Give the trial a spin and tell me if a waiting session ever actually slips past you — if it does, that's a bug in my book and I want to hear about it.

0
回复

Finally something that solves the constant tab-switching annoyance with Claude Code. The auto-surfacing when input is needed feels genuinely native, not bolted on.

0
回复

@erengilsed11978 

Thank you! "Not bolted on" is exactly what I was going for — the app only speaks up when a session actually needs you, and gets out of the way the rest of the time.

And "native" goes all the way down: it's Swift (SwiftUI + AppKit), the whole download is ~2.5 MB, and everything runs locally on your Mac. The tab-switching tax was the whole reason I built it, so it's great to hear it lands.

0
回复

Missing a Claude Code session waiting for input is such a specific but real annoyance once you're running multiple sessions. Does it notify across all active sessions simultaneously, or do you need to check in on each one individually?

0
回复

@ark_y_k 

Exactly the annoyance that made me build it 😄 And no individual checking needed — every session registers automatically and they all live in one floating window, each with its own state light. The moment any of them starts waiting, the window surfaces with that row pulsing amber; when nothing needs you, it hides itself. The menu bar shows a live count too, so even with the window hidden you can see "3 running, 1 waiting" at a glance.

Small detail I'm fond of: if several sessions hit "waiting" at nearly the same time, the sound alert has a global debounce so you get one meow, not a chorus. From the window it's one click to jump to the exact terminal (iTerm2 pane, VS Code window, etc.) that needs the answer.

0
回复

A click-to-jump shortcut is great, but it would be even better if I could pin certain sessions to stay visible even when idle, since I often switch between two long-running tasks and lose track of which window belongs to which project.

0
回复

@senargf3 

Good news on half of this: sessions never leave the list while they're alive — idle ones stay right there as grey rows, and every row is labeled with its project folder name. So "which window belongs to which project" is exactly what the list answers: glance, click, and you're in that project's terminal. The window itself is also one ⌥Space away at any moment, even when it's auto-hidden.

The other half — keeping it visible through idle instead of auto-hiding — is a fair ask. Auto-hide is the default because "quiet when nothing needs you" is the app's whole personality, but a pin option for people juggling two long-running tasks wouldn't fight that. Adding it to the list — thanks for the concrete use case, that's the kind that actually shapes the roadmap!

0
回复

Congratulations on the launch! One question: does it also notify when a session finishes cleanly, or only when one is blocked waiting for input? I run long pipeline jobs, and knowing when they are done matters as much as knowing when they are stuck.

0
回复

@alieksia 

Thank you! 🙌 And yes — a clean finish is a first-class event, not just blocked-waiting. When a session finishes, its row turns green and the floating window surfaces the same way it does for waiting; it only hides when there's nothing left that needs your attention. The menu bar also shows a live count per state, so you can see "2 running, 1 done" at a glance.

One detail that matters for long pipeline jobs: "done" isn't premature. If the main response stops but background subagents are still working, the session stays "processing" and only flips to done when the last one actually finishes.

The only thing reserved for waiting is the sound alert — a finished session surfaces silently, so a long run ending doesn't interrupt whatever you switched to. If you'd want an optional sound for completions too, that's an easy toggle to consider — tell me and I'll put it on the list!

1
回复

Really I'm here for the cat sounds. Though I'd want per-session pitch, so I know which one is meowing.

0
回复

@aidan_codefox 

Finally, someone here for the right reasons 😄

And honestly… that's not even a joke to me. The meow is procedurally synthesized (no audio files), so giving each session its own pitch is genuinely doable — your refactor session as the deep old tomcat, the test runner as the anxious kitten. Adding it to the list, and I can't believe this is now a roadmap item. Thank you.

0
回复

'hides when all is clear' is the half most tools skip, they just add alerts. with several sessions waiting at once, does it rank which to surface first?

0
回复

@andrewzakonov 

Thanks — "hides when all is clear" was actually the starting point of the whole app. The alert part is easy; the quiet part is what makes it livable.

Honest answer: no ranking. All sessions live in one list in a stable order (by launch time), and rows never reorder on state changes — waiting ones just light up with a pulsing amber stripe. That's deliberate: with a handful of parallel sessions, reordering would break your spatial memory of "top row = the refactor, second = the tests", and you'd misclick jumps. A stable list you can scan in a second beat every ranking scheme I tried.

In practice you glance at the amber rows, click the one you care about, and you're in the right terminal. If you'd find explicit prioritization (e.g. "waiting longest first") genuinely useful though, I'm open to it — that's exactly the kind of feedback I'm here for.

0
回复

The 'waiting for input' detection is the whole value here — auto-surfacing and jumping to the exact iTerm2/Cursor pane is what I'd actually use daily. How does it read session state under the hood: scraping terminal output / the TTY, or hooking the Claude Code process directly? Mostly wondering how it avoids false 'waiting' signals when a session is just mid-stream on a long tool call.

0
回复

@noctis06 

Great question! It's neither — no TTY scraping, no process hooking. AgentManager uses Claude Code's official hooks: it registers them in ~/.claude/settings.json, Claude Code invokes a tiny binary on every lifecycle event (prompt submit, tool use, notifications, stop…), that writes per-session state to a local JSON file, and the app watches it with FSEvents. No polling, no daemon, no network — everything stays on your Mac.

False "waiting" can't really happen because state is never inferred from output or silence. "Waiting" only comes from explicit signals — a permission prompt, a plan-approval dialog, or an AskUserQuestion call. A long tool call just fires PreToolUse and stays "processing" until it finishes, no matter how long it streams. Trickiest edge case was background subagents that outlive the main response — those are tracked too, so sessions don't flip to "done" early.

Happy to go deeper if you're curious!

0
回复

This is exactly the annoying Claude Code failure mode. The floating session state is useful, but the one-click jump back to the right terminal is the part that makes it feel workflow-native instead of just another alert layer.

Also love the cat room, as long as Simple mode exists for when I need to pretend I am serious.

0
回复

@aditya_harish_2002 

Thank you! And you nailed the design bet — alerts alone just move the problem ("now I have unread alerts"). The jump is what closes the loop, so I spent a silly amount of time on precision: iTerm2 lands on the exact pane, Terminal and Ghostty on the tab, VS Code / Cursor / Zed on the right project window.

And yes, Simple mode exists precisely for screen-sharing-with-your-manager moments 😄 The cats will wait for you.

1
回复
#13
Grindoro
Focus with others through synchronized Pomodoro sessions
124
一句话介绍:Grindoro 是一款通过同步番茄钟工作时段,让用户与全球陌生人共同专注、消除独自工作孤独感的生产力工具,核心解决“自律难坚持、专注感孤立”的痛点。
Productivity Task Management Social Networking
社交番茄钟 同步专注 体双效应 习惯养成 游戏化 咖啡能量 每周咖啡馆 生产力工具 iOS应用 社区行动
用户评论摘要:用户普遍认可“同步专注”解决孤独感的设计,并关注其差异化细节。部分用户关心:迟到能否直接加入当前同步时段(答:可立即加入每15/30分开始的轮次);咖啡能量是否为硬性限制(答:是真实预算,但可通过完成工作解锁无限模式);缺席惩罚(答:每日咖啡刷新,每周只需登录5天收集咖啡馆,有补签盾)。建议增加专注计划与实际完成度对比的复盘功能。
AI 锐评

Grindoro 聪明地将“体双效应”(body doubling)这一被研究验证的认知策略,产品化为“全球同步番茄钟”,避开了传统番茄钟App“孤独计时”的叙事陷阱。其核心价值不在于计时,而在于“共同在场”的感觉——用户并非与陌生人社交,而是共享一个起止节奏,这恰好利用了他人在场带来的微弱社会压力,却又通过隐藏活动流、禁止聊天等方式消除了多数社交App带来的分心。这种“轻社交、重存在”的定位,比“专注森林”等游戏化工具更直接地触达了自律的核心矛盾:意志力薄弱往往源于缺乏外部锚点。

然而,产品的可持续性面临挑战:全球同步依赖足够大的同时在线用户基数,若用户量不足,即时加入的“共同感”将大幅削弱。此外,“咖啡能量”作为硬性预算的设定存在风险——它被包装成“让每次专注更刻意”,但实质上是一种额外的认知开销:用户需管理可用能量水平,并在能量用尽时被迫停止,这与“想专注就专注”的核心需求存在内在冲突。尽管可以通过“磨豆机”奖励解锁无限模式,但这使游戏化系统变得复杂,可能让用户为了“解锁”而机械完成任务,偏离专注本身的目标。评论中关于缺席压力、被动社交的需求,提示开发者需持续平衡“轻盈的社会责任”与“令人上瘾的游戏化”。长期看,Grindoro 能否从“小众新奇工具”进化为“日常习惯基础设施”,取决于它能否在用户基数增长前保持“任何时候加入都不孤单”的核心体验。

查看原始信息
Grindoro
Grindoro is a Pomodoro focus app that makes productivity feel less lonely. Join synchronized focus sessions, turn habits and goals into tasks, manage your daily coffee energy, and collect weekly cafés as proof of your consistency.
Hey Product Hunt 👋 I built Grindoro because I was tired of productivity apps that feel lonely. Most Pomodoro timers help you track time. Grindoro is designed to make focus feel like a shared ritual. Every session starts in a global rhythm — at the top of the hour or the 30-minute mark — so you’re not just starting a timer, you’re joining other people who are also trying to focus right now. Inside Grindoro, you can: – Start synchronized focus sessions – Use Light, Medium, or Dark Roast focus modes – Build routines, single tasks, and bigger goals – Manage your daily coffee energy – Earn XP and boosters – Collect weekly cafés as part of your focus passport The idea is simple: discipline feels easier when it feels alive. I’d love your feedback on the concept, the onboarding, and the weekly café mechanic. Thanks for checking it out ☕
2
回复

body doubling for focus work actually has real research behind it so this isn't just a gimmick layered on top of a timer app. the coffee/cafe collection as a consistency proof is a nice touch too, gives you something to look back on besides a streak number that resets your motivation to zero the day you miss it. do the synchronized sessions show who else is in them or is it fully anonymous, just a shared clock ticking?

1
回复

@omri_ben_shoham1 Thanks, I really appreciate that!
Especially your point about consistency needing something more meaningful than a streak number.

The sessions aren’t fully anonymous. Before joining, you can see a live feed of who’s there and what they’re working on. Tasks are public by default, but users can turn that off in Settings.

Once you join the session, the feed disappears so it doesn’t become another distraction. You still know others are focusing with you, but the interface gets out of the way.

0
回复

The synchronized-on-the-hour hook is the interesting part — that s the difference between a timer and a room. But it cuts both ways for a solo user: if I sit down at 2:14 and the next global session starts at 2:30, am I waiting 16 minutes, or can I start solo right away and get folded into the next synced block when it comes around? And is the daily coffee energy an actual cap on how many sessions I can run, or purely cosmetic XP?

1
回复

@leo404 At 2:14, you wouldn’t need to wait until 2:30 — the next Light Roast 15/5 session starts at 2:15, so you could jump in almost immediately.

You can also start a Solo session or a free-form Timer at any time. Those stay separate rather than automatically folding into the next global block, so the choice is between joining the shared rhythm or starting immediately.

Coffee is a real session budget, not just cosmetic XP. I designed it to make each focus session feel intentional, rather than something you start and abandon without thinking.

At the same time, every completed session charges the Grind Machine, which can unlock unlimited coffee for a day or longer as a reward. That creates a kind of release valve: periods of intentional limits, followed by a more relaxed stretch where you can focus as much as you want.

0
回复

I wanted to share a little more about two community mechanics in Grindoro - Shared Coffee and Live Momentum.

Shared Coffee - coffee is the energy used to start focus sessions. If someone is running low, another person can send them a cup while focusing. Sometimes you support someone else, and sometimes a coffee arrives exactly when you need it.

Live Momentum - the more people focusing at the same time, the stronger the shared XP boost becomes. Your work remains personal, but showing up together creates extra momentum for everyone.

I built both because I wanted the social layer to be useful without becoming distracting. There’s no need to choose a group, introduce yourself, chat, or manage another feed. You simply focus - and your presence can quietly help someone else keep going.

1
回复

Love that you turned collecting weekly cafes into proof of consistency, it reframes focus as something social and worth savoring rather than a lonely grind.

1
回复

@ilko_kacharov Thanks, Ilko! Really glad that resonated!

0
回复

Synchronized pomodoro with others is simple but actually effective for accountability. I tend to work better knowing someone else is in the session. Is there a way to see what people are working on or just the timer?

1
回复

@abdurrahman_fakhrul Thank you! You can see more than just the timer. There’s a lightweight activity feed with other people’s avatars and what they’re currently working on, while the session itself shows how many people are focusing alongside you.

I’ve intentionally kept it minimal, no chat or busy social feature, so it adds presence and accountability without becoming another distraction.

0
回复

the coffee-energy and weekly-cafe mechanic is cute, but I'd want to know what happens on a day you just don't show up - does the streak/energy reset hard, or does it forgive a skip? Sergey's comment below about not wanting this to become a streak-pressure machine is the right worry, since a lot of these habit apps end up optimizing for not-breaking-the-chain instead of actual focused work getting done.

1
回复

@galdayan That’s a fair concern - I don’t want Grindoro to punish people for missing a day.

Coffee refills daily, so skipping a day doesn’t create a hard energy penalty. For the weekly café, simply opening the app counts as a visit, and you only need 5 visits per week - so you can miss 2 days and still collect it.

There is a daily streak, but missed days can be covered by streak shields. You collect shield fragments after focus sessions and combine them into shields, so one busy day doesn’t automatically erase your progress.

The goal is to reward returning and doing meaningful focus work - not create anxiety around maintaining a perfect chain.

0
回复

That synchronized focus sessions idea is genuinely clever, makes long pomodoro stretches feel way more motivating when you know others are grinding alongside you.

1
回复

@mertrstemobygx 

Thank you! That sense of quietly grinding alongside others is exactly what I wanted to create.

There’s also a small community-support mechanic around coffee. During a session, you can share a coffee with someone who’s running low, and sometimes another person may send one to you. It’s a lightweight way to help each other keep going without adding chat or turning the app into a social network

0
回复

This is neat. Does the group sync adjust if someone joins late to a session?

1
回复

@dhiraj_patel5 Yes! If you join late, you enter the session at its current point, so you sync to the shared timer instead of restarting it for yourself.

You can see that in the Live Activity too: the faded part on the left shows the portion of the session that had already passed before you joined, while the highlighted part shows your active segment. The same idea also carries into your stats, so it’s clear when you joined and how much of the shared session you actually completed.

0
回复

Synchronized focus sessions add a nice social accountability layer that solo Pomodoro apps usually lack. How does matching work — do you get grouped with strangers working on similar goals, or is it more for pre-formed groups/friends?

1
回复

@ark_y_k Right now, there’s no goal-based matching or private friend groups. Everyone joins the same global session at the scheduled start time - usually on the hour or half-hour - alongside whoever else is focusing then.

I also don’t want users spending time deciding which room or group to join. Their main task is to focus, so the social layer should require almost no setup.

You can see the shared participant count and a lightweight view of what some people are working on, but everyone stays in the same shared rhythm. The idea is accountability without adding more decisions or distractions.

0
回复

The synced focus sessions are such a thoughtful touch, turning solo Pomodoros into something that actually feels shared. Love that the café collection ties consistency to a tangible reward instead of just another streak counter.

1
回复

@hava3cfd Thank you! That’s exactly the feeling I was aiming for: enough shared presence to make focusing feel less lonely, without adding the distractions of a social network.

And yes, the cafés are meant to turn consistency into something you can actually collect and look back on, together with your focus stats from each week.

0
回复

Is it something like a tracker where you can see other peers or?

1
回复

@busmark_w_nika It’s more like a shared focus room than a peer tracker. You track your own tasks, sessions, and progress, while Global mode shows how many other people are focusing at the same time.

There’s no chat or social feed to distract you, the idea is simply to make a solo session feel like you’re working alongside others.

0
回复

The synchronized-session angle is smart because most Pomodoro apps optimize solo discipline, not the show-up-with-others part. I would be curious to see a lightweight post-session recap that separates planned focus time from completed rounds, so the social energy helps accountability without turning the app into another streak-pressure machine.

1
回复

@sergbmw That’s a really good distinction. Right now, the recap focuses mostly on what was completed, but showing planned focus time alongside completed rounds could make it feel more honest and useful.

I especially like that it adds accountability without punishing people for breaking a streak. I’m adding this idea to the post-session recap list.

0
回复

Grindoro makes the social Pomodoro idea feel more concrete than just “study with friends.” Starting everyone on the hour or half-hour is a nice constraint because it gives the session a shared rhythm instead of becoming another solo timer.

The weekly cafés mechanic is the part I’d want to try. Are cafés mostly collectible proof of consistency, or do they change how people match into sessions and build routines over time?

1
回复

@aditya_harish_2002 Thanks! That shared rhythm is exactly what I wanted the fixed start times to create.

Right now, cafés are mostly collectible proof of consistency. Each week has a new café, and completing at least one focus session during that week adds it to your seasonal passport.

Each café also keeps a snapshot of that week: focused time, sessions, visits, and cups used. You can swipe through the café carousel later and look back at both the places you collected and your progress during each week.

They don’t currently affect session matching. Everyone still joins the same shared focus rhythm, while cafés add a longer-term reason to return each week. I’m curious whether cafés shaping smaller groups or shared routines would make the mechanic more interesting to you.

1
回复
#14
Redential
A developer credential that proves what you built, NDA safe.
124
一句话介绍:Redential通过本地扫描Git历史生成可验证的开发者履历,并支持实时答辩,解决NDA下工作经验无法可信证明的招聘痛点。
Open Source Developer Tools Career
开发者可信履历 Git历史分析 NDA安全 技术招聘 技能验证 开源CLI 实时答辩 防伪凭证 隐私保护 人才筛选
用户评论摘要:用户关注防伪机制(如反对手选提交)、私人仓库合规性、企业IT限制、挤压/变基历史的影响、作者邮箱错误归属、以及LLM提问深度不足可能被糊弄的风险。建议增加按角色展示最相关工作的演示功能,并强调纯本地扫描对合规场景的天然优势。
AI 锐评

Redential精准切中了技术招聘中“简历可信度崩坏”与“NDA下隐形贡献无法量化”的双重死穴。其核心价值不在于扫描Git历史——这仅是基础数据层——而在于“本地扫描+实时答辩”构建的信任闭环:本地扫描降低伪造门槛(但无法根除),答辩则用认知拷问筛选出真正理解代码的人。这种设计聪明地绕开了代码泄露风险,却带来了新的博弈——造假者只需伪造历史并背熟剧本,而资深开发者因记忆衰减反而可能在答辩中吃亏(用户已精准指出)。产品目前停留在“提高造假成本”而非“杜绝造假”的务实阶段,但若提问引擎仅依赖元数据生成泛泛架构问题,其效度将迅速被市场驯化。更致命的隐患是:它排斥“手选提交”这种反造假设计,反而可能让拥有大量个人项目的Junior,比在单一企业深耕NDA项目的Senior获得更光鲜的Profile。此外,团队侧API的“无审核权”设计,本质将责任全部外包给面试官,对用人单位而言,这依然是一个“辅助判断工具”,而非“背调替代品”。作为招聘链条中的中间件,Redential的商业化能否跑通,取决于它能否在“防伪承诺”和“用户体验”之间找到不牺牲可信度的平衡点,以及是否愿意接入更多如代码审计日志、AI辅助提问可信度评分等抗博弈层。

查看原始信息
Redential
Your best work is probably under an NDA, so nobody can see it. Nobody trusts CVs anymore, they are so fakeable. Redential reads your git history on your own machine and turns it into a profile of what you actually built. Never your code, and you see exactly what gets shared before it does. Then you defend it live, answering questions about your own work. The result: a credential recruiters can trust more than a CV. The CLI is free and open source.

Hey Product Hunt 👋

I'm Juan, founder of Redential.

Here's the problem that made me build this: Nowadays a CV can be faked and polished, a github can be cloned, none of it shows anyone that you really know how to do the job.

So the question became: what can't be faked? The answer: work you actually did and the memory of doing it.

Here's how Redential works:

🔍 Scan your work, locally (Open Source CLI).
Run 'npx redential scan' on any repo, no account needed. It reads your git history on your own machine and detects what you actually built with: not just "Stripe is installed", but real patterns, like a webhook flow with signature checks and duplicate-payment protection, connected.

You can also connect the github app to your repos if you don't choose the CLI path.

🔒 Your code never leaves your machine
Only a small summary does: languages, activity, skills, and you see the exact data before anything uploads. This isn't a promise: the privacy rules are tests inside the open source repo, and you can run them yourself.

🪪 Get a public profile
Your work becomes a credential you can share with recruiters: what you built, with what, for how long. Even the work nobody was allowed to see.

🎤 Defend it live
This is the part that's hard to fake. You answer questions about your own work, in real time, generated from your own history. If you did the work, you remember it. If you copied a history, you don't.

⚖️ Honest by design
Anything from your own machine can be faked — so we label that evidence as exactly what it is, and never call it verified until it's defended. No inflated claims, ever.

Getting started is free: scan your repos, build your public profile, and run your first live defense. For teams that hire, there's a dashboard and API to request and review defended profiles. Want it for your team? Book a call: https://cal.com/juan-redential/3...

The CLI is fully open source: https://github.com/redential/redential-cli - if you're technical and you can break the trust model, that's the feedback I want most. I want more contributors to help me improve the evidence and make a stronger credential.

👉 Try it now: npx redential scan or redential.com

Thanks for reading, excited to hear what you think! 🙌

Juan

2
回复

Congrats on the launch!!

Verification is the interesting part here. How do you prove someone actually built what they claim without access to private repos?

1
回复

The proof happens on the dev's own machine: the open-source CLI reads the git history right where it lives and turns it into metadata (skills, activity, time spans, never code). Then they defend that work live, answering questions about their own decisions. So the claim gets tested without the repo ever going anywhere.

Btw if you ever feel like helping us harden the evidence, the CLI is open source, happy to have you there: github.com/Redential/redential-cli

0
回复

This is a really clever approach to the trust problem in hiring. One thing that would make it even more useful would be letting users pick which repos or specific commits get included, since someone might have a dozen personal projects but want to lead with their most relevant work for a specific role.

1
回复

@kaplaneylu22927 Hey eylul! So today it's repo by repo: nothing enters your profile unless you scan and submit that repo, so you already control what's in.

We stopped there on purpose though: the moment evidence can be hand-picked commit by commit, it stops being evidence and turns back into a curated portfolio, which is kinda the problem we're replacing.

Curious about your use case though: is it more about hiding noise, or highlighting the most relevant work for a role? Because the second one is a presentation feature we could build without touching the evidence, and tbh it's a good idea.

0
回复

This hits close to home — I'm self-taught and just launched my first product today. My git history is basically the only proof I have that I actually learned to build something real, since I don't have a traditional resume line to point to. Turning that into something shareable is such an obvious-in-hindsight idea. Congrats on the launch!

1
回复

@luyuan_liu This is exactly who we built it for tbh. Self-taught devs have the strongest proof there is (the work itself) and nowhere to point to it. Run npx redential scan on your repo, takes a minute, no account needed, and you'll see what your history says about you. And congrats on YOUR launch today, big day for both of us!

0
回复

reading git history instead of asking people to write a portfolio is a genuinely good idea, most portfolios are curated fiction anyway. the "defend it live" step is what sells it for me - that's the part that actually separates someone who understands their own commits from someone who just merged a PR someone else wrote.

1
回复

@omri_ben_shoham1  Thanks Omri! Really appreciate it. The defense is where we're putting most of the work, first one is free if you want to try it on your own repo.

If you want to contribute to our open source repo, happy to have you there!

0
回复

A really interesting approach to the problem. How (if at all) can you prove past work after leaving a company and losing access to its code? Also, can employers view shared profiles freely?

1
回复

@mateuszkonik Timing matters tbh: the CLI reads local git history, so the move is to scan while you still legitimately have the repo on your machine (most devs scan before wrapping up a job).

Once you've lost all access there's nothing local to read, we can't conjure evidence from nothing. What you keep forever is the credential you built while you had it: the bundle lives on your profile, not in their repo.

And yes, sharing is the whole point: your profile is a link, anyone you send it to can open it, no account needed on their side.

Actually you just gave me an idea: I opened an issue for this ("how do you preserve evidence before losing access?") and see what the community comes up with. That's how our best improvements happened so far.

0
回复

The local-scan-then-defend-live split is the right trust boundary here — self-reported evidence stays labeled unverified until you actually answer for it. Before trusting a profile though, how does the scan attribute authorship inside a repo where a chunk of the committed code isn't yours: vendored deps, generated files, or co-authored/pair commits? Curious whether those inflate the skills summary, or whether you filter to authored diffs before anything gets summarized and uploaded.

1
回复

@hi_i_am_mimo Yep, filtered before anything is summarized. Skills only attribute from lines YOU actually added: vendored deps, generated files and other contributors' code never enter the analysis (author filter + path exclusions, all public in the repo). Co-authored commits inherit git's one-author-per-commit model, that's the honest limit. And it's all Attested-tier anyway: labeled claims, not verdicts.

1
回复

Congrats on the launch. The interesting part here is the defense step, not the scanning. Anyone can generate a skills list from commit metadata, and once that's a market signal people will farm it. The live technical interview about your own code is the actual hard-to-fake bit, because you can't defend architecture decisions you didn't make.

Which raises the question I'd want answered before signing up: who's asking the questions, and how deep do they go? If it's an LLM prompted from the repo summary, a good bluffer with a rough understanding of the codebase probably gets through, and the credential is worth roughly what a CV is. If it's genuinely probing ("why did you pick exactly-once here instead of at-least-once with dedupe, and what broke when you didn't"), that's worth something.

Also curious how you handle the ghostwriting problem now that a big chunk of shipped code is agent-written. The git history says it was committed under your name, not that you understood it. Maybe that's fine and the defense catches it, but it seems like the main thing standing between this and a CV.

1
回复

@aidan_codefox Depends on the evidence. From a CLI bundle the engine only sees metadata, so questions are architecture-level (why Redis, what breaks at 10x). From a connected repo it sees real files and mined fix history, so it gets specific, exactly your "why exactly-once here" type.

First question is deterministic from your detected capabilities; follow-ups are generated one at a time from the live transcript, so it drills into your last answer (capped at 3 levels per topic, server-side). Max 10 questions in 10 minutes, including one counterfactual that inverts a real fact from your history to catch people who agree with anything. Every positive verdict needs a literal quote from what you said, rules can only lower a verdict, and the model alone can never grant the top tier: Verified requires deterministic commit evidence. No human review today, that's the honest state.

On ghostwriting: we think the axis moved. The question isn't "did you type it" but "did you direct it and can you defend it". The bundle carries honest agent-involvement signals, and the defense tests exactly the part a ghostwriter can't sit for you.

Thanks for the feedback btw happy to hear more

0
回复
@jpbelmo How does the CLI handle private repos with strict compliance rules, and is there any plan to support more platforms beyond GitHub?
1
回复

@tehreem_fatima5  The CLI is git-based, not GitHub-based: it reads local git history on your machine, so GitLab, Bitbucket and self-hosted repos already work. Only a small metadata summary you review ever leaves, no code, no file names, which is exactly the point for strict compliance setups. If you meant the sign-in or the connected-repo GitHub App, those are GitHub-first today with more planned. Which part did you have in mind?

1
回复
@jpbelmo Thanks for clarifying! I was mainly asking about the sign in and connected-repo integration side of things. Good to know more platform options are planned! Also, the metadata-only approach for local history gives great confidence for compliance.
0
回复

What happens when you work in close enterprise git mostly and you are not allowed to install any unapproved apps.

1
回复

@shahrukh_khan39  Tbh that's a real constraint and I won't tell you to sneak it past IT, that's not our vibe. Good news imo: this is the easiest kind of tool to get approved. Nothing installs (npx runs it once), scan makes zero network calls, and it's fully open source so your security team can literally read every line and run the privacy tests themselves. Btw a few folks have gotten it greenlit exactly that way, worth a quick ask to your sec team.

You can check our open source repo here.

1
回复
The interesting tension here, git history proves a history, not that you're the one who lived it. Someone could scan a repo they only reviewed, then still answer basic questions about it correctly if they read the PRs closely enough. The live defense narrows that gap but probably doesn't close it for someone who prepped hard. More practical question: what happens with squashed/rebased commits or a repo migrated from a different VCS? A lot of real production repos have messy or rewritten history curious if that shows up as "can't verify" or if it just quietly produces a thinner profile.
1
回复

@thys_beesman Yeah, so local evidence can't prove you lived it, so it's our weakest tier and labeled that way. Two things help tho: scan only counts commits YOU authored (a reviewer scanning that repo gets a thin profile), and claimed emails get checked against your verified ones. The hard-prepper gap is real. Our bet: memorizing months of someone else's decisions well enough to survive live follow-ups is almost as much work as doing the work. Narrows, doesn't close. Honest state of it.

Squashed/rebased repos: no "can't verify" verdict, you just get a thinner profile (fewer commits, capabilities still detected from the diffs). And squash-merge workflows are explicitly documented as benign in the forensics, what looks bad is years of history committed in one sitting. Migrated repos land as weaker evidence, not invalid. It's all in docs/schema.md so if you wanna poke holes, best kind of contribution for us.

One more thing: replayed histories usually leave every committer date clustered in one sitting even when author dates span years, and the bundle ships both spans. A careful forger can fake both dates though, git allows it. That's exactly why local evidence is the weakest tier by design and the live defense sits on top. We never claim to detect all forgery, we claim to label what we can't verify.

0
回复

You asked for someone to try to break the trust model, so here is where I think it gives.

The live defense is generated from the same history it is meant to validate. If I fabricate a history, the questions come from my fabrication, so I answer them easily, because I am the author of that fiction. Your line holds literally, if you did the work you remember it, but writing a convincing fake history is also work and I remember that too. The defense has no ground truth independent of the artifact under test.

The second half is the part that hurts your actual users. Someone who genuinely shipped a payments integration three years ago under NDA has forgotten the specifics. Someone who generated a repo last night has them fresh. The test rewards recency, and the people it penalises are exactly the senior devs whose invisible work you built this to surface.

Is question generation weighted by commit age, or does a 2019 repo get interrogated at the same depth as one from last month?

1
回复

@abdullah_javaid3 You're right: at the Attested tier there's no ground truth outside the artifact, that's why it's the floor tier and labeled that way. What the CLI does is raise the price of the fake: signed commits can't be backdated without your key, replayed histories leave a committer-date signature the bundle ships, and the structural signals need real connected code (a webhook flow that verifies, writes and dedupes across files), not plausible-looking commits.

Also the session throws you a counterfactual (inverts a real fact from your history) and drills into your own answers, so a fake has to stay coherent under pressure. At that point faking it is basically.. doing work.

You got us on the recency thing. Today a 2019 repo gets the same depth as last month's. You basically wrote our next design issue. Wanna write it up for real? Issue #28 in the repo, that's exactly the contribution we want most.

Thanks for your feedback

1
回复

Congrats on shipping. You said the git-history trust model is what you want people to try to break, I got curious about the dashboard side instead: once a team requests a defended profile, is that visible only to the team that requested it, or could a different team account pull up someone else's pending review before it goes public?

1
回复

@vollos Thank you Chalermpon! So teams don't request defended profiles, and there's no pending review that later goes public. Devs defend their own work, and the only public thing is the dev's profile, which the dev owns and shares.

What teams get is an API: invite a candidate with a magic link, read the result of your own sessions only. Every API key belongs to one company and every read is filtered by it. Plus a private workspace to review your own candidates, with the defense audio and evidence behind each credential.

So no, another team can't pull up someone else's pending review. There's simply no shared pool to pull from.

Happy to hear what you think, and if you like to chat just schedule here

1
回复

coming at this as a CTO rather than a candidate: the "never your code, only a summary" pitch protects the code, but the summary itself can still be a leak. "webhook flow with signature checks and duplicate-payment protection, connected to auth via X" tells a competitor exactly how your payments system is architected, even with zero lines of code shown. that's the kind of detail I'd actually care about in an NDA, not the literal source. does an employer get any say over what a former engineer's scan surfaces about company repos, or is that entirely the individual's call once they have local access to the history?

0
回复

@galdayan hey gal, thanks. one factual fix first: the bundle never says "connected to auth via X", that example overshoots. What actually leaves is a closed-vocabulary slug (payments/payment-webhook-flow) and that's literally the whole disclosure. No architecture graph, no system connections, nothing.

So: the vocabulary IS the ceiling. it can only name industry-standard patterns (the taxonomy is public in the repo), so what gets out is "their front door has a lock" level stuff. Anything bespoke straight up can't be expressed, there's no slug for it, and adding one takes a public discussion plus schema ceremony. The leak surface is fixed and auditable, not open-ended.

On employer say: honest answer, nope. It's the individual's call, same as listing skills on a CV always has been, and the scan makes them confirm they're authorized. the way we think about protecting the company side isn't a veto, it's the ceiling: the tool literally can't surface more than CV-level facts. Does that cover what you'd actually worry about, or is there a case you're seeing that I'm not?

Btw this is exactly the feedback that's been shaping the product all week, keep it coming. And give the scan a try on any repo, takes a minute, no account: npx redential scan

0
回复

The failure mode I would want to know about is the committer email.

Mine is wrong. My global user.email was a colleague's address for a long stretch of work, so those commits are credited to his GitHub account and my own graph is empty for all of it. The work is real, it sits in the history, and every tool that reads git hands it to someone else.

That is the same gap you describe with NDA work, except self inflicted, and far more common than people admit.

Since Redential runs locally, it could be the one tool that gets this right. GitHub has to trust the email because all it ever sees is the push. You are reading repos I actually hold. So does attribution come from the commit email, in which case it inherits the same wrong answer, or from the fact that the repo is on my machine and I can prove I have it?

Asking because if it is the second, that is a stronger pitch than the NDA angle for a lot of people.

0
回复

@abdullah_javaid3 Attribution is by author email.

Possession can't be identity: everyone who clones a repo "holds" it, so possession-based attribution would let anyone claim any history they can download. Your case is the painful flip side: from the outside, "my commits carry my colleague's email" is indistinguishable from "I'm claiming my colleague's commits", and we refuse to guess.

So yes, today the CLI inherits git's wrong answer for that stretch, same as every tool. The one thing we add: you can claim multiple emails that are YOURS (they get checked against your verified addresses), which covers the common multi-email mess, just not the someone-else's-email one.

It's a real gap and more common than people admit, like you said. If you've got ideas for an honest fix (colleague attestation? signed retro-claims?), that's exactly the kind of thing issue #28 exists for.

0
回复

Dev credential based on what you actually built is better signal than certifications tbh. GitHub activity shows some of this but not the full story. How do you verify the contribution is genuine?

0
回复

@abdurrahman_fakhrul Two paths today: the open-source CLI (for NDA work you can't show) and the GitHub App for repos you can connect. The App gives stronger verification because we can read the actual code. The CLI earns a lower, clearly labeled tier because local evidence is forgeable by nature, and that's exactly why it's open source: so people can red-team it, and we actively invite that in the repo.

The genuine check is the live defense: video and audio, recorded on your profile. Our AI generates questions from your real history (possible bugs, how you'd solve X, why you made this decision) and matches your answers against your commits. With a connected repo that matching is strong. With the CLI it's harder, which is why it stays the lower tier and why we harden it in public.

Before any of that, the basics: only commits YOU authored count, the emails you claim get checked against your verified addresses, signed commits can't be backdated without your key, and replayed histories leave a date signature the bundle ships.

Happy to have you as a contributor in our open source repo

2
回复
#15
Overflight
Identify every aircraft in your sky.
115
一句话介绍:Overflight 是一款将 iPhone/iPad 作为随身雷达、Apple TV 作为常亮环境屏的飞行识别工具,精准解决“头顶飞过的是什么飞机”这一即时好奇痛点,提供实时航线、航司、高度及目视方向指引。
iOS Apple TV Travel
飞行雷达 飞机识别 环境显示 Apple TV 常亮屏 实时航班追踪 航空爱好者 iOS应用 增强现实 小众工具
用户评论摘要:用户高度认可 Apple TV 常亮模式与 iPhone 快查场景的差异化设计。部分用户担忧 iPhone 指南针精度影响指向功能,开发者回应已融合陀螺仪与加速计数据。多名用户建议增加 AR 叠加模式,并询问 Android 版本计划;开发者表示若 iOS 成功将考虑移植。
AI 锐评

Overflight 的价值不在于“功能多”,而在于“场景准”。它精准切割了飞行追踪市场中的两个极端需求:一是“拿起就放下”的瞬时查证(iPhone),二是“开启就忘记”的环境沉浸(Apple TV)。这种基于设备形态的体验分层,比大多数多端同步应用更懂用户实际行为——iPhone 会话平均不到一分钟,Apple TV 却长达半小时,数据本身已经证明“对的设备给对的功能”比“全平台功能一致”更重要。

产品的巧妙之处在于商业化路径。付费点(Pro)几乎全部集中于 Apple TV 场景(航线、语音播报、常亮模式、军事飞机提醒),而 iPhone 快查作为免费入口培养习惯。正如评论者所言,“容易付费的屏幕(iPhone)不常驻,常驻的屏幕(Apple TV)才是付费动机形成的地方”——开发者无需在快查环节打扰用户,而是让用户在高频沉浸感的 TV 屏上自然产生“值得升级”的念头。这种将意图捕获从入口迁移至沉浸场景的设计,比传统免费试用到订阅转化更克制、也更有效。

不过,AR 叠加层的缺失是明显短板。当前“指向哪里看”完全依赖二维地图和手机罗盘,在城市楼宇间精度极易失效。若能实现摄像头实时画面叠加飞行标签,将从“工具”跃升为“体验”。此外,作为单人作品,其长期更新风险与国际化(特别是 Android 缺席)将成为扩展瓶颈。总体来看,Overflight 是一个“小而美”的典范:它不试图覆盖所有天空,只野心勃勃地接管你家门头顶的那一片。

查看原始信息
Overflight
A live radar scope of the sky above you. Every aircraft you can actually see, plotted, with airline, flight number, route, altitude, and which way to look. iPhone and iPad in your pocket, Apple TV as an always-on ambient display. Free, with Pro for routes.
Hey, PH, dev here. I'm Charlie Wood, and DGR Labs is my a one-person shop. You might remember me from Spanning Sync, Spanning Backup, or Numerous. Overflight started with a personal compulsion: when a plane flies over my house I have to know what it is. Like, always. But every flight tracker wants to show me a map of the whole country and I wanted the opposite: just my visible sky and details about the aircraft in it. During beta testing, the Apple TV version emerged as its own thing: an ambient, always-on instrument. iPhone sessions average under a minute. Apple TV sessions average around 30 minutes. Huh! I use them both all the time: the iPhone version to answer the "What's that?" question, and the Apple TV version in my office to give me ambient awareness of my local airspace. The spoken traffic feature is IMO the hidden gem. Try it and see if you get hooked. Also try to "Precise location & direction" feature on Apple TV. I think it's pretty novel. The app is free, with a Pro tier for routes, spoken traffic, always-on mode, and "inbound interesting", a special mode for incoming military aircraft. Lots more info at https://dgrlabs.co/g/id4orc Let me know what you think.I'll be around in the comments.
2
回复

no mention of Android or a web version anywhere, is that international given the Apple TV ambient angle or just where bandwidth ran out first.

2
回复

@kyle_bennett6  TBH, I apply the Willie Sutton rule. (When asked why he robbed banks, he replied, “Because that’s where the money is.”) :) If it becomes a hit on iOS/tvOS I will look at an Android port, but until then it’s tough to justify the development/testing expense and complexity.

0
回复

"inbound interesting" mode sounds like it could get concerning fast.

2
回复

@dustin_warren Ha, yeah. But the truth is it’s mostly military transports. Fighters typically don’t advertise their location.

0
回复

the 30 min avg Apple TV session vs under a minute on iPhone tells you everything about who this is for. one's a quick lookup, other's basically a screensaver you actually want on.

2
回复

@derek_julian I’m not so sure—I think it’s just two different usage modes. I myself use the iPhone app many times daily, but only for a few seconds each time, and always to answer the question, “What’s that plane hear?” But I also leave it running on the Apple TV in my office for hours at a time.

0
回复

the way it pulls live flight data and overlays direction-of-look on the map feels really polished, like the team actually thought about how a plane spotter holds their phone.

1
回复

Love that it treats Apple TV as an always-on ambient display instead of just a second screen. Turning the TV into a living radar scope while you sit on the couch is such a thoughtful use of the form factor.

1
回复

@azizvwhr Thanks! It’s been really fun to work on.

0
回复

This is one of those apps you dont think you need until you try it. Identifying planes overhead is surprisingly satisfying. Does it work in real time or is there a delay from the ADS-B data?

1
回复

@abdurrahman_fakhrul There’s a tiny bit of latency, but it’s near-real-time.

0
回复

the "which way to look" part depends entirely on phone compass accuracy, and phone compasses drift constantly especially near buildings or in a car. have you had to build in some kind of correction for that, or does it just trust the raw magnetometer reading? that seems like the thing that'd make people distrust the app fastest if it's ever pointing you at the wrong patch of sky.

1
回复

@omri_ben_shoham1 Fair concern. I use the iOS system heading as-is, but it’s not just raw magnetometer data. CoreLocation combines the magnetometer with the gyro and accelerometer, which handles most short-term drift. If it becomes an issue, I could use the heading-accuracy estimate reported by iOS, but no one (other than you) has reported trouble with it so far. Are you seeing heading accuracy issues? If so, what device are you using? Thanks for the feedback!

0
回复

the always-on Apple TV mode is such a smart move, honestly feels like something out of a flight ops room but you know, in your living room. finally a use for that screen.

1
回复

@cemgmlcnehrtm Thanks, that’s the feel I was going for.

0
回复

The session split is not just two audiences. It puts your funnel and your engagement on opposite screens.

Almost everything in Pro is Apple TV value. Always on, ambient routes, spoken traffic, inbound interesting. That is the surface with the thirty minute sessions, so it is where the reason to pay actually forms. But the sub minute iPhone session is the one where someone answers what is that and puts the phone away, and that is the surface where paying is easy. The screen that earns the upgrade is the awkward one to buy on, and the screen that is easy to buy on is the one nobody lingers in.

Which makes the interesting move handing the moment off rather than trying to close it on the TV. Capture the intent while the ambient session is running, because that is when it exists, not eight seconds into the next lookup on the phone.

Does the tvOS app take the purchase directly, or does it send people to the phone to finish it?

1
回复

@abdullah_javaid3 That split was observed during the beta when all features were available to all users. Yes, you can purchase directly thru the Apple TV interface—no phone handoff required. Thanks for the comments!

0
回复

Love the Apple TV angle here. Turning local airspace into an ambient display feels way more natural than another big map you have to actively check.

"inbound interesting" is hilarious and a little terrifying.

1
回复

@aditya_harish_2002 Interesting inbounds don’t happen very often for most people, but when they do they’re pretty exciting. (The spoken traffic voice announcing it is even excited. 😊) thanks for the comments!

1
回复

Finally tried this on my iPad and the Apple TV ambient mode is genuinely cool, just watching planes drift across the living room screen while I work. Route info for transcontinental flights has been spot on so far.

0
回复

@koglu5218 Glad to hear you’re liking it! I’m curious: do you have spoken traffic turned on? I do that and then the volume down low. It’s nice background noise while I work, and just the right amount of interesting. :)

0
回复

Would love to see a simple AR overlay mode where you hold up your phone and it draws arrows right onto the sky pointing toward each plane's actual position, basically mapping the radar data onto the camera view so you don't have to figure out bearings from a 2D screen.

0
回复

@hafizekang6uei That would cool! I’ll put that in the feature-request hopper.

0
回复
#16
Lattics
Brain-like knowledge base with AI writing & deep research
112
一句话介绍:Lattics是一款将卡片知识库、视觉化思维导图与AI写作深度结合的桌面应用,专为学术研究、论文写作和长文创作场景设计,解决用户在分散工具间反复切换导致思路断裂的痛点。
Productivity Writing Artificial Intelligence
知识管理 AI写作 学术论文 Zettelkasten 本地优先 引文管理 思维导图 隐私保护 桌面应用 非线性格
用户评论摘要:用户赞赏本地存储和AI集成,但质疑AI处理多文档时能否准确归因来源;关注AI操作(如改写、@提及)是否为云端处理,以及能否选用本地模型;好奇AI生成的笔记能否自动纳入知识图谱,并询问PDF翻译是否保留原版式及支持扫描页。
AI 锐评

Lattics的定位聪明地卡位在了“严肃创作者”与“本地优先”的交叉点上。它的核心价值并非单纯的AI写作或知识管理,而是试图解决这两个领域结合时最棘手的两个问题:归因与隐私。

从评论看,用户的担忧非常精准。当通过“@提及”让AI综合多篇文档时,如果缺乏清晰的来源标注,那所谓的“知识综合”就变成了AI的“不负责任的黑箱操作”,对学术工作的风险极大。团队目前的回应“AI增强自动化,但人类保留控制”稍显模糊。真正的护城河在于,能否在AI输出时,对每一个观点或句子的原始出处做到精确且可见的关联,让写作不仅是“生成”,更是“溯源”。

另一个关键点是AI的部署。Lattics支持Ollama等本地模型,这是它相比Notion AI等竞品的杀手锏。但“支持”与“好用”之间差距巨大。本地小模型在理解复杂上下文、执行长文本改写时的质量,往往远不如云端大模型。如果用户为了隐私必须忍受AI质量的明显降级,那么这个“支持”就沦为了营销噱头。

至于“脑状知识库”,目前更多是手动建联,AI自动感知和插入仍是“下一步”计划。这意味着其知识图谱的智能程度目前还很初级,远未达到“系统自动生长”的理想状态。

总体而言,Lattics是一个思路清晰、定位成功的产品,但目前更像是一个“本地优先的、带AI功能的强大笔记与写作工具”,而非一个真正的“脑状AI写作系统”。它在AI与知识管理的融合深度上,尤其是归因和自动联结层面,还有很长的路要走。说“重新定义写作”为时尚早,但它确实为严肃写作者提供了一个值得关注的、更干净的起点。

查看原始信息
Lattics
Three years ago, we launched Lattics as a "brain-like" knowledge management tool for writers and researchers. Today, we're back with something much bigger — AI-assisted writing that's deeply integrated into your research and creative workflow. *** 30% OFF DISCOUNT for Produnct Hunt, Limited for 50 *** Please check the discount code in the link :https://lattics.com/specialoffer/7fGKwQZEhV, You will get 30% discount code of Lattics Pro
Lattics is a desktop writing app (Mac + Windows) designed for academic research, thesis writing, analytical reports, and creative writing. It combines a card-based knowledge base (Zettelkasten), visual mind maps & timelines, professional citation management (8000+ CSL styles), interactive charts, and PDF translation — all with privacy-first local storage. ### ✨ What's New — AI Assistance 1. AI Writing Tool (Right-click) Select any text and right-click to access AI actions: translate, proofread, rewrite, continue writing, or start a conversation about the content. No switching contexts. 2. Collaborative AI Editing 🆕 Unique AI-generated text is fully editable — you can modify it directly, add your own notes, and track changes with built-in version control. Unlike tools that treat AI output as final, Lattics lets you iterate. 3. Batch Processing with @ Mentions 🆕 Unique Type @ in AI Chat to mention multiple articles and ask the AI to analyze, summarize, or compare them together. Cross-document intelligence, no copy-pasting.
1
回复

The @ mention batch processing is the standout feature here, comparing sources without manual copy-paste is a real time-saver. But for academic work, that convenience carries a real risk: if the AI conflates or misattributes a point across multiple papers, that's not just noise, it's a citation error.

How does Lattics handle attribution when synthesizing across documents? Does it flag which source each point actually came from, or is verifying that still on the writer?

0
回复

The local-first angle is the strongest part here. For research and thesis writing, having cards, timelines, citations, PDFs, and drafts in one desktop workspace sounds much better than bouncing between a notes app, reference manager, and chat window.

The main thing I’d want to understand is how the AI layer handles privacy. When using right-click rewrite or @ mentions across multiple articles, is that always sent to a cloud model, or can users choose what leaves the machine?

1
回复

@aditya_harish_2002 You can use Ollama or your own deployed model as AI provider, and only you choose content send to model, unless you authorize AI acess some projects, but it is all under your control.

2
回复

the privacy-first local storage plus AI writing combo is the part I'd want more detail on - the moment I right-click and ask for a rewrite or start an AI chat, that selection has to leave the machine to reach whatever model is doing the work, right? curious whether that's opt-in per action, and whether there's ever a local/offline model option for people who picked this specifically for the local-storage privacy angle in the first place.

1
回复

@galdayan You can use Ollama or your own deployment model as AI provider

0
回复

curious how the AI writing layer plays with the brain-like graph structure over time - does it auto-link a new AI-drafted note into the right place in your existing knowledge graph, or do you still have to manually connect it after generating? that manual reconnection step is usually where these tools quietly turn back into a flat notes app once you've got a few hundred entries.

0
回复

the way lattics weaves AI directly into the research flow instead of bolting it on feels really thoughtful, like the team actually understood how writers get stuck before shipping anything. excited to dig into this one.

0
回复

@tekelioglu84719 We welcome you to try Lattics and write a review. Here is a redemption code for a Lattics Pro monthly membership: fkHbk03usf, which can be redeemed at this link: https://lattics.com/en-US/reseller/redeem

0
回复

The "brain-like" framing is interesting — most knowledge-base tools are still fundamentally folders and tags with search bolted on. Curious how the underlying structure actually works: is it building real semantic connections between notes over time (more like a knowledge graph), or is "brain-like" more about the retrieval/writing experience on top of a fairly standard store? That distinction matters a lot for how well it holds up as someone's knowledge base grows into the thousands of notes.

0
回复

@abhineetarora Currently, users can create content non-linearly and visually, manually establishing connections between different parts. In the future, AI will enhance automation, but humans will still retain control.

0
回复

Hey Louis, that scattered feeling of jumping between a dozen windows is exactly what wrecks my focus when I settle into a long piece of writing. Bringing calm to that mess really appeals to me, and knowing my work stays on my own computer is a lovely bonus.

0
回复

@robin_de_lacroix Yeah, same fellings. I am the heavy user of Lattics, 5 years ago, we start design and develop it, based on our expectations for long-form and serious writing

0
回复

Brain like knowledge base is a bold claim but the concept of non-linear note organization is genuinely useful. How does the AI writing side integrate with the knowledge graph, does it pull context automaticaly?

0
回复

@abdurrahman_fakhrul Integrate with knowledge graph is our next step, and will be context perception automatically

0
回复

The way Lattics pulls my scattered notes into a visual map while I draft has genuinely changed how I approach long projects. The AI suggestions feel tuned to what I'm actually working on, not generic fluff.

0
回复

I noticed the PDF translation feature. Does the translated version stay aligned with the original document layout? And can it also translate text inside images or scanned pages?

0
回复

@jerome2000 Yes, It keeps the layout same as original docs, and include scanned pages

0
回复
#17
Garmin CIRQA™ Smart Band
The screenless health + fitness tracker with no subscription
112
一句话介绍:Garmin CIRQA™ Smart Band是一款无屏幕的健康与健身追踪手环,通过自动活动检测与长达10天续航,解决用户无需频繁充电和订阅付费即可全天候监测健康状态的痛点。
Health & Fitness Wearables Fitness
无屏幕智能手环 健康追踪 健身追踪 自动活动检测 压力追踪 长续航 无订阅 Garmin 可穿戴设备
用户评论摘要:用户普遍认可无订阅模式,但核心疑问集中于:无屏幕下活动误检能否在App内手动修正?有无振动警报反馈?是否开放API?睡眠追踪精度如何?建议增加实时震动提醒功能。
AI 锐评

CIRQA™在概念上精准切中了可穿戴市场的两大痛点:订阅疲劳与屏幕冗余。在Oura、Whoop等品牌将健康数据“按月出租”的当下,Garmin凭借硬件基因打出“无订阅”牌,的确是一记有力的差异化重拳。但无屏幕设计是一把双刃剑——它砍掉了续航焦虑和交互成本,却把信息反馈的“最后10米”推给了手机。从评论反馈看,用户真正关心的不是“有没有屏幕”,而是“无屏下能否获得等效的即时反馈”:活动误检能否在App内事后修正?心率异常能否靠振动即刻打断用户?这些细节决定了CIRQA是“极简的健康黑箱”还是“聪明的无声助理”。Garmin在专业运动传感器上的积累是天然护城河,但若将自动检测做成“闷声算数”,并在交互上仅靠App单向输出,反而可能放大用户的信噪比焦虑。真正值得期待的,不是去掉屏幕本身,而是Garmin能否用更聪慧的算法与更精细的触觉反馈,让“无屏”变成一种更专注的体验迭代,而非功能的妥协。

查看原始信息
Garmin CIRQA™ Smart Band
CIRQA™ Smart Band is a screenless smart band that can automatically detect activities, has 24/7 health monitoring, stress tracking, and up to 10 days of battery.

Love the screenless approach and the 10 day battery is genuinely impressive. One thing I'd love to see is a small vibration alert for irregular heart rate or high stress moments, so the health monitoring actually nudges you in the moment instead of just showing it in the app later.

1
回复

@boran4dda I love the idea of a small vibration alert for irregular heart rate or high stress moments. I hadn't considered a screenless fitness tracker before, but a feature like that would make me consider it.

0
回复

DOES IT HAVE API??!

FINNALY!!!

1
回复

@sayyids an API would be great.

0
回复
Automatic activity detection without a screen is the part I'd stress test hardest that's usually where these bands quietly fall short. Garmin's watches already do decent auto detect, but a lot of screenless bands either misclassify similar movements (cycling vs. spin bike, walking vs. light jog) or lag by several minutes before locking in the activity type. Since there's no screen to glance at and correct it in real time, any misclassification just sits wrong in your data until you sync to the phone. Would be curious if CIRQA lets you manually retag an activity after the fact from the app, or if you're fully trusting the auto detection with no override that's the trust question a screenless design lives or dies on.
1
回复

No subscription for health tracking is really refreshing. Most wearable companies find ways to lock core features behind monthly fees. How accurate is the sleep tracking compared to full Garmin watches?

0
回复

The no-subscription model is the most interesting part here. Wearables have gotten weirdly expensive once the hardware turns into a monthly fee.

For a screenless band, I’d be curious how CIRQA handles feedback in the moment. If it auto-detects an activity wrong, can you retag it later in the app, or is the idea that the band gets accurate enough that users should not need manual correction?

0
回复

no subscription is the headline that matters most here - Whoop and Oura basically both turned into monthly fees on top of hardware you already bought, so a Garmin-quality band that skips that is a real differentiator, not just marketing. curious how you handle the no-screen part in practice though - is checking your stats mid-workout purely a phone-app thing, or is there some minimal on-band feedback (vibration patterns, LED, anything) so you're not pulling your phone out constantly.

0
回复
#18
AI Agents in Chat
Your Chat UI Just Got an AI Roommate
106
一句话介绍:CometChat 的 iOS UI Kit 让开发者无需从零构建聊天界面,即可在现有应用中嵌入支持流式响应、对话历史和自定义UI的AI智能体,解决“AI好做,聊天UI难搭”的长期痛点。
Messaging Developer Tools Artificial Intelligence
iOS UI Kit AI智能体嵌入 流式聊天 对话历史 可定制UI 开发者工具 消息平台 成本控制 群聊智能体 自助部署
用户评论摘要:用户关注AI对话成本控制与断线恢复;建议增加分支/线程转录、多用户共享AI群聊历史、无代码Webhook路由、原生Slack/Teams迁移工具。另有反馈强调自定义能力、已有auth/历史继承、跨平台配置一致性,开发团队对部分建议进行了深度回应。
AI 锐评

CometChat 这次切中的不是“AI能力”,而是“AI落地中那个最无聊却最致命的环节——聊天UI打磨”。在AI Agent产品扎堆卷模型、卷推理的当下,它清醒地意识到:企业用户缺的不是一个聪明的大脑,而是一个能直接塞进现有应用、不破坏已有体验、开箱即用的“聊天壳”。这个定位既务实又锋利。

从用户反馈也能看到,真正深入一线开发的工程师不会纠结AI多智能,而是关心“断线后消息怎么恢复”“token成本谁来管”“群聊里Agent怎么交接”——这些都是产品能否从Demo走到生产环境的关键门坎。CometChat 在回复中展现了足够的成熟度:单向对话已完备,但成本控制、群聊Agent上下文共享、跨平台一致性等问题尚未有明确路线图,这些是未来可能被对手反超的薄弱点。

此外,产品目前仅主打iOS UI Kit,Web和Android仍属计划内而非已就绪,这限制了它在多端应用场景下的即时吸引力。正面看,它的模块化架构(Agent Builder、BYO模型、MCP、前端动作)构建了可扩展的生态,真正价值在于让企业“以最小迁移成本”把AI聊天变成一个可控、可定制、可审计的业务组件,而非一个黑盒玩具。如果接下来能在成本监控、群聊协作、多端统一上给出完整方案,它将不只是UI工具,而是AI聊天的标准化基座。

查看原始信息
AI Agents in Chat
Your AI agent belongs in the chat you already built. CometChat's iOS UI Kit lets you turn your existing messaging experience into an AI conversation with streaming responses, chat history, and a fully customizable UI built in. Connect your agent, match your brand, and ship without building AI chat from scratch.

Hey Product Hunt!

Amit from CometChat here 👋


Every time someone tells us they're adding an AI agent to their iOS app, the same thing happens. They think the hard part is the AI. It's not. The hard part is that building a decent chat UI eats weeks you didn't budget for. Message bubbles, streaming, history screens; none of it is hard, it's just a slog, and it's the slog nobody warns you about.

So we stopped watching people rebuild the same thing and put it in our iOS UI Kit.

You get the actual chat experience out of the box: streaming responses, suggested prompts, and conversation history so people can pick up an old chat instead of starting from zero every time. Wire up your agent, restyle it to match your app, ship.

It's 1:1 AI conversations today. Agents in group chat is what we're building next and if that's something you'd actually use, tell me, because we're deciding how far to push it right now.

If you kick the tires, I want to hear where it annoyed you as much as where it worked. We'll be in the comments all day.

9
回复

The platform looks solid for bringing chat, voice, and video under one roof. One thing I'd love to see is a built-in branching or threaded transcript view for voice and video sessions, so it's easier to jump back to a specific moment or quote during a follow-up. Right now most tools flatten that history, and it can be a real headache for support or sales teams reviewing calls later.

1
回复

@ilk_l13311 that is quite something and I don't want to hand-wave it. What you're describing isn't a small add or a feature behavior tweak. A branchable, navigable transcript where you can grab a specific moment and carry it into a follow-up is a real product surface: timeline, quote-linking, threading a call moment back into chat. Genuinely interesting, and not at all trivial :)

Today: we've got live captions / transcription (β) for calls (English), that's the raw material the flow you're describing would sit on, but it's the groundwork, not the whole thing.

 If you can walk me through how your team would actually move through it right from reviewing a support call, to pulling a sales quote, and what a "jump back" looks like in practice I'll take that straight to the team.

A concrete flow is worth far more than a feature request here. Really appreciate you thinking at this depth 🙏

0
回复

Most platforms in this space want you to use their AI layer exclusively
so letting teams plug in their own agent logic without rewiring everything around new infra is a much friendlier

especially when i do have my own workflow

great work team

1
回复

The iOS UI Kit angle feels very practical. Streaming responses and conversation history are exactly the parts that make an AI chat feel like a real product instead of a demo screen.

For group support, I’d be curious how you’re thinking about agent handoff and context. Would multiple users be able to share one AI thread with the same history, or is the first version more about adding AI into existing group chats?

1
回复

the UI Kit part makes sense as the thing nobody wants to rebuild, but I'm curious about the layer above it: once the agent's wired in, who's watching the token spend? a chat UI invites people to just keep talking, and unlike a fixed feature, an LLM conversation doesn't have a natural stopping point. is there a built-in per-user rate limit or cost cap on the agent side, or is that something the app builder has to bolt on themselves once real usage shows up?

0
回复

The part that ate our time wasn't the bubbles, it was what happens when the stream dies halfway. iOS backgrounds the app, the socket drops, and you're left with half an assistant message that the server thinks it finished sending. We moved to persisting server side and treating the client as a replay of a message id, which killed the duplicates but made resume feel slower. Where does the UI Kit keep the partial, and on reconnect does it resume the same message or start it over?

0
回复

Dropping an agent into the messaging stack teams already shipped, instead of standing up a separate bot surface, is the right call — migration cost is usually what kills these integrations. One concrete thing: does the agent run through my existing CometChat conversations so it inherits the auth, moderation, and message history I already have, or is it a parallel agent service I wire up separately? And is the agent layer iOS-UI-Kit-only right now, or the same config across web and Android?

0
回复

A native Slack and Teams importer would be a huge win for anyone migrating from those platforms. Right now the switch feels heavier than it should, and losing years of chat history is a real blocker for teams considering CometChat as a replacement.

0
回复

@bnyaminndejjqk  

Okay this is the comment I was hoping someone would leave 😀

You're describing two things: a real replacement for Slack/Teams, and not losing your history in the switch.

We've got both.

The replacement is CometChat Airhttps://www.cometchat.com/air - team chat that never leaves your walls. Channels, DMs, threads, voice/video with recording, full-text search, in-chat AI - self-hosted or fully air-gapped, built for teams that can't put their conversations in someone else's cloud. And because Air is aimed squarely at regulated environments, it's built to the compliance bar those team actually need — ISO, HIPAA (with BAA), GDPR, AICPA/SOC, PIPEDA, COPPA.

If "considering CometChat as a Slack replacement" is the thought, Air is that product.

 On the years of history, really not the blocker it feels like. Migration is white-glove: our team brings your existing messages, users, and groups across for you, so you land in Air with your past intact instead of a blank workspace. You don't DIY it.

If you tell me where you'd be coming from (Slack? Teams?) and rough scale, I'll get you connected to the right folks at CometChat.

0
回复

Embedding agents directly into chat UI makes the interaction feel more natural than switching between tools. How customizable is the agent behavior from the developer side?

0
回复

@abdurrahman_fakhrul Thanks for seeing it the way we intended it :)

On developer control, it goes pretty deep.

Two ways in: build the agent natively in our Agent Builder, or bring your own (OpenAI, LangGraph, Mastra, CrewAI, Vercel…) and run it headless or with our UI. It's model-agnostic, so you can swap providers/models without touching the customer experience.

  
Behavior itself is shaped by:
- Instructions: tone, guardrails, hand-off rules, how it should respond. You can @-mention specific tools/MCP resources inline so the agent knows exactly when to use each one.
- Tools: call any REST endpoint, with templated context ({{uid}}, {{role}}, message metadata) injected into the request.
- MCP endpoints: connect external services/knowledge bases for richer, live context.
- Frontend actions: the agent can drive your UI (open a modal, navigate, fire a notification), not just reply with text.

So it's not a fixed bot you skin, it's the behavior, the tools, the model, and how it acts on your app, all in your hands.

Docs if you want to dig in:   

  - Overview: https://www.cometchat.com/docs/ai-agents

  - Agent Builder: https://www.cometchat.com/docs/ai-agents/agent-builder/overview

  - Instructions: https://www.cometchat.com/docs/ai-agents/agent-builder/instructions

  - Tools: https://www.cometchat.com/docs/ai-agents/agent-builder/tools/overview

  - MCP: https://www.cometchat.com/docs/ai-agents/agent-builder/mcp/overview

  - Frontend actions: https://www.cometchat.com/docs/ai-agents/agent-builder/frontend-actions/overview


What are you looking to build? Happy to point you at the exact layer.

0
回复

Love that moderation is already baked in, that's saved us headaches before. One thing that would seal the deal for me is a no-code webhook builder for routing events to tools like Zapier or Make, so non-dev teammates can wire up workflows without pinging engineering every time.

0
回复

@hasanfia0 Thanks for the love :) Yes that was a very conscious decision from the early stages of our roadmap.

The event plumbing is already there, webhooks fire on messages, user and group activity, calls, even moderation results, so the signal you'd want to route absolutely exists today. What's missing is the part you're actually pointing at: a no-code layer on top so a non-dev can wire "this event → that tool" without standing up an endpoint and pinging engineering every time.

What are some workflows that you were thinking of solving?

0
回复
#19
UltraPod
Turn your old iPhone into a music-first dumbphone
102
一句话介绍:UltraPod 将闲置旧 iPhone 变为以 Apple Music 为核心的音乐优先“功能机”,保留健康、相机、阅读等实用工具,屏蔽社交信息流与注意力陷阱,为用户提供无需购买新硬件的数字排毒方案。
Health & Fitness Music
数字排毒 功能机 旧手机改造 音乐播放器 注意力管理 物理外设 iOS 限制破解 健康追踪 电子垃圾回收 专注工具
用户评论摘要:用户普遍认可旧机再利用的环保与实用性。核心反馈聚焦于:**如何在不依赖 iOS 系统级限制下强制执行专注模式**?开发者回应采用 **3D 打印磁吸手机壳** 实现物理隔离,并配合快捷指令实现敲击背部快速切换。用户担忧短信/2FA 通知会因专注模式被错过,开发者承认需用户主动摘壳(10秒内可切换)。多数用户建议保留**相机、笔记、离线地图**作为核心非社交功能。
AI 锐评

UltraPod 的聪明之处在于,它没有试图说服用户“戒断智能手机”,而是教你如何给旧手机“穿上紧身衣”。其核心价值并非软件功能,而是那个 **3D 打印的磁吸手机壳**——这是整个产品最“硬核”的交互设计。在 iOS 封闭生态下,任何试图通过软件实现专注的方案都会面临被用户自己“三秒关闭”的窘境。而物理外壳的摩擦感,恰恰利用了人类“怕麻烦”的心理弱点,让退出专注状态比进入它更痛苦,这才是真正的行为设计。

但风险同样来自这个天才设计:为了不丢掉短信和验证码,用户必须频繁摘壳,这恰恰与“专注”的初衷背道而驰。开发者承认需要“10秒内切换”,但这恰恰暴露了产品的思维分裂——它既想当功能机,又不敢切断智能手机的急救通道。更值得警惕的是,用户反馈中提到的“内置小程序”和“社区构建”想法,正一步步将 UltraPod 推向它试图反抗的“全能应用商店”。一旦它开始让第三方开发者入驻,这就不再是一个数字排毒工具,而是另一个“精简版应用市场”。

本质上,UltraPod 是一件完美的**硬件减负配件**,却披着软件的皮。其真正价值不在于“App 做得有多好”,而在于“如何让用户舍不得卖掉旧 iPhone”,从而在环保和戒瘾之间找到一条消费主义的逃逸路线。但若创始人抵挡不住添加功能的诱惑,这个“第三条路”最终只会沦为通往另一个轻量级智能手机的岔路口。

查看原始信息
UltraPod
UltraPod turns an old iPhone into a music-first dumbphone without abandoning the useful parts of a modern device. Stream with Apple Music, keep HealthKit workouts, reading, camera, and essential local tools, while stripping away social feeds and attention traps. It is a focused daily companion for people who want a digital detox without buying another piece of hardware.
I built UltraPod because the usual digital-detox choice felt unnecessarily extreme: keep a distracting smartphone, or buy a separate dumbphone and give up streaming music, health tracking, reading, and a good camera. UltraPod takes a third path. It gives an old iPhone a focused, music-first interface built around Apple Music, then keeps only the modern tools that genuinely add value: HealthKit-powered workouts, reading, capture, and a small set of local utilities. There are no social feeds and no engagement loops. The goal is not to make the iPhone less capable. It is to make it more intentional—and to reuse hardware you already own. I would love to hear which essential, non-social function you would keep on your own dumbphone.
3
回复
Love this idea. Reusing an old iPhone instead of buying another device for a digital detox feels like a much more practical approach. I’m curious—what has been the biggest challenge? Keeping users focused while still allowing enough essential apps to make it usable every day?
2
回复

@minh_nguyen84 Actually, this is the most interesting part. I designed a magnetic 3D-printed phone case. Once installed, you can only focus—even if you want to browse TikTok, the protective case will make you give up. This isn't a coercive measure, but the psychological hint will constantly remind you of unconscious behaviors, quickly interrupting addictive patterns. After launching in China, I had over 10 testers, and their feedback was quite effective. When you really need to work or want some appropriate entertainment, you can attach the phone to the back and return to the smart device.

1
回复

cool idea Chris! I like the friction aspect of it. even a simple magnetic case is enough for our brain to give up on those behaviors.

1
回复

@gianmarco_speroni  I came up with this idea while solving my own problems.

0
回复

the physical-case-as-the-real-lock detail is honest and makes sense given iOS won't let you build a true restricted mode. the thing I'd want to know day-to-day: while the phone is in dumbphone mode under the case, do texts and calls still come through normally in the background, or does going music-first mean you'd also miss a 2FA code or an urgent text unless you pop the case off? that's usually the dealbreaker for using something like this as your actual daily phone instead of a side device.

0
回复

@galdayan  Yes, there must be compromises. There is no perfect solution yet, so the design of the phone case is crucial. Being able to switch within 10 seconds when the user needs it is the original design intention of my phone.

0
回复

What do you think about the balance between reducing digital distractions and adding more functionality to UltraPod? At what point does a focused tool risk becoming another smartphone?

0
回复

@felipe_nery  I'm already conceptualizing it. I've designed a mini-program feature inside it, where users can upload their own mini-programs (not yet open currently), which requires community building.

1
回复
@gigass Turning old hardware into a focused music and health device is brilliant! Offline navigation and emergency calls would be my top non social essentials.
0
回复

@tehreem_fatima5  The project is still being refined, haha. I've invested a lot of time in this.

0
回复
@gigass All that time invested really shines through in the execution! Rooting for UltraPod as you refine it further. Let us connect on LinkedIn so I can follow your journey!
0
回复

Makes sense — and Guided Access is manual per session (triple-click to arm each time), so the real everyday enforcement is the case, not software. That actually fits the friction-nudge-not-a-lock philosophy you re going for. For v2, would you wire a Shortcut/automation that arms Guided Access when the UltraPod app opens, so the software side matches the physical commitment?

0
回复

@leo404  I created a shortcut inside the phone where tapping the back of the device three times automatically enters the app and enables Guided Access mode, which has proven very effective.

0
回复

This is such a clever reuse play. Just curious... what was the hardest iOS limitation to work around in turning a locked-down phone into a dedicated music device?

0
回复

@derrickshowers  It should be about input methods; developing a good built-in input method requires too much engineering effort. There's also the issue with text messages being unable to access the iPhone's internal data.

0
回复

This is such a clever way to give old iPhones a second life.

I also really appreciate the pricing, it feels accessible, and the "Full Companion App" included in the Ultimate plan is a nice touch.

The only problem is... I wish I still had my old iPhone 😂

0
回复

@philipjpj iOS 16 and above works perfectly, haha, I'm working hard to share this interesting project

1
回复

Repurposing old iPhones for focused music listening is clever. Reduces e-waste and gives the phone new purpose. Does it work offline or needs internet connection for the music library?

0
回复

@abdurrahman_fakhrul  You can absolutely stream music online, listen to downloaded streaming content offline, or play local music files. You can even upload your local music to the appmusic cloud, with support for mixed playback. If you enter offline mode, undownloaded music will appear grayed out, won't be included in shuffle queues, and direct playback will prompt you to connect to the internet. You can also search for online music within the app and add it to my favorites.

0
回复

The reuse angle is the strongest part for me. A separate minimalist device can become one more thing to buy, charge, sync, and justify. Turning hardware already in the drawer into a music-first tool is the more practical version of opting out: keep the parts that serve you, remove the parts that harvest attention.

0
回复

@krekeltronics Actually, just taking an old iPhone and turning it into a usable device is already appealing enough to people.

0
回复

Love the concept of repurposing an old iPhone this way, the HealthKit and camera retention is a really smart touch. One thing that would make it stick for me is a built-in bedtime mode that automatically locks the phone into reading or music only after a set hour, so the detox actually kicks in when willpower drops.

0
回复

@n_eslem68217 Actually, I've already developed part of the sleep mode functionality, which is easy to integrate. The project is still ongoing, and I want to create something even more interesting.

0
回复

The 'third path' framing is what sells it — reusing an old iPhone instead of buying a separate dumbphone is the practical version of a detox. Since iOS won't allow a true custom launcher, how does UltraPod actually enforce the music-first setup on day one — a configuration profile, Screen Time, Guided Access? And is going back to a normal iPhone a clean one-tap revert, or does it need a full reset?

0
回复

@leo404 Similar to a launcher, the app is already live—search for UltraPod. Combined with the magnetic 3D-printed phone case, you can switch between smartphone mode and dumbphone mode whenever you need to focus. The physical protective case keeps you continuously disconnected from information streams, helping you return to real life.

0
回复

This feels like the right middle ground. Music-first, no social feeds, but still keeping camera, workouts, reading, and local tools makes way more sense than buying another device just to escape your current one.

For my own dumbphone, I’d keep music, camera, notes, and maps. Everything else can fight for its place.

0
回复

@aditya_harish_2002 The project is still in its early stages, and there's a lot of work to be done before it reaches true perfection, but I feel this direction is absolutely right!

1
回复

Thank you for building and sharing this — I really love what you've made here. 

To answer your question — if it were me, I'd probably keep music, the camera, and notes for jotting things down. That combination alone already covers most of what I actually need when I want to step away from everything else.

UltraPod is such a lovely idea. It genuinely took me back to how things felt ten-plus years ago. In an era where we run into an endless flood of new things every single day, being able to shut out the noise and sink back into your own little world might be the real luxury now.

0
回复

@kryptonite_wei Actually, I've already developed part of the sleep mode functionality, which is easy to integrate. The project is still ongoing, and I want to create something even more interesting.

1
回复
This is a great idea. It would be useful for parents who may want to pass on an old iPhone, without the socials. However can you still make calls?
0
回复

@os_ishmael Sure, I designed 3D-printed phone cases with magnetic attachment. When you need smartphone functionality, you can simply switch back to smartphone mode.

0
回复
#20
AGINE Academy
A story-driven game for learning Claude by doing
102
一句话介绍:AGINE Academy 是一款将 Claude 学习融入故事驱动游戏的互动平台,通过完成真实任务(如桌面自动化、Claude Code)解决用户“看完教程却不会用”的痛点,让技能学习直接产出可用的工具或自动化脚本。
Productivity Education Artificial Intelligence
AI学习平台 故事驱动游戏 Claude实战 任务式教学 Prompt工程 AI自动化 产品构建 低代码 技能迁移 互动教程
用户评论摘要:用户普遍认同学以致用和故事化设计,但核心质疑围绕内容时效性(Claude更新快)、评分机制(LLM裁判是否公正)、故事路径是否自适应、AI导师边界(能否回溯旧知识),以及课程是否真正支持“零门槛”构建真实产品。
AI 锐评

AGINE Academy 的亮点不在于“游戏化”,而在于它解决了AI工具教学中最隐蔽的痛点——行动瘫痪。用户不缺教程,缺的是在空白聊天框里打出第一句话的勇气。“故事+任务”的设计本质是提供了结构化的行动脚手架,把模糊的“学会Claude”拆解成77个有明确产出的微行动,这是对抗“看完就忘”的实用武器。

但产品真正的价值有待验证。第一,内容时效性是硬伤。Claude 更新快,77节课若不能实现高频自动迭代,很快会从“最前沿”沦为“历史课”,这对依赖课程收费的商业模式是致命打击。第二,评价机制的坑比想象中深。评论里提到“LLM裁判会奖励长答案”是典型的技术债务——如果玩家为了升级而投其所好,而非真正理解逻辑,游戏就变成了另一个对抗而非学习。第三,故事线的“伪非线性”风险。用户询问路径是否自适应,如果只是把固定章节套上剧情外壳,那和传统课程并无本质区别,只是多了个机器人伙伴。

最犀利的拷问是:产品是帮人“学会用Claude”,还是“学会用AGINE Academy”?如果用户离开你的平台仍然对着空聊天框发怵,那所有任务和故事只是在玩你自己的游戏。真正的成功指标不是100票,而是用户离开后能否独立写出一条高质量Prompt并调试成功。当前这款产品有成为“AI界的Duolingo”的潜力,但必须狠心解决内容永续性和深度迁移能力,否则就是一款精心包装的“高级教程”,而非改变认知的“学习引擎”。

查看原始信息
AGINE Academy
AGINE Academy makes learning Claude a story-driven game. You finish real missions inside Claude, from Desktop automations to Claude Code and your own products, with an AI mentor kept to each lesson. Every lesson leaves you a working result.
We built a game for learning Claude. AGINE Academy started with a problem we kept seeing: people watch AI tutorials, save a few prompts, and still freeze when they open an empty chat. We wanted learning Claude to feel like moving through a story, not sitting through another course. You complete short practical tasks and level up UNIT, a robot companion that grows as you learn. Every lesson ends with something you keep: a working tool, an automation, a configured skill. It starts at installing Claude Desktop and goes all the way up: Artifacts and Projects, connectors for your inbox and calendar, assistants and plugins, Claude Code, content, and building your own sites and products. 77 lessons so far, the first one is free and you can start without signing up. There's an AI mentor inside each lesson, but it's deliberately restricted to the current topic. We learned the hard way that a clever answer is useless when it has nothing to do with what the student is trying to finish. Students have already built real things with it — a marketplace CRM that auto-replies to reviews, AI sales agents on amoCRM, content factories, morning briefs that run on a schedule. Start with the free lesson: https://academy.agineai.com/ We're especially interested in where the game helps you keep going and where it gets in the way. If anything feels slow, confusing, or unnecessary, tell us directly.
1
回复

@konstantin_konovalov1 Claude changes fast, new features, workflows, even UI shift every few months. That's the real challenge with any structured course in this space.

How are you keeping 77 lessons current as Claude evolves? Is this continuously rebuilt, or do some lessons risk going stale between updates?

0
回复

How do the tasks get graded? An open prompt has a hundred passable answers, so if a model is scoring them, players quietly learn what the grader likes rather than what works. We ran an LLM judge over our own agent output for a few weeks and it turned out to be rewarding longer answers, which took ages to spot because the scores looked healthy the whole time. Is UNIT leveling up against a rubric you wrote, or against a model's opinion?

0
回复

Are these prompt design challenges?

0
回复

Learning-by-doing through a story framework is a smart way to teach something as open-ended as agentic coding — most tutorials just walk through static examples that don't transfer well to real, messy codebases. I've been building my own product with Claude Code this month, and the actual skill gap isn't syntax, it's learning to write context-rich prompts and briefs. Curious how much of the game focuses on that "how to talk to the agent" skill versus more traditional coding concepts.

0
回复

Story-driven learning for AI tools is actually smart approach, way more engaging than docs or tutorial videos. Is the content focused on Claude specifically or will it cover other models too?

0
回复

Love this idea, I think you could introduce a lot more non-tech oriented users to Claude this way, or any other AI tools for that matter.

0
回复

Learning by doing through a story format is a smart way to lower the intimidation factor of picking up Claude. How much of the "story" adapts based on what the learner actually does, versus being a fixed narrative path?

0
回复

The story-driven angle is great here. Making people finish real Claude tasks instead of just collecting prompts feels much closer to how this stuff actually sticks.

Also like the choice to keep the AI mentor scoped to the lesson. That constraint probably matters more than it sounds.

0
回复
The "AI mentor restricted to the current topic" choice is smart most learn by doing tools let the AI just solve the task for you, which defeats the point. Curious how strict that restriction is in practice: if someone's stuck on lesson 12 because of something they misunderstood back in lesson 3, does UNIT nudge them back to review, or just stay silent on anything outside scope? Also you mention students already shipped real things (marketplace CRM, sales agents on amoCRM). Are those built during the course itself, or do people finish the 77 lessons and then go build separately? Curious how much of "shipping something real" happens inside the game vs. after it.
0
回复