Product Hunt 每日热榜 2026-06-05

PH热榜 | 2026-06-05

#1
SellerClaw
A team of AI agents that runs your stores across channels
438
一句话介绍:SellerClaw 是一支由人工智能代理组成的团队,通过“监督者”模式协调专门代理(选品、店铺管理、广告投放),自动化运营 Shopify、eBay 等多渠道电商店铺,解决跨境卖家“一人多店、手动管理疲惫”的核心痛点。
SaaS E-Commerce
AI电商代理 多平台店铺管理 自动化运营 跨境卖家工具 产品上架优化 广告投放优化 库存同步 定价管理 Shopify工具 人机协同
用户评论摘要:用户高度关注跨平台库存同步一致性、广告预算冲突解决机制、无API渠道浏览器自动化的“静默失败”监控。核心问题包括:如何保证实时库存不超卖?监督者如何裁决广告提价与定价降利的矛盾?系统如何验证无API渠道的操作是否生效?热切期盼接入Amazon、TikTok等平台。
AI 锐评

SellerClaw 的亮点不在于“AI帮你干活”,而在于它试图构建一个“可驾驭的代理联邦”。它没吹嘘无脑全自动,反而强调了“人工回环”和“监督者模式”,这恰恰是当前AI落地企业级场景最务实的路径。用户评论中聚焦的“静默失败”、“跨平台库存时差”、“冲突裁决”等硬核问题,说明目标人群是真正的资深卖家,而非小白韭菜。

**价值点在于“分权与制衡”**。它将选品、广告、客服等解耦成独立Agent,再通过“监督者”基于经济模型(如单位经济学)而非逻辑优先级来裁决冲突。这比单一全能Agent更具可解释性和可控性,本质上是一套可配置的SOP自动化系统。

**隐患与挑战同样明显**。首先,“浏览器自动化”作为无API渠道的补救,其可靠性存疑,团队也承认未完全解决静默失败问题,这在扣费即亏损的电商场景中是不可承受之重。其次,高度依赖API壁垒的打通和平台政策变化,一旦Shopify或Amazon推出类似原生功能,其跨平台优势将被削弱(评论中已有此疑问)。最后,产品处于极早期(仅Shopify、eBay在线),离Demo中描述的“7x24小时代理团队”尚有差距。

总体而言,SellerClaw 切中了一个高价值但极难做的领域。它不只是一个工具,而是一次将AI代理管理与传统电商SOP深度结合的工程化尝试。真正的考验在于:当代理规模扩大、平台数量增多后,其“监督者”是否能保持决策效率,以及消费者端是否会出现因AI误操作导致的服务事故。如果能解决好“信任与归责”问题,它或将成为DTC卖家的新一代基础设施。

查看原始信息
SellerClaw
Running even one online store is a full-time job. SellerClaw is a team of AI agents that runs it for you: specialized agents for product sourcing, store management, and advertising, coordinated by a supervisor you direct. Tell it what to sell — the agents build listings, manage ads and pricing, and handle fulfillment and support across Shopify, eBay, and more. You stay in control: every action is visible and approvable, and you set how much runs on its own. Free to start.

Hey Product Hunt! 👋

Artem here, co-founder of SellerClaw. My journey in e-commerce started when I was 18, selling on Amazon. Later, I co-founded Zonesmart, where we helped over 1,000 sellers scale across borders before exiting in 2022.

Despite all the tech we had, one thing never changed: running a store is a relentless, manual grind. Sourcing, pricing, managing ads, and handling support across multiple channels like Shopify and eBay is a 24/7 job. My co-founder Kamil and I realized that to truly scale, we didn't need more tools - we needed more hands.

That’s why we built SellerClaw - a team of AI agents that actually runs the store for you.

🛠 Key Features:
👉 Specialized Agents - Dedicated AI agents for product sourcing, store management, advertising, and customer support.
👉 The Supervisor Model - You direct a "Supervisor" agent who coordinates the rest of the team.
👉 Multi-Channel - Seamless operations across Shopify, eBay, and more.
👉 Human-in-the-Loop - You stay in control. Every action is visible and can be set for your approval before it goes live. No "rogue" AI pricing.

Who is this for?
- Solo-founders who want to run multiple stores without hiring a massive team.
- Cross-border sellers struggling with local payment methods, customs, and international support.
- Growing brands that want to automate the "boring stuff" to focus on brand strategy.

Our Goal & Offer 🎁 We are in the early stages and launching to real sellers this week. We believe this changes the economics of e-commerce - decoupling revenue growth from headcount.

SellerClaw is free to start, and no credit card is required.

We need your feedback! As we refine our agents, we want to know:

Which part of your e-commerce workflow is the biggest "pain in the neck" that you'd hand over to an agent today?
What specific platforms (beyond Shopify/eBay) should our agents learn next?

We’ll be here all day to answer your questions. Let us know what you think! 👇

30
回复

@artem_kosilov Congrats on the launch, Just a quick que: when agents operate cross-border, how do you handle multi-currency and local payment method edge cases without forcing the user to manually override per market?

11
回复
@artem_kosilov love the story behind SellerClaw! Moving from running a store at 18 to building specialized AI agents that handle sourcing, ads, and support is incredible. The multi-channel coordination via a 'Supervisor' agent is a very smart approach to solving the scaling problem for solo-founders. All the best with the Product Hunt launch!
0
回复

@artem_kosilov How are ad budget limits set?

0
回复

I remember my first job -I used to sell apples when I was a kid. Now everything is sold online, and agents can handle everything for you. I already know who I’ll recommend your service to.

Good luck with today’s launch! ✨

9
回复

@maria_anosova Commerce has moved fast since those apples. Thank you for the support. When they try it, we'd want to hear how it lands.

5
回复

Most "AI runs your store" pitches mean "AI drafts your listings and reminds you to restock." The multi-channel part is where it gets genuinely hard, because inventory state has to stay consistent across Amazon, Shopify, wherever else, and any lag there turns into oversells or missed repricing windows. Curious whether SellerClaw owns that sync layer directly or sits on top of something like a middleware feed. Also wondering how it handles channel-specific policy differences, like Amazon's title length rules versus what Shopify tolerates, when the same SKU needs to live in both places.

8
回复

@fberrez1 SellerClaw connects directly via API where it exists, and through browser automation where it doesn't, without a feed aggregator in between.

Inventory and pricing sync runs through the Supplier Agent on that layer. The lag window question across Amazon and Shopify running simultaneously is a fair stress test and worth a detailed answer in the comments here. On channel-specific content: listings are generated per-channel, not from a single template. How Amazon's title length constraints get handled versus Shopify is a good follow-up to press on.

7
回复

@fberrez1 

Really good question. On inventory sync: yes, we update stock levels across platforms. If someone buys a unit on Shopify and you're running FBM (fulfillment by merchant), the count on your warehouse drops by one everywhere it's listed.



The nuance is your shipping model. If you use a third-party fulfillment provider, that service usually has its own stock-sync layer already. Where SellerClaw really earns its keep is when you ship from your own warehouse and don't have that middleware, our software handles the cross-channel sync directly. And we also have our own fulfillment solution, so if you want it fully connected, we can wire the fulfillment and the agent together end to end.


On channel-specific policy differences: our agents are trained specifically for the requirements of each platform. They account for the full set of rules a given marketplace imposes and adapt every listing to fit them, so the same SKU lands correctly whether it's on Amazon or Shopify. On top of that, we pull from external data sources to build SEO descriptions, not just the marketplace's own algorithms. Today we use DataForSEO on Shopify, and we're rolling out Helium 10 keyword data for Amazon next, so your listings are optimized against real external analytics and have a shot at ranking number one.

6
回复

@fberrez1 Thank you. We look forward to hearing your feedback 👍

4
回复
On the ‘agentic shopping’ side (ACP/UCP-style channels), what’s the hardest operational problem to keep consistent—final totals (tax/shipping), availability, returns/disputes, or post-purchase status updates—and how are you engineering the system so the promise stays consistent from chat → checkout → delivery?
7
回复

@curiouskitty 
Great question, though worth clarifying which side we're on. SellerClaw is the seller-side agent, it works in place of a marketplace manager (listings, pricing, ads, inventory, customer replies), rather than the buyer-side agent shopping through chat.


That said, the core challenge you're describing, keeping a promise consistent from chat → checkout → delivery, is exactly the kind of thing we're building toward. Our current focus is SellerClaw, but a follow-up product, SellerCart, is designed to address the agentic-shopping side directly.


Even today the same discipline applies on our end: the agent never invents numbers, it reads from the platform when they're needed, inventory stays synced, and settlement runs on the marketplace's own rails. So what the agent promises always reconciles with the system that actually fulfills it.

4
回复

@curiouskitty Thank you for your support; it means a lot to us.

4
回复

@curiouskitty Totally agree — keeping the promise consistent end-to-end is the hard part.

On the seller side, I think the biggest operational risk is availability and post-purchase state staying in sync across systems. Totals/tax/shipping can usually be computed deterministically at checkout, but inventory, fulfillment status, cancellations, and returns are where things drift fast if the agent is not grounded in live system data.

Our approach is to keep the agent decision-making layer separate from the source-of-truth transaction layer: the agent can decide and act, but the final numbers, inventory state, order status, and settlement always come from the platform or connected system itself. So the promise is only made off live data, and every action has to reconcile back to the system that will actually fulfill it.

4
回复

I like that Autonomous is not forced from day one

7
回复

@artem_anikeev 

Glad that resonated, it's a deliberate choice.


We have an advisory mode where the agent asks for your approval before each action. Most people lean on it early, and that's exactly right: like any new hire, an agent needs to be trained and checked before you trust it to run on its own. There's no magic where you flip a switch on day one and money rains down.


Think of it as an upfront investment. You put in a bit of supervision at the start, and what you get on the other side is essentially an employee who then works 24/7, no vacations, no sick days, and does it consistently. You decide when it's earned enough trust to take the wheel.

5
回复

@artem_anikeev By design. You should see how the agent behaves before giving it more room to run.


5
回复

Well done team! Question on whether the browser-automation half failing silently (eg - the marketplace ships a UI tweak, the agent thinks it repriced or paused a listing and nothing actually landed). How do you verify an action took effect on the channels you drive through the browser? How fast do you catch it when a layout change breaks the flow?

7
回复

@artstavenka1 The browser path exists for channels where there's no API to hit. For those, actions are logged and failures come through as Telegram notifications before the next cycle runs.

When a marketplace ships a UI update that breaks a flow, the agent reaches an error state and flags it. How fast that gets caught depends on how the failure looks. A clean error surfaces faster than a silent wrong-page interaction. We haven't fully solved that and won't pretend we have.

6
回复

@artstavenka1 Sharp question, and exactly the right thing to poke at.


The browser path only exists for channels that don't expose an API, so it's the fallback, not the default. For those, every action is logged, and failures surface as Telegram notifications before the next cycle runs, so a broken step doesn't just disappear into the void.


When a marketplace ships a UI change that breaks a flow, the agent hits an error state and flags it. How fast we catch it honestly depends on how the failure presents: a clean error surfaces immediately, while a silent wrong-page interaction (the agent "thinks" it repriced but nothing landed) is the harder case. That's the exact gap you're pointing at, and we haven't fully solved it, won't pretend we have. Tightening post-action verification on the browser-driven channels is an active area for us.


Appreciate you raising the silent-failure case specifically. It's the one that matters most and the one most people don't think to ask about.

6
回复

@artstavenka1 Thank you✨

6
回复

Congrats on the launch! SellerClaw tackles a real pain point, managing multiple e-commerce channels is overwhelming, and having a team of AI agents handle the heavy lifting while keeping you in control is a smart approach. My question for you: how does the supervisor agent resolve conflicting priorities between specialized agents, say when the advertising agent pushes for higher spend while the pricing agent recommends thinner margins?

7
回复

@davitausberlin The Supervisor agent doesn't resolve this unilaterally. When two agents are pulling in opposite directions on something consequential, it surfaces the conflict for your approval rather than making the call. Neither agent can push through an irreversible action without confirmed context.

You set the budget rails and the goals each agent operates within. Genuine conflicts outside those rules come to you.

5
回复

@davitausberlin Great question. The short answer: the supervisor resolves it through unit economics, not by picking a favorite agent.


The moment we list a product, we map every non-recoverable cost: category commission, fulfillment and shipping. (Those differ by model, FBA carries Amazon's own fees, while self-fulfilled orders on Shopify or other channels run on USPS, UPS, FedEx, or FBA rates depending on item dimensions and the state you ship from.) Once that full P&L picture is in place, the supervisor knows the allowable ACoS/TACoS and exactly how much of the margin can go to ads.


So in your example, profitability comes first. The pricing and advertising agents don't fight, they operate inside the same economic envelope. The system won't green-light ad spend that pushes a SKU into the red. The one deliberate exception is an investment window, where the supervisor may approve heavier ad spend on purpose to gather early reviews and build organic ranking that pays back later. But that's a conscious call, not an agent winning a tug-of-war.


In short: economics sets the boundaries, and the agents optimize within them rather than against each other.

4
回复

Spent most of this launch figuring out where to draw the line in the messaging between what SellerClaw does and what you still own. Each agent has a defined scope, irreversible actions require confirmed context, there's a full log. That's the architecture that earns the trust to expand the scope.

7
回复

Cool product! Good luck guys!

6
回复
5
回复
4
回复

@dmitry_zakharov_ai Thank you, we appreciate your support!

4
回复

How many of these features are actually available? What can we connect and actually start using now.

6
回复

@kn0wn All of the features we’ve mentioned are already live — the main limitation right now is the number of integrations. At the moment you can connect Shopify and eBay as sales channels, Meta Ads and Google Ads for ad platforms, and CJ Dropshipping for suppliers. Amazon is coming in the next 2–3 weeks, and TikTok is on the way as well.

5
回复
Nice product, I had a convo with my friend few days ago about product like this, we believe it’s great idea. But what if Shopify will launch same internal tool for their customers? How to compete with them?
6
回复

@ponikarovskii Shopify will almost certainly build something. When they do, it'll work well for sellers who live entirely on Shopify. The sellers we're building for mostly don't. SellerClaw runs across Shopify, eBay, and Amazon simultaneously. Same SKU, different rules on each channel, inventory state that has to stay consistent across all of them. A Shopify-native tool doesn't coordinate across those boundaries. I ran a cross-border platform with over a thousand sellers. Most of them were on multiple channels. The hard problems were never inside any single one.

5
回复

@ponikarovskii Appreciate that, and it's the question we ask ourselves too.

Shopify could ship a native tool, but it would only ever see Shopify. The whole point of what we're building is cross-platform: repricing and sourcing across Shopify, eBay, and CJ in one place. A seller competing on eBay's Buy Box or pulling margin from CJ can't get that from a Shopify-only feature, by definition.


Platforms optimize for their own ecosystem. We optimize for the seller wherever they actually sell. That gap is the business.

4
回复

Congrats on the launch, team! Can it work with any (even custom) stores or only with the integrations that you have (ebay, shopify, etc)?

6
回复

@danshipit 

Thanks for the question! Right now it works with our existing set of integrations, eBay, Shopify, and the others, which is fixed for the moment, rather than any arbitrary custom store.


That said, we already have a roadmap of marketplaces and ERP systems we're planning to add soon. And since we're still in active development, we're genuinely open to requests, so if there's a specific platform you'd want supported, tell us. We'd rather hear it now while we can still shape what gets built and tailor solutions to what customers actually need.


What store are you running? Happy to note it down.

4
回复

@danshipit Beyond the listed integrations, SellerClaw can work through browser automation when an API isn't available. So it's not locked to the platforms on the list. Custom setups depend on what access you can give it.

5
回复

A classic project, guys! Nowadays, anyone who doesn't implement agents risks bankruptcy. That's the reality. Good luck!

6
回复

@artem_anikeev 
Thank you, really appreciate that!


You're touching on something real. Even though the US market is relatively mature in e-commerce, with fairly predictable costs, the global trend is clear: marketplace fees keep climbing, and so do the costs of goods, production, and labor. Margins get squeezed from every direction.


That's exactly the problem we're focused on, helping sellers optimize and stay profitable while the ground keeps shifting under them. The pace of change is only accelerating, and doing things the old manual way gets harder every quarter.


Thanks for the kind words, and good luck to you too!

4
回复

@artem_anikeev Thank you! The urgency is real. For most sellers it shows up as margin getting eaten by operational overhead before revenue catches up. That's the problem SellerClaw is built around.

4
回复

Thank you for being here. This is a big day for us.

We always give 500 credits on signup. Today, for the Product Hunt community specifically: enter promo code PH1000 and get an extra 1,000 on top. 🎁 1,500 credits to put the agents to real work.

6
回复

@artem_kosilov Now that's what I call a real bonus! Thanks

4
回复

Hey Product Hunt!

The SellerClaw SaaS is a really game-changer in e-commerce market.

I've been working with Amazon, eBay, Etsy and other e-commerce platforms for 7 years and it's really challenging to monitor all day-to-day activitities across several platforms since they're completely different. I know this pain as a seller and as a CEO of the marketplace agency.

E-commerce has drastically changed over the past few years.
In 2015 you would create a listing and wait for sales. Just ship the order and receive funds to your bank account.

Now it's a huge scope of work
- market research
- import / custom clearance, duties / VAT
- changing legal environment
- certifications
- trademark issues
- product listings (photo & video content, SEO description)
- pricing (unit economy, competitior analysis)
- order fulfillment
- returns
- customer feedbacks
- loyalty programs
- ADS (marketplace traffic and external)
- UGC, influencers

And that's not the final list of the seller's activities. It requires daily control and resources.

SellerClaw will change the marketplace department in a company and will facilititate the growth when you're focused on a strategy goals instead of tracking orders shipped by USPS or changing the price manually since the FBA is becoming annualy expensive in Oct.

SellerClaw professional agents can close all the issues.

Wishing good luck to the project and we're waiting for a demo with you.

Best,
Gleb

6
回复

@tolstov_gleb How does the credit system map to actual work? Trying to predict roughly what a month costs if I run repricing daily across a few hundred listings.

5
回复

@maurya_abhiranjan Credits map to work done: small tasks use a few credits, heavier workflows use more. Our examples are ~2 credits for a customer reply, ~15 for a listing, ~40 for market research, and 100 credits = $1. For daily repricing across a few hundred listings, the exact monthly spend depends on how often you run the loop and how much reasoning is involved, but the dashboard shows estimates before big tasks and exact usage after.

3
回复

When multiple agents are working across the same business workflows, how are you handling shared context and coordination?

We've found that keeping different agents aligned on the same view of reality can be harder than the individual tasks themselves.

5
回复

@zaid_mallik1 Great question, and you're right that the coordination is the hard part.


We run a supervisor architecture. There's a supervisor agent that owns the shared view of the business and sits above four specialists: Product Scout for sourcing, an eBay manager, a Shopify manager, and an Ads manager. The specialists don't talk to each other directly. The supervisor assigns the work, hands each one the context it needs, checks the results, and resolves conflicts when two of them would otherwise act on stale or competing info.


So there's one source of truth at the top instead of four agents each guessing at reality. The Product Scout finding a SKU, the channel managers pricing and listing it, the Ads manager promoting it, all of that stays in sync because the supervisor is the one coordinating, not the agents negotiating among themselves.

4
回复

@zaid_mallik1 Great question — we’ve found the same thing. The hard part is usually not the individual agent skills, it’s keeping them aligned on one shared state.

Our approach is to have a supervisor agent coordinating specialized agents. That supervisor manages task routing, shared context, and execution order, while the underlying systems remain the source of truth for things like inventory, pricing, and order state. So instead of every agent maintaining its own view of reality, they operate through a coordinated layer that keeps decisions and actions reconciled.

4
回复

congrats on the launch! Is Amazon support already live, or are Shopify and eBay the main focus right now?

5
回复

@rustam_khasanov thanks! Shopify and eBay are live now. Amazon is coming.

5
回复

@rustam_khasanov Thanks! Shopify and eBay are the main focus right now, so those are fully live. Amazon's next — we've already been granted developer access to the SP-API and we're working on bringing it into the product soon. Stay tuned.

4
回复

@rustam_khasanov Thanks! Shopify and eBay are the main focus live right now. Amazon support is in progress and should be coming in the next 2–3 weeks.

4
回复
Love the idea and the product.Very big Big congratulations on your launch. I have just one question like how the competitive pricing side works in practice. Is it pulling live competitor prices via marketplace APIs (like Amazon SP API or eBay Browse API), or using a third party repricing data feed? And when the agent sets a price, is it optimizing for margin, Buy Box win rate, or velocity? Asking because the difference between a one time price suggestion vs. a real time repricing loop matters a lot for resellers competing on thin margins.
5
回复

@veerhunt_agai Thank you for breaking that down so specifically. Pricing is part of the agent scope and you configure what it optimizes for, with guardrails on what it can change on its own.

6
回复

@veerhunt_agai Yes — it depends on where the competitor pricing lives. If the platform exposes usable pricing data, the agent can pull it via API; if not, it can use browser automation to read prices directly from the competitor’s site.

On the optimization side, the user can choose the target metric, but the default logic is to maximize Buy Box win rate while still respecting the minimum margin threshold set by the user. So it’s not just a one-time suggestion — it can operate as an actual repricing loop within the guardrails you define.

5
回复

@veerhunt_agai Good question. Short answer: API where the marketplace exposes pricing, browser automation where it doesn't, so you're covered even on platforms that lock their data down.


Default logic targets Buy Box win rate but never below the margin floor you set, and you can switch the target metric per your strategy. It's a live repricing loop inside your guardrails, not a one-off suggestion.


Happy to go deeper on the thin-margin case if useful.

4
回复

Hey Artem! This is awesome cause online stores are usually managed manually and it's gonna change the game. Wish you all the best here!

5
回复

@german_merlo1 Thank you, I really appreciate your support!

4
回复

@german_merlo1 Thank you so much! You nailed it, running an online store today is still mostly manual, repetitive work, and that's exactly what we're out to change. Really appreciate the kind words and the support!

4
回复

Wow so I can run whole ecom business in the background :D

5
回复

@malithmcrdev That's the goal. You set what to sell and how, agents handle the rest.

5
回复

@malithmcrdev 
That's the dream we're building toward :) The way to think about it: you stay the owner making the calls, and the agents handle the daily grind in the background, sourcing, listing, pricing, orders, customer replies. You set the direction and approve the big moves, they do the heavy lifting around the clock. Less running the store, more running the business.

4
回复

This seems especially useful for dropshipping, where supplier issues show up every day.

5
回复

@alena_medvedevaa 
Exactly — the supplier side is what we built the workflow around. The agent runs the full loop: source a product from a supplier into your catalog, list it, and when an order comes in, buy it from the supplier and fulfill it. The part that matters for daily supplier issues: it keeps re-checking supplier stock and prices across your catalog on its own, automatically syncs stock-outs down to your live listings so you don't sell what can't ship, and won't complete a purchase if the supplier's cost has jumped — so price spikes don't quietly eat your margin. Proactively chasing stuck or lost shipments is still on our roadmap.

4
回复

@alena_medvedevaa Dropshipping breaks in the handoffs more than anywhere else. Supplier stock changes, data goes stale, and by the time an order lands the inventory is already gone. The Supplier Agent stays on top of that layer continuously.

5
回复

@alena_medvedevaa 
You've both said it better than I could, so I'll just put a bow on it.


Dropshipping doesn't break at the big, visible steps. It breaks in the quiet handoffs between them, the moment supplier stock shifts, a price ticks up, or data goes stale between the order landing and the fulfillment going out. That gap is where margins leak and customers get disappointed, and it's invisible until it's already cost you.


That's the whole reason the Supplier Agent exists: to live in that gap and watch it continuously, so the boring-but-critical work of re-checking stock and prices and syncing it down to your live listings just happens, without you babysitting it. You spotted exactly the pain we set out to kill.


Thanks for getting it, this is the part we're most excited about.

4
回复

Can it edit product images too, or is the listing work mostly text?

5
回复

@annmast 
A bit of both, and the image side is growing fast.


Today we have photo sourcing: if you don't have images for a product, we can find them by SKU/barcode across web sources and match them to your listing. Those tend to be the standard white-background factory shots, so think of it as covering the basics.


What's coming soon is more interesting. Based on competitor analysis and the product description, we'll generate proper infographics that highlight the product's unique selling points, sized automatically to each platform's requirements. So not just sourcing an image, but building listing visuals that actually sell.


And the step after that is A/B testing, letting you compare which visuals and content perform best according to the metrics, so the listing keeps improving over time.

4
回复

@annmast Listing work is primarily text right now: copy, titles, descriptions, structured for each channel. Image editing isn't part of it at the moment. Best way to see what it does: there's a short video on the listing, and if you want to run it yourself, use promo code PH1000 at signup for 1,500 credits free.

4
回复

If it can update tracking and message customers, that removes a lot of daily admin.

5
回复

@artyom_zhuravlev 
Exactly, that's a big chunk of the daily admin gone.


Worth noting how tracking actually reaches the buyer: on marketplaces, they see status and tracking numbers right in their account, and on Shopify they get the shipping email with the tracking number and an estimated arrival. So the baseline updates are already handled by default.


Where we add value is being smart about what's worth a message. We don't ping the customer at every leg of the journey ("now it's in this state, now the next one"), because that's just noise that ends up annoying people. Instead we notify on the moments that matter, shipped and delivered, plus we flag the exceptions: a tracking number that stops updating, a damaged label, or a cross-border parcel stuck at customs. Those are the cases where the buyer actually needs to step in, so that's where the agent reaches out.


So less routine busywork for you, and fewer pointless notifications for your customers.

4
回复

@artyom_zhuravlev Exactly — and if you have a dropshipping supplier connected, it can also fulfill orders fully autonomously, which is where it starts saving a lot of real operational time.

4
回复

@artyom_zhuravlev That's exactly the kind of admin SellerClaw takes off the plate. Fulfillment tracking and customer communication are both covered.

4
回复

Nice launch! Curious how quickly credits get used in a real store.

5
回复

@solodnev 
Thanks! Honest answer: real-world burn rate is something we're still gathering data on as more stores come online, so I won't give you a made-up number.


What I can tell you is how it works under the hood: different operations run on different models matched to the task. Routine, high-volume actions use lighter models, while the more analytical work runs on heavier ones. So credit usage scales with what your store actually does, its size, how much daily routine you automate, how much content gets generated, and so on, rather than a flat rate.


As we get more usage data, we'll be able to share real benchmarks. Great question to keep us honest on.

4
回复

@solodnev Depends on the workflow. Active sourcing and listing across multiple channels uses more than a store running mainly support. We're on a credit-based model: three plans with larger monthly balances at higher tiers, plus top-up packs if you run over. The 1,500 from PH1000 are enough to run a sourcing-to-listing workflow and see where credits go.

4
回复

Does it learn brand voice from existing listings and past replies? @kamil_bagaviev

5
回复

@shepovalovdenis 
Yes. The agents use your existing listings and past customer replies as context, so the tone they produce matches how your store already sounds rather than a generic default. You can also set explicit brand guidelines if you want to steer it further, but the baseline is learned from what you've already published.

4
回复

Your agent can integrate with Woocommerce

5
回复

@lovik1468 
Not yet, WooCommerce isn't integrated at the moment. It's something we can look at adding soon, though. Right now our focus is on Shopify as the primary platform, alongside the leading marketplaces. Appreciate you flagging it, demand like this is exactly what helps us prioritize what comes next.

4
回复

@lovik1468 WooCommerce is definitely on our priority list for upcoming SellerClaw integrations. It’s one of the platforms we’d really like to add next.

4
回复

For me the key question is control. If SellerClaw can source products, publish listings, place supplier orders and reply to customers, I’d want clear guardrails for refunds, margins, ad spend and publishing. The Advisory mode makes sense as the first step.

5
回复

@andrew_white_13 Advisory mode covers publishing and orders: every action goes to your review queue before anything runs. Budget rails handle ad spend. Margin floors and refund-specific controls are on the roadmap. Appreciated the specifics. That's useful input for us.


5
回复

@andrew_white_13 
Control is the right thing to anchor on, it's exactly how we think about it too.


Advisory mode is the foundation: publishing and supplier orders both flow into a review queue, so nothing executes until you approve it. Ad spend runs inside budget rails you set, so the agents can't exceed what you've allowed. Margin floors and dedicated refund controls are on the roadmap, the same guardrail logic, applied to those two areas next.


The principle is simple: the agents do the heavy lifting, but you set the boundaries and hold the final say on anything that matters. Really appreciate you spelling out the specifics, refunds, margins, ad spend, publishing, that's genuinely useful input as we build the guardrail layer out.

4
回复

Shopify only for now, or can this handle mixed stores too?

5
回复

@konstantin_alkhimov Not Shopify only. Shopify and eBay are live now, with more channels coming. The same agent logic runs across whichever stores are connected.

4
回复

@konstantin_alkhimov 
Mixed stores, not Shopify only. Here's where things stand right now:


Sourcing / dropshipping: CJ Dropshipping
Sales channels: Shopify and eBay (Amazon and Etsy are next on the roadmap)
Ads: Google and Facebook


And it's not one store per account, you can connect several Shopify stores and several eBay accounts to a single SellerClaw login and run all of them from one window. So if you're juggling multiple storefronts across both platforms, they all live in one place.

4
回复

Looks interesting. For me, the most important thing would be knowing what the AI is doing and why. I'd want to see the reasoning behind things like price changes, ad spending, or supplier choices, and be able to approve bigger decisions before they happen.

If I can easily track the results and stay in control when needed, I'd be much more comfortable letting the AI run parts of my store👍

5
回复

@valeriya_vovk That's the core of how SellerClaw is built. Every action is logged with context, so you can see what ran, what changed, and why. For bigger decisions, advisory mode sends them to your review queue before anything goes live. Results come through the dashboard and agent reports, so you're not digging across platforms to understand what's working. The autonomy level is yours to set and adjust anytime.

5
回复

@valeriya_vovk This is exactly the part we obsess over, and it maps really well to the day-to-day of managing marketplaces.


Every meaningful action comes with its reasoning attached: why a price moved, why ad spend shifted, why a supplier was chosen, all tied back to the unit economics behind it. So instead of manually pulling numbers to justify a decision, you have the "why" already documented. The bigger moves sit behind your approval: the agents propose, you confirm, and you set how much autonomy to hand over.


Where this really pays off in your role is reporting. SellerClaw rolls everything up into integrated reporting you can slice different ways, by channel, by SKU, by ad spend vs. margin, by period. So when leadership asks "what changed and why," you're not stitching together exports from five tools at 11pm. You get a single, defensible view of what the agents did, what it cost, and what it earned, which makes the conversation with management much easier.

Track the results, stay in control, and have the numbers ready when you need them. That's the goal.

4
回复
#2
Minimi
Your ambient memory for Claude
388
一句话介绍:Minimi是一款Mac端背景记忆工具,它能静默捕捉用户的文档、通话、消息和标签页等操作,为Claude提供即时上下文,彻底消除用户每次重复解释的痛点。
Productivity Artificial Intelligence Tech
AI助手 Mac应用 上下文记忆 Claude插件 隐私优先 本地部署 MCP协议 生产力工具 工作流自动化 智能检索
用户评论摘要:用户高度赞赏其“免提示”和“本地优先”设计,但提出三大焦点:1)如何区分信号与噪音,避免提供无关背景;2)Windows版及选择性删除特定记忆的路线图;3)长期使用后内存的时效性与准确性维护。
AI 锐评

Minimi精准击中了AI工作流中一个被长期忽视的“软肋”——上下文断裂。 当前大模型竞赛集中在参数规模和推理能力,但普通用户面对的问题根本不是模型不够聪明,而是每次都像和一个失忆的同事开会,重复解释带来的摩擦成本远超想象。Minimi用“被动监听+本地向量检索”的方案,把摩擦降到了零,这比任何“一键总结”工具都更贴近真实工作习惯。

但光鲜背后有几个不简单的坎。一是“信号vs噪音”的工程取舍。把所有操作一股脑喂给Claude,短期看着炫酷,三个月后就是失控的混沌。团队没有回避这个难题,选择用BEAM基准测试来量化准确度,54%对36%的成绩说明他们抓对了矛盾重心——上下文的质量远比数量重要。二是“长时记忆”的时效性危机。截图和聊天记录会过期,昨天定下来的决策今天可能就被推翻。目前Minimi用时间戳让Claude自己推理新旧,这不算长久之计;长期看,必须引入显性的记忆更新和冲突消解机制,做成类似个人知识图谱的结构,才能支撑真正“智能”的长期协作。三是“Windows缺失”的战略盲区。当前回应称“不兼容Accessibility API”是技术借口——大企业用户恰恰多数在用Windows,而收割B端才是收费转化最快的路径。先吃Mac重度用户没问题,但若迟迟不补全平台矩阵,只能是小圈子玩具。

最后说隐私:本地向量数据库确实比云端存文本安全得多,但“所有内容只存本地”也意味着单机丢失即永失。一份真正的“记忆”产品终归要考虑跨设备同步,那时如何“隐私优先”又将是新考题。Minimi已经证明,为AI建“第二大脑”的方向是对的,但把“大脑”做牢还要走很长的路。

查看原始信息
Minimi
Every great Claude response starts with context. Minimi listens across your Mac - docs, calls, messages, tabs - and gives Claude the full picture. No prompting. All on-device and private.
I've been living inside Claude for most of my workday, and the one thing that always frustrated me was having to re-explain myself every single session. "Here's what I'm working on. Here's what happened in my last meeting. Here's the email thread you need to know about." Minimi fixes that. It sits quietly on my Mac, reading what I read, hearing what I hear - and then feeds all of that to Claude as live context. So when I open a new chat and ask "what should I follow up on from this morning?", Claude already knows. No briefing. No copy-paste. Just the answer. A few things I love about Minimi: 1. On-device memory - your context never leaves your Mac (the vector DB lives locally). We benchmark at 54% on BEAM vs the previous SOTA's 36%. 2. MCP-native - one link, paste it into Claude's custom connector, done. No new app to live in. 3. Granular control - you pick which apps it can see. Pause anytime. If you use Claude and you work on a Mac, this is a no-brainer install. Three steps and it just works.
25
回复

@jay_gadekar so excited to have built it alongside you and our team! <3

15
回复

@jay_gadekar Many congratulations on the launch! :)

Really, really beautiful landing page, so cute, and I love the branding!

Minimi is your ambient memory for Claude, a Mac app that quietly captures everything you do on your computer (every tab, document, call, and Slack thread) and feeds it to Claude as live context.

Instead of manually briefing Claude or hunting through your history, you can just ask questions like "Who sent me the screenshot about the bug?" or "What did we decide in yesterday's meeting?" and Claude will know.

I endorse it because it's 50% more accurate than previous memory systems (54% vs 36% on the BEAM benchmark), keeps your memory on-device in a local vector database with nothing stored on the cloud, and lets you skip the extra prompting to get straight to answers.

This is exactly what AI assistants have been missing, true long-term memory that actually works while protecting your privacy.

11
回复

@jay_gadekar Congrats! Love the idea, especially that it's ambient (aka frictionless). Sadly, I'm PC - any chance you'll be doing a PC version soon?

5
回复

Minimi catches literally everything. Claude basically now has my personal context and knows everything. Helps across the workday. Remembers weeks of conversations and makes work way more productive than earlier. I can ask it "What are the tasks I should handle urgently" and it knows. I can ask it "Who all did I talk to today" and it will tell me the names, platform the conversation happened on and the context. I also use it to remember followups. Really deep use-case.

7
回复

@niketrajdwivedi - yup! Proud to have built this together :))

2
回复

@niketrajdwivedi minimi started with such a small spark - to now see it become an awesome side project. Crazy stuff.

2
回复

the context bottleneck is real. most bad AI output i see is a missing-context problem, not a model problem, so this direction makes a lot of sense. the part id be curious about is signal vs noise. passively capturing everything across docs/calls/tabs is powerful, but the risk is feeding Claude confidently-irrelevant context. how you decide what's actually worth surfacing feels like the real moat here. on-device + private is a smart trust call too. nice work.

6
回复

@ozandag Hi Ozan, even @zaid_mallik1 asked me the same question!

Here was my answer:

In Minimi - updates, contradictions, and temporal order are handled as core behavior, not patched on.

It's why we measure ourselves on BEAM rather than the older recall-only benchmarks. BEAM runs at 1M and 10M token scale and can't be solved by a bigger context window, so it directly tests the staleness question.

We're at 54% vs the prior 36% SOTA, with most of the lead on the over-time tasks.

Short version: maintaining an accurate picture beats retrieving more, every time!

2
回复

@ozandag Anyone can capture everything; the value is in what you choose to surface. We optimize for an accurate picture over raw recall, which is why we benchmark on BEAM and LongMemEval rather than recall-only tests — these run on very long conversations where the retrieval system has to surface only the relevant pieces. And keeping it on-device.

1
回复

@ozandag We are super accurate with what to surface. The underlying tech of Minimi helps with the accuracy.

1
回复

The fun technical bit: it's all local-first. Your context gets embedded and stored on your Mac, retrieval runs locally, and Claude pulls it over a single MCP connector. Nothing leaves your device. I know because I built it :))

6
回复

@vineet_gupta20 - great job, Vineet :))

2
回复

@vineet_gupta20 Being local is a big relief!

1
回复

Nice product! The on-device, you-pick-what-it-sees approach is the part that I think makes this actually look really usable. I spend my time in the Claude ecosystem too (building governance tooling around skills/access), so the granular per-app control especially caught my eye. Quick question: when you pause it or revoke an app, does the context it already captured from that app stay in the local store, or get dropped?

6
回复

@tom_palmer_ux - thanks for writing back. When you pause - say for 5 or 10 min, your memory won't be created for that duration. Please feel free to ask more queries. Good day! :)

4
回复

@tom_palmer_ux thank you for trying out Minimi! Please share your feedback with us soon :)

3
回复

@tom_palmer_ux By pausing, we don't capture anything from that window from the moment you turn it on

2
回复

Minimi is the most delightful part of my day. It has even made me a better, more thoughtful gifter haha 😛

SUPER stoked that others can now play around with it.

Here are some fun and work related things you can try doing!

Fun

  1. "What should I get Jay for his birthday?" and it actually knows, because it remembers the offhand thing he wanted three weeks ago on a call.

  2. "What was that restaurant someone raved about last month?" No idea who, no idea when. Minimi finds it.

  3. "Recommend a movie for tonight" and the pick actually is awesome, because it knows what I've genuinely been into lately.

Work

  1. "Draft a follow-up from my call with Niket" and it pulls exactly what we discussed.

  2. "What did we decide about the UX copy?" answered in one line, across scattered Slack threads, docs, and calls.

  3. "Catch me up on what I missed" after a long deep work session, so I walk back in already knowing where things stand.

Do try and let me know what you built <3

6
回复

@ojasvika_sahu the gift example — remembering an offhand thing someone said on a call weeks later — is exactly the magic tbh. but if its hearing everythiing, how does it know that one line mattered vs the 99% thats just background chatter? curious if thats tuned or you just store it all and let retrieval sort it out

3
回复

@ojasvika_sahu  - yup! Proud to have built this together :))

2
回复

@ojasvika_sahu The usecases are so insane!

0
回复

so cool!!!!!!!! kudos to the team

5
回复

@nainabajaj27 Thanks Naina!

0
回复

@nainabajaj27 Thank you as always Naina <3

0
回复

@nainabajaj27 Thanks Naina!

0
回复

A lot of memory systems seem useful while a conversation is active, but the harder test is what happens after weeks of accumulated context.

How are you thinking about memory quality over time? Is the bigger challenge helping Claude retrieve more information, or helping it maintain an accurate picture of what's still true versus what's become outdated?

5
回复

@zaid_mallik1 In Minimi - updates, contradictions, and temporal order are handled as core behavior, not patched on.

It's why we measure ourselves on BEAM rather than the older recall-only benchmarks. BEAM runs at 1M and 10M token scale and can't be solved by a bigger context window, so it directly tests the staleness question.

We're at 54% vs the prior 36% SOTA, with most of the lead on the over-time tasks.

Short version: maintaining an accurate picture beats retrieving more, every time!

3
回复

@zaid_mallik1 Hope Ojasvika's answer has clarified your question. Feel free to ask if there's anything else, Zaid.

1
回复

@zaid_mallik1 Really the right question - and honestly the harder engineering problem. Retrieval is mostly solved. Accuracy over time will need more work.

The way we think about it: Minimi captures chronologically, so context has a timestamp. Claude can reason about recency - what you discussed last week vs last month - rather than treating everything as equally current. We're also working on explicit memory updates, where newer context can surface and deprecate older facts.

The bigger unsolved problem is knowing what you consider still true. That's more personal signal than technical - we're exploring ways to let users flag it directly.

0
回复

The "no re-explaining yourself" pain point is so real — I spend a chunk of every session giving Claude context it had yesterday.

Love the on-device angle too. Privacy-first local storage is the right call when your context includes work meetings and personal projects.

One question: any Windows roadmap? That's my main blocker for trying it today.

5
回复

@dynatrading Hi Andy - the infrastructure we rely on - Accessibility - is not currently reliable for Windows, thus we have not gotten around making a windows version.

However building ambient memory for Windows is something we are absolutely going to get on very soon!

2
回复

@dynatrading  We went Mac-first to get the capture quality right every app, zero integrations, completely passive. Replicating that on Windows takes time to do properly.

2
回复

@dynatrading Hopefully soon, Andy!

1
回复

Wow. This is exactly what I need. Will come back and ask questions but excited to check this out!

5
回复

@jessica_w204 - glad to hear. Please feel free to reach out any time. Good day! :)

2
回复

@jessica_w204 Looking forward to your feedback!

0
回复

@jessica_w204 Love to hear it Jessica would love your feedback once you've tried it

0
回复

Top team, Top product.
Congrats on the launch guys!!!

5
回复

@divyansh_shourie Hi Divyansh! Thank you for your kind words - looking forward to hearing your feedback on Minimi :)

3
回复

@divyansh_shourie - thanks Divyansh :))

3
回复

@divyansh_shourie Thank you, Divyansh!

0
回复

Congrats on the launch. Most memory tools that 'always listen' wave their hands at the delete path, so I went looking for it here. When I revoke an app or delete a memory, do the vectors already sitting in the local store actually go? That's the real privacy question I believe for something that's on by default

5
回复

@artstavenka1 - great question, and you're right to push on this. When you block an app or domain, Minimi stops capturing from it going forward. Revoking or pausing fully stops all capture.

On deletion - you can't yet delete specific memories granularly, but a full app uninstall wipes the local store entirely, vectors included. Selective memory deletion is on our roadmap.

Keen to hear your feedback once you've tried it.

5
回复

@artstavenka1 Great question. It stops creating memory after you pause Minimi or block an app. The earlier memory stays but we are planning to allow selective deletion.

1
回复

Fr. Giving context to every LLM for the same thing I had it do yesterday and the day before is frustrating. About time someone built a plug-and-play memory layer and relieved me of the annoying ritual. Great work, team. Rooting for you.

5
回复

@kritarthmittal Thank you for trying out Minimi Kritarth! Really appreciate your support!

3
回复

@kritarthmittal - thanks Kritarth - for always being a true believer in us. Hope you enjoy Minimi :)

3
回复

@kritarthmittal Thanks, Kritarth!

0
回复
Woohoo! All the best team 🚀
5
回复

@suhasmotwani Thanks as always Suhas! <3

3
回复

@suhasmotwani - thank you! Do try using and share more feedback :)

3
回复

@suhasmotwani Thanks Suhas!

0
回复

Super cool product. Congrats on the launch team.

5
回复

@nikhilsheoran - thank you! Do share feedback :)

2
回复

@nikhilsheoran looking forward to hearing your feedback!

2
回复

@nikhilsheoran Thanks, Nikhil!

0
回复

Honestly....I was so tired of giving my LLM context about everything I was working on 😣
My projects, stuff about myself, my choices, my working patterns, pasting screenshots from old chats, sharing the same docs again and again.

Minimi SOLVES ALL OF IT. No need to give any context to your LLM about what you're working on or your past conversations. It captures it all, everything on your screen, and keeps the data on your device, so it's completely safe and local. If you live inside your LLM, this is the upgrade you didn't know you needed.

Do give this superpower tool a try ⚡️

5
回复

@dhanishta_likhar  - yup! Proud to have built this together :))

2
回复

@dhanishta_likhar those daily reports you create via minimi are so awesome!

2
回复

@dhanishta_likhar It indeed feels like Jarvis!

0
回复

Been using Shram for a while now and it is making my life a lot easier. Minimi is a crazy upgrade and i am loving it

4
回复

@vikrambhandari thank you Vikram - looking forward to hearing your feedback on Minimi :)

2
回复

@vikrambhandari - thanks Vikram :))

2
回复

@vikrambhandari So glad! Your feedback has helped us a lot!

0
回复

Neat idea. Can you tell Minimi to skip certain apps it shouldn't capture context from?

4
回复

@dhiraj_patel5 yes yes! You can exclude apps!

5
回复

@dhiraj_patel5 Yes absolutely! You can block apps as well as websites.

1
回复

@dhiraj_patel5 - yes, you can block apps to not make memory from on your Minimi home page :))

1
回复

Have been lucky to get early access to Minimi and my god it’s powerful! From getting random, small insights that I forgot from my meetings to tracking my work output to remembering things that I did 2 weeks ago. Minimi is like magic

4
回复

@prannay_kedia your initial feedback was critical for us to build ahead. Thank you for supporting us so early on!

2
回复

@prannay_kedia - thanks Prannay for being amongst our earliest users!

2
回复

@prannay_kedia Always glad to have your feedback, Prannay!

0
回复
Many people already try “memory” via manual notes or lightweight MCP memory servers. What’s the key product bet behind ambient capture across tabs/docs/messages/calls—and where does that approach win or lose versus a more intentional, user-curated memory workflow?
3
回复

@curiouskitty - the bet is on zero friction.

Manual notes and curated workflows ask you to decide what matters in the moment. Which means you're one busy day away from a gap. Most people don't take notes on the tab they skimmed or the offhand thing mentioned on a call - but that's exactly the context Claude ends up needing.

Ambient capture removes the decision. You don't curate, you just work - and Minimi builds the picture in the background.

Intentional memory is great for things you know you'll need. But most context isn't that - it's just ambient. That's the gap.

3
回复

@curiouskitty Manual/Curates memory/workflow will miss out on things by default. Humans aren't perfect and hence Minimi ambiently capturing everything helps.

0
回复
#3
Leni
The world’s most accurate AI for investors
350
一句话介绍:Leni是一款专为投资和商业地产团队打造的“准确性优先”AI平台,通过结构化数据检索、可验证的决策溯源和确定性计算,解决了大模型在高风险财务工作中“看似合理但数字出错”的核心痛点。
Investing Artificial Intelligence Data & Analytics
金融AI 投资决策 商业地产 数据验证 决策溯源 模型无关 确定性计算 投资研究自动化 企业级AI 财务分析
用户评论摘要:用户高度关注其“准确性”在真实工作流中的验证方式,核心质疑集中于:当数据源冲突时AI如何处理(是隐藏还是显式报告?)、如何防止错误假设被固化进“机构记忆”、以及人类审批与AI生成的结合点(谁批准了哪些假设?)。创始团队的回答强调了“显式冲突展示”和“可版本化的源头追溯”。
AI 锐评

Leni的聪明之处在于,它没有试图去和GPT或Claude比“谁更聪明”,而是直接刺穿了整个金融AI领域最虚伪的泡沫:“听起来对的答案”。大多数竞品财报分析工具的本质是“漂亮的摘要”,Leni则把赌注压在了“肮脏的溯源”上——它明白,在投资工作中,一个错误数字的危害远大于一个空白答案。

它的核心竞争壁垒不是某个“地表最强模型”,而是那套由21000+决策痕迹训练出来的“防胡诌引擎”和确定性数学计算器。这解决了业界一个长期存在的错位:企业客户C端购买的是“智能”,但在财务审计场景下,他们真正需要的是“可推翻的反刍”——Leni把决策过程变成了一个可以被检查、被版本化、被审批的“链”,而非一个黑箱。这恰恰是传统AI系统在百万美金级别的交易面前最薄弱的环节。

但需要警惕,Leni的定位使其高度依赖“冷启动”口碑。其价值完全体现在用户对其“反幻觉”能力的信赖上,一旦出现未被检测出的严重数字错误,这种信任的崩塌将是毁灭性的。同时,对于小型开发商或家庭办公室,其搭建的复杂“定义治理层”和“语义层”可能过于沉重。最终,Leni真正的对手或许不是OpenAI,而是那些已经嵌入用户业务流程、拥有海量原始数据的传统ERP或物业管理系统。它必须证明,自己不仅仅是“最准的AI”,更是团队内部数据协作不可替代的底层操作系统。

查看原始信息
Leni
Leni is the most accurate and verifiable AI for serious investment work. Built on 21,000+ decision traces and processing 100M+ rows daily, it delivers finance-grade outputs with full auditability through source links, timestamps, and grounded comps. Leni outperforms GPT, Claude, and Manus on independent benchmarks for accuracy, modeling, and valuation while giving teams the trust they need when millions are on the line. Leni is part of Google Startups and a serious machine for investors.

Hey Product Hunt 👋

I’m Arunabh, Co-Founder & CEO of Leni.

Three years ago, we started with a simple observation:

The smartest people in investing were spending an absurd amount of time moving data between systems, fixing spreadsheets, validating reports, and checking the outputs of tools that were supposed to save them time.

Everyone was talking about AI.

But when real money was involved, most professionals still didn't trust it.

And honestly, they were right.

In high-stakes work, "mostly correct" isn't good enough.

A wrong number, a missed assumption, or a hallucinated fact can cost millions.

So instead of building another chatbot, we spent years working alongside sophisticated investors, operators, lenders, and asset managers to understand what trustworthy AI actually looks like.

Since then, we've supported more than $80B in assets, processed over 100 million rows of investment data every day, built proprietary verification systems, and tested relentlessly against real-world workflows.

The result is Leni.

THE most reliable and accurate AI infrastructure platform for investors and back office work that can analyze hundreds of files simultaneously, reason through complex tasks, validate its outputs, and deliver finished work instead of just generating responses.

In independent testing, Leni now ranks among the top AI systems for spreadsheet analysis, reasoning, and resistance to hallucinations. That work also led to our selection as one of the few companies invited to Google's Gemini Forum, where we've had the opportunity to collaborate with the DeepMind team.

But what excites me most isn't a benchmark result.

It's seeing professionals finally trust AI with the work that actually matters.

Huge thank you to our team, customers, advisors, investors, and everyone who helped us get here.

We’re excited to finally put Leni and its API portal into the hands of the broader Product Hunt community and see what you build with it.

We'll be here all day answering questions, gathering feedback, and learning from the community.

My team and I are here all day. Ask us anything 🙌

P.S. 🎁 Exclusive for the Product Hunt community: Try Leni.co directly on the platform or via APIs today with code PHLENI to get 90% off your 1st month's subscription on any plans, valid till the end of the day!

36
回复

@arunabh_dastidar The focus on auditability and source-backed outputs is really compelling, especially for high-stakes investment decisions. How do users typically challenge or validate Leni's conclusions when they disagree with them? Awesome work on this!

0
回复

@arunabh_dastidar "In high-stakes work, 'mostly correct' isn't good enough." — this hits the nail on the head! 🎯 Standard LLMs are great for brainstorming, but relying on them for millions of rows of investment data is a different story. Incredible to see Leni ranking at the top for resistance to hallucinations. Congrats on the launch, Arunabh! This is exactly what the financial and back-office world needs. 🚀

0
回复
@arunabh_dastidar excellent
1
回复

@arunabh_dastidar Congrats on the launch!!

Two things I'm curious about. The model-agnostic routing, how does Leni decide which LLM handles what? Is it task-based, like one model for number-crunching and another for writing memos, or something more dynamic? And does the user get any say in that or is it fully behind the scenes?

Also, as a founder myself, I'm curious how you got the first few institutional customers to actually trust AI with real money decisions. That's probably the hardest cold start problem in enterprise AI. Did you have to start with low-stakes work and earn your way up, or did one customer go all in early?

15
回复

@devanandb thank you! Devanand, really thoughtful questions. Let me answer in two parts.

1) “Model-agnostic routing”: how does Leni decide which LLM handles what, and does the user have a say?

At a high level, it’s more dynamic than “one model for math, one model for writing,” but we do use that spirit (specialization) under the hood.

We use a planner/executor architecture:

Planner: breaks your request into a typed step graph (for example, “retrieve the right source data,” “compute and reconcile numbers,” “write the memo,” “validate outputs”).

Executors: each step is dispatched to the best “worker” for that job, which can be a different model (or tool) depending on the step’s requirements (reasoning depth, speed, cost, strictness, context window, etc.).

Results flow back to the planner, which can adapt the remaining plan based on what came back, including running verification passes.

So yes, it’s task-aware, but also context-aware and adaptive step-by-step, not a static mapping.

On user control, we do both:

For most people, it’s behind the scenes with a strong Auto default (so you don’t have to become an LLM ops engineer).

But we also believe enterprises should be able to standardize on approved models/providers, and in some cases force routing constraints for compliance, security, procurement. The workflow should not change when your firm’s model policy changes.

2) How did we get the first institutional customers to trust AI with real money decisions?

You’re exactly right: trust is the cold start problem.

The honest answer: we didn’t start by asking anyone to “trust the AI.” We started by earning trust operationally, in a few deliberate steps:

Start where the pain is high but the blast radius is controlled

Early use cases were time-sink analyst work (data pulls, consistency checks, first-pass drafts, reconciling numbers across sources) where the team could still review outputs quickly.

Win on verifiability, not vibes

Institutions don’t care if the answer sounds smart. They care if it’s right, and if you can show why. So we focused on:

deterministic data retrieval from their systems,

explicit calculations,

consistency checks,

trust-building behaviors like surfacing assumptions and tying outputs back to source artifacts.

Meet them where their data already lives, and keep it secure

A lot of early trust came from being able to operate inside the reality of institutional stacks (property management, reporting systems, deal docs, models) and being clear on security boundaries (no cross-client leakage, no training foundation models on client data, etc.).

Expand scope only after repeated “no-surprise” outcomes

Once teams saw the same level of quality across multiple cycles (monthly reporting, portfolio monitoring, underwriting support), they naturally moved from “low-stakes” to “real decisions,” because the system had already proven it could behave like a reliable analyst.

So no single customer “went all in” on day one. It was more like: prove accuracy, prove security, prove repeatability, then scale.

If you want, I can share a concrete example of what a routed step graph looks like for something like “build an IC memo, tie-out numbers to the model, generate a lender-ready package.” That tends to make the routing concept click fast.

15
回复

@arunabh_dastidar  @devanandb Firstly, love this question, especially the second part. On getting the first enterprise customers: honestly, there was no magic trick. It was hard.


A few tactical things helped us early, though. We leaned on our network, started with people who already trusted us as operators/founders, and were very clear about expectations. We didn’t go in saying “trust AI with million-dollar decisions on day one.” We positioned it more as: let us help with the painful, repetitive work first.


We also made a deliberate choice to pursue a few more respected institutional names early, even though those cycles were harder. The thinking was: if we could earn trust with sophisticated teams first, it would create confidence for everyone else later.

10
回复

Hello Product Hunt, excited to be live today with Leni. I'm Gaurav, co-founder at Leni.

Leni is an accuracy-first AI platform for investment finance and real estate teams. It helps you go from messy documents & siloed systems to structured, verifiable answers with analysis you can actually trust.


AI tools optimize for fluent responses. Leni obsesses over accuracy.

• With verification layers that validate outputs instead of "guessing."
• Decision traces so you can see how an answer was formed and what it was grounded in
• A context graph + Unified Data Model (UDM) that keeps information consistent across documents, models, and entities
• A focus on retrieval + extraction (getting the right facts) before generation (writing the response)


If you work in investments, asset management, credit, capital markets, valuation, or any workflow where a single wrong number can derail a deal, Leni is for you.


Over the years, especially in the last 6 months, it's been rewarding to see skeptics become believers. Teams that started with us as an experiment now rely on Leni for mission-critical work. That trust came from obsessing over accuracy, building robust verification systems, and learning through real implementations.

We'd love feedback from the Product Hunt community:

  1. What workflow are you trying to make "AI-native" today?

  2. Where do existing tools break down on trust/accuracy?

Thanks to our customers, team, advisors, investors, and early supporters who believed in us before this became obvious.

We're here all day so fire away with questions 🙌

P.S. 🎁 Exclusive for the Product Hunt community: Try Leni.co directly on the platform or via APIs today with code PHLENI to get 90% off your 1st month's subscription on any plans, valid till the end of the day!

15
回复

@gaurav_madani05 wait does it actually remember a project's context over time, or does every question start from a cold search? been burned by "knowledge" tools that forget everything the second i close the tab

6
回复

@gaurav_madani05 Really excited to see Leni pushing the industry in this direction. Congrats on the launch!

0
回复

@gaurav_madani05 Congratulations, Gaurav! Excited to see Leni bringing a new perspective to how teams work with information and make decisions.

0
回复
How does your “decision trace” and private context graph work over time—what gets stored, how do you prevent bad assumptions from becoming institutional memory, and how do you handle changing definitions (e.g., NOI, occupancy, same-store) across teams?
9
回复

@curiouskitty Great question. We treat both decision traces and the private context graph as versioned, auditable artifacts, not “free-form memory.”

1) What gets stored (and what doesn’t)

We store evidence + decisions, not guesses: extracted facts (with source pointers), intermediate calculations, and the final outputs/claims.

We also store the reasoning structure as a trace (what was considered, what was ruled out, and the assumptions made), but we don’t blindly promote it into reusable “truth.”

Anything that’s low-confidence, speculative, or user-specific can be tagged as ephemeral (session-scoped) vs. durable (approved to persist).

2) Preventing bad assumptions from becoming institutional memory

Every node/edge in the graph carries provenance + confidence + freshness (where it came from, how sure we are, and when it was last validated).

Nothing becomes “institutional” without a gate: either explicit human approval, or repeated confirmation across independent sources / repeated workflows.

We use contradiction detection and “challenge” steps: if new evidence conflicts with something previously stored, we don’t overwrite silently — we create a fork / flag and force reconciliation.

3) Handling changing definitions (NOI, occupancy, same-store) across teams

Definitions are treated as first-class, versioned objects (basically a “semantic layer”): NOI_v1, NOI_v2, etc., each with its formula, inclusions/exclusions, and scope (fund, asset class, team).

Any metric in the graph is linked to the definition version used at the time. So historical analyses remain reproducible, and you can re-run with a new definition intentionally.

Cross-team alignment is done via mapping + governance: you can declare “Team A NOI” ↔ “Team B NOI” with explicit deltas, rather than forcing a single ambiguous field.

Net: the system behaves more like an audit-grade knowledge + semantic layer with controlled memory persistence that's transparently shown to teams, versioned, and reversible.

14
回复

This is a strong space to build in, especially because Leni seems to sit close to real investment workflows like underwriting, portfolio reporting, market research, document review, and IC memo creation.

I was curious about the governance layer around this.

You already mention structured traces and an institutional context graph, which is interesting. How are you thinking about the human approval side of that trace?

For example, when an AI-generated underwriting note or market risk flag influences an investment memo, how do you capture who reviewed it, what assumptions they accepted, and why the team trusted that output at that point?

In investment workflows i feel like auditability feels less useful if it only shows source links, timestamps, and model history. The harder part is tracing the judgment around the decision.

Would love to understand how you are approaching this.

9
回复

@grover___dev Great question. I think about this every day - Leni’s governance protocol is a set of technical controls that keep outputs accurate, bounded, and auditable in real investment workflows:

Access & action controls: role/permission-based data access plus strict tool allowlisting so the AI can only take approved actions within defined scopes.

Provenance by default: every output carries traceability to the exact source documents/data (and versions) and the extracted fields/mappings used.

Deterministic math: financial calculations run in a controlled execution environment so numbers are reproducible and inspectable, not guessed.

Verification gates: outputs must pass consistency, definition, and reconciliation checks before they ship; if evidence is insufficient or conflicting, the system surfaces the gap instead of improvising.

Definition governance: key metrics and underwriting/reporting conventions are encoded as explicit definitions/rules so the system uses institutional meaning, not generic interpretations.

Closed-loop feedback governance: users correct issues in-line; feedback is structured into error types and updates the context graph + verification layer to prevent repeat errors over time.

Auditability & evaluation: end-to-end logging of inputs, sources, computations, and verification outcomes supports audits and continuous quality measurement.

8
回复

@grover___dev this is exactly the layer that gets interesting after source traceability.

A source link shows where a number came from, but why a team accepted the assumption, who reviewed the risk, or what changed between the first memo and the version that went to IC.

The way we think about it is that the trace has to capture more than evidence. It should capture the decision context around the evidence: accepted assumptions, open questions, reviewer notes, definition versions, and the point at which a flagged item moved from “needs review” to “approved for use.”

In underwriting, for example, a market rent assumption might be supported by comps, but the insight relies on why the team chose one comp set over another. That is a part absolutely worth preserving, because it is what people come back to later when a deal changes, a lender asks a question, or the next committee wants to know how the conclusion was formed.

And you can have that answer at your fingertips, it's just one ask away.

4
回复

Arunabh congrats on the launch! The accuracy-first framing really stands out. Most tools chase fluency and quietly hope the numbers are right, so flipping that order (retrieval and extraction before generation) feels like the correct instinct for finance. Curious how the verification layer handles a conflict - if two source documents disagree on a number, does Leni surface the discrepancy or resolve it for you?

9
回复

@tom_palmer_ux thank you! Tom, really appreciate that, and +1 on “retrieval/extraction before generation” being the right instinct for finance.

On conflicts: Leni surfaces the discrepancy first rather than silently picking a winner. When two sources disagree, we:

Show both values side by side, tied back to the exact underlying excerpts (and source docs) so you can see why they differ.

Run a verification step that tries to explain the conflict (for example, timing/cutoff dates, pro forma vs. actual, consolidated vs. property-level, unit mix changes, definition mismatches like NOI vs. NCF, etc.).

If it’s resolvable with clear rules/evidence, we’ll propose a recommended resolution (with rationale + trace). If it’s not, we’ll flag it as “needs human judgment” and let you choose which to carry forward (or keep both with a note).

12
回复

Arunabh  @tom_palmer_ux the important thing here is that we don’t think resolution should always mean the system picks one number and moves on. In finance, two conflicting numbers can both be “right” depending on context. One might be from the latest model, another from a signed operating statement. One might be actuals, one might be budget.


So the first job is to make the conflict obvious and explain why it may exist. Then, where there is enough evidence, Leni can recommend which number to use and why. But if the conflict requires judgment, we’d rather surface it rather than hide it behind a confident answer.

6
回复

Arunabh  @tom_palmer_ux 

Thanks Tom! We appreciate it!

This is one of the exact places where finance workflows break if the AI tries to be too clever. We don’t want the system to hide disagreement. Surfacing the conflict is often more valuable than forcing a single answer too early.

6
回复

Hey PH fam 👋

I've been watching AI stumble in high-stakes professional work for a while now. Real estate and investment teams can't afford hallucinated numbers in an underwriting model or a memo that cites something that doesn't exist. The cost of that mistake isn't a slap on the wrist. It's a blown deal.

That's the exact problem Leni was built to solve.

Leni is an AI agent built specifically for real estate and investment teams. Not a general-purpose chatbot pointed at your files. A purpose-built system designed to handle the kind of work where accuracy is non-negotiable.

Here's what makes it different:

🏗 It connects to the actual systems your team already uses. Yardi, Entrata, RealPage, AppFolio, ResMan and more. No manual data wrangling. No explaining your world from scratch every time.

🔍 It doesn't just generate. It verifies. Multi-agent architecture cross-checks work and reduces hallucinations before anything lands in your hands.

📊 It delivers finished work products. Underwriting models, IC memos, lease abstracts, market research with cited sources. Not a 25-message thread you have to babysit.

🔐 It's built for sensitive data. Containerized models, strong guardrails, and a private institutional context graph that gets smarter about your firm over time.

And it's model-agnostic. Use your favorite LLM or let Leni route across models automatically for the best output. You're not locked into one model's limitations.

For anyone who's been burned by AI that sounds confident but gets the numbers wrong, this one's worth a serious look.

Big respect to @arunabh_dastidar and the Leni team for tackling one of the hardest problems in enterprise AI: not just being smart, but being trustworthy.

Check it out and drop your questions below!

9
回复
@thisiskp_ great
1
回复

@arunabh_dastidar  @thisiskp_ I see amazing comments and feedback, and we are not half way through the day yet!

Keep them coming, guys!

I'm having the time of my life answering all these questions 🧐

1
回复
What are your subscription plans? Where can I find more information? I browsed through your site, but it wasn’t obvious.
8
回复

@lakshminath_dondeti - its free to try once you are in you can see all the plans starting at $25 a month.

7
回复
@arunabh_dastidar cool. Will try it out.
5
回复

Love the positioning. In investing, accuracy matters more than speed alone. A wrong model or uncited assumption can cost real money. Turning scattered docs into verified, cited memos feels like the right workflow for investment teams.

What’s the strongest early use case so far: acquisition memos, underwriting models, or portfolio reporting?

8
回复

Great question, Thami. The biggest opportunity we're seeing is investor reporting and market research, the work that drives sound investment decisions. That's where scattered, uncited data costs the most, and where verified, source-backed output changes the game. Acquisition memos and underwriting matter too, but reporting and research are where teams feel the lift first.

6
回复

Thanks@thamibenjelloun! The thing we’ve seen is that reporting (internal or external) tends to be a great wedge because it's recurring, painful, and very easy to judge. When Leni helps a team pull together a cleaner investor update, explain portfolio movements, cite the right sources, and catch inconsistencies before the meeting, the value is realized immediately.

4
回复

@thamibenjelloun I’d add one reason reporting tends to be a strong early use case: it repeats.

An acquisition memo may be high-value, but reporting creates a cycle where the same team has to explain performance, variances, leasing movement, capex, occupancy, budget vs. actuals, and portfolio changes again and again.

When Leni helps carry forward the structure, pull the right sources, and show what changed since the last period, the workflow improves each cycle. That repeatability is where teams start feeling the lift quickly

1
回复

Very proud to see Leni launch today 🎉 Working with real estate and investment data has reinforced that accuracy starts long before analysis. Good data foundations make trusted answers possible, and it has been rewarding to contribute to that work.

7
回复

@ye_tao1 thanks for being the integral part of the journey.

5
回复

@ye_tao1 really appreciate you saying this, and thank you for being part of that foundation!

1
回复
Super interesting product, congrats on the launch!
7
回复

@manuelabarcenas Thank you!!

5
回复

Appreciate it,@manuelabarcenas!

5
回复

@manuelabarcenas thank you!

5
回复

Most "most accurate AI for finance" claims fall apart the moment you ask something that requires reasoning across multiple time periods or reconciling conflicting signals in the data. What's the actual benchmark here, accuracy against what baseline, on what types of queries? And I'm curious how Leni handles cases where the underlying data sources disagree, like when reported earnings differ across filings or analyst estimates conflict with management guidance.

Congrats for the launch tho

7
回复

@fberrez1 thanks and great pushback. You’re right that for others “most accurate AI for finance” claims usually collapse as soon as you introduce (a) multi-period reasoning and (b) conflicting sources.

Here’s what we mean, concretely:

1) Benchmark, baseline, query types

We benchmark on investment workflow tasks, not generic QA. These are multi-step jobs that require retrieval, extraction, and deterministic computation across periods. Examples include roll-forwards, bridge analyses, same-store calculations, covenant and debt schedule math, lease abstraction, and memo reconciliation across multiple documents.

Baseline is frontier LLMs + standard RAG. Those stacks are strong at summarizing, but they tend to fail on exactly these tasks because they silently fill gaps, miss constraints, or do “plausible math.”

2) Why our accuracy claim holds (and why it doesn’t break on multi-period / reconciliation)

We don’t treat the LLM output as the final answer. Leni runs a separate verification step using our own verifier models trained on 31k+ decision traces, built specifically to check outputs for inconsistency, missing evidence, and violated constraints.

For numeric work, we use a deterministic math engine so calculations are executed and reproducible instead of “LLM arithmetic.”

3) What happens when sources disagree

When filings, models, memos, or third-party data conflict, Leni does not choose one silently. We:

surface the conflict explicitly (what differs, by how much, and where it came from)

keep provenance down to the exact source artifact (document section / table / cell where possible)

apply policy-driven resolution if the customer configures it (for example “audited overrides unaudited,” “newer period overrides older,” “filing overrides deck”), otherwise we flag it as requires review and route it as an exception

Benchmark coverage (third-party writeups)

https://briefglance.com/articles/niche-ai-platform-leni-outperforms-openai-google-on-key-benchmarks

https://www.wallstreetmojo.com/leni-benchmark-results-financial-research/

https://dupple.com/blog/how-leni-beat-genspark-and-manus-on-gaia-benchmark

Net: the claim is not “we have a magic model.” It’s that we built the workflow so the model cannot get away with being plausibly wrong, especially on the multi-period and reconciliation cases you mentioned.

8
回复

@fberrez1 one more way we think about accuracy: it has to hold at different layers of the workflow.

There is extraction accuracy, calculation accuracy, reconciliation accuracy, and delivery accuracy. A system can be strong at retrieval and still fail when the answer depends on tying a debt schedule to a model, or reconciling actuals vs. budget across periods.

That’s why we care about benchmarks that test different failure modes: GAIA for long-horizon task execution, SpreadsheetBench for cell-exact Excel work, Bullshit Benchmark for rejecting false premises, and DRACO for research quality. The common thread is whether the system can be checked at each step.

Please test it hard and let us know your feedback! :)

6
回复

Hey Product Hunt 👋

I’m Zain, co-founder at Leni.

A lot of our work on Leni has come from sitting close to real investment and commercial real estate workflows and seeing where AI actually breaks.

It usually isn’t the final paragraph.
It’s the step before it:

• Which rent roll did this number come from?
• Did the model use the right NOI definition?
• Why does the OM say one thing and the T12 another?
• Is this based on the latest file, or the one someone uploaded two weeks ago?
• Can this survive a partner review, lender question, IC memo, or investor update?

That is the bar we built around.

Leni helps investment and real estate teams move from scattered docs, spreadsheets, systems, and research into structured work products: underwriting support, market research, IC memos, portfolio reporting, diligence trackers, and source-backed answers.

The part I’m most proud of is that Leni is designed to slow down in the right places.

If the evidence conflicts, it should show the conflict.
If the assumption is missing, it should ask.
If a number is calculated, it should be reproducible.
If a definition changes, the system should know which version was used.
If the answer cannot be supported, it should say so.

That sounds less flashy than “instant AI answer,” but it’s what serious teams kept asking us for.

Commercial real estate teams taught us what accuracy really means in practice: numbers that tie back, assumptions that can be reviewed, sources that are easy to inspect, and outputs that hold up when real decisions are being made.

We delivered against that standard, and then pushed ourselves to take it further across spreadsheets, research, reporting, and multi-step workflows.

Excited to finally share Leni with the Product Hunt community today.

Would love to hear what you would test first:

• underwriting?
• investor reporting?
• market research?
• document review?
• internal knowledge / Q&A?
• something else entirely?

We’ll be here all day answering questions and learning from the feedback 🙌

P.S. Product Hunt community gets 90% off the first month with code PHLENI, valid today.

6
回复

@zain_nj The Product Hunt community is in for a treat. Congrats!

0
回复

@zain_nj I like everything that I’m reading. I’m impressed that Leni is going after the unsexy part of investment work. Everyone talks about AI writing memos faster, but the painful part is usually before the memo: checking if the rent roll matches the model, figuring out which version of a file is current, catching when a number moved between reports, and making sure nobody is building a conclusion on the wrong assumption. If Leni can make that part easier to inspect, that’s genuinely useful. Congrats on the launch!

0
回复

Hallucinated figures are one of the top reasons that actually helpful platforms aren't adopted. Glad the Leni team has listened to the real estate and investment teams specifically to create something custom built and something you can trust in to give you dependable results every time.

Congrats on the launch today! 🍾

6
回复

@ashamplifies - thank you, my man!! We're here to hear every feedback and suggestion!

5
回复

@ashamplifies - Thank you! And Leni is all yours to use and abuse, would love for you to try it and share your feedback :)

5
回复

@ashamplifies thank you, really appreciate it!

Commercial real estate teams were very clear with us about what “accuracy” means for them: numbers that tie back, assumptions that can be reviewed, sources that are easy to inspect, and outputs that hold up when someone is making a real decision.

We listened, built for that standard, and then challenged ourselves to take it further across spreadsheets, research, and reporting workflows.

Excited to have Leni live and get feedback from people outside our usual real estate and investment circle too!

3
回复

Congrats on the launch! 🎉

Curious — what was the biggest challenge in building an AI that investors can actually trust with high-stakes decisions? Was it the accuracy, the auditability, or getting users comfortable relying on AI for investment research? 👀

Looks like a really ambitious product. Wishing the team a successful launch day!

6
回复

@suryansh_tiwari2 Thank you!

Biggest challenge was getting reliable accuracy under real-world messiness, then making that reliability provable. Accuracy and auditability are tightly linked: you need strong extraction/reasoning, plus verification checks and traceability so an investor can see what drove the answer and where it came from.

The “comfort relying on AI” part comes last in our experience. Once the outputs are consistently correct and explainable, trust follows.

5
回复

@suryansh_tiwari2 I’d answer it a little differently: the hardest part was teaching the system when to stop.

In investment work, the dangerous failure mode is not always a bad final sentence. It is an earlier assumption that slips through and then infects the model, memo, market read, or reporting narrative downstream.

So a lot of the product work went into decision boundaries: when should Leni continue, when should it ask for a missing input, when should it run a check, and when should it say “this needs review”?

That is also why benchmarks matter to us. The useful test shows whether the system can complete multi-step work and still be checked.

This writeup covers some of that: https://dupple.com/blog/how-leni-beat-genspark-and-manus-on-gaia-benchmark

Comfort from users comes after they see that behavior: the system knows when the answer is not ready yet and will flag it.

2
回复

@arunabh_dastidar congrats on the launch! How well does this handle source data quality issues / discrepancies / missing data / disparate sources that tends to always appear in middle-market private M&A transactions? This is part of the automation puzzle I feel is the most difficult - it's whether the source data at the bottom is any good and how to efficiently correct it if it isn't

5
回复

@arunabh_dastidar  @millwiller this is one of the hardest parts of applying AI to M&A.

The system has to treat source quality as part of the job. In a real data room, the CIM, QoE, model, exports, and management deck may all be “official” in different ways, but they may not agree.

So Leni should first build a source map: what file says what, which period it refers to, whether it reconciles to the model, and where the breaks are.

If revenue by customer does not tie to the financial model, that should become an exception to resolve, not a part of a confident summary.

That is also why we obsessed with benchmarks around grounded research and spreadsheet execution.

The output has to be checkable across documents and calculations, especially when the source package itself is imperfect:

https://briefglance.com/articles/niche-ai-platform-leni-outperforms-openai-google-on-key-benchmarks

5
回复

@millwiller Totally agree. In mid-market private M&A, the hardest part is often not the model, it’s the messiness of the source layer.

Where Leni does well is treating “data quality” as a first-class problem, not an afterthought:

  • Provenance + traceability: we keep citations back to the exact source snippets used, so you can see what the system relied on and where it came from (and quickly spot when a source is wrong or outdated).

  • Discrepancy detection: when multiple sources disagree (or values don’t reconcile), we flag it explicitly rather than forcing a single answer. You get a “here are the competing values + confidence + why” view.

  • Missing/partial data handling: we’ll return structured outputs with gaps clearly marked, plus a list of the specific fields and documents that would resolve the gaps (vs. vague “need more info”).

  • Efficient correction loop: the practical win is that once you correct or confirm a value, that resolution becomes a reusable context for the next deal and next report, instead of re-litigating the same discrepancy every time.

Net: the goal isn’t to pretend the bottom-of-funnel data is clean. It’s to (1) surface what’s unreliable, (2) quantify the uncertainty, and (3) make the “fix” workflow fast and auditable.

If you have a concrete example (e.g., NOI, rent roll, debt schedule, capex history) where sources commonly conflict, I can tell you how we typically set up the checks and the correction flow for that pattern.

2
回复
The combination of AI-powered research and clear source attribution is refreshing. Looking forward to seeing how investors and analysts incorporate this into their workflows.
4
回复

@tanjum Thank You, Tanjum.

0
回复

Most AI tools help you find answers. Leni seems focused on helping users understand why those answers make sense. That distinction is incredibly important in investing. Great launch!

3
回复

@1mirul - Thank you!

0
回复

Congrats on the launch. The auditability angle is what got me, since that's the part most finance AI tools gloss over. Curious how you're thinking about third-party verification down the road, like giving an auditor or LP a way to independently confirm an output and its sources. Either way this looks really strong.

1
回复
#4
Veltrix AI
AI finance copilot for cash flow, margins, and growth
271
一句话介绍:Veltrix AI 是一款面向创始人及财务团队的AI财务副驾,通过自然语言问答,实时整合跨平台数据,解决财务数据分散、报表滞后导致的决策延迟痛点。
Artificial Intelligence Data & Analytics Data
AI财务副驾 现金流管理 实时数据分析 自然语言查询 QuickBooks集成 Shopify整合 财务自动化 商业智能 智能预警 中小企业工具
用户评论摘要:用户普遍认可其简化财务分析的潜力。主要疑问包括:是否具备真正的现金流预测能力,而非仅事后查询;具体支持哪些数据源及如何处理混乱数据;能否管理多店铺及多账本;历史数据回溯深度及移动端规划。团队回应称已能主动预警异常,支持CSV/PDF等杂乱数据,多工作区管理可行,移动端在规划中。
AI 锐评

Veltrix AI 切中了一个非常实在的痛点:财务数据分散在QuickBooks、Shopify、HubSpot等多个系统中,导致决策滞后。其“用自然语言问财务问题”的交互设计,确实降低了财务分析的门槛,尤其对于没有专职数据分析师的中小企业主而言,这比传统BI工具或Excel更友好。

然而,值得警惕的是,许多用户评论中体现出的核心疑虑并未被彻底打消。当被问及“是否具备前瞻性预测能力”时,团队的回答仍停留在“主动预警异常趋势”上,这本质上仍是对历史数据的模式识别而非真正的预测建模。对于高敏感度的现金流管理,预告式的“预警”与“预测”之间的鸿沟,恰恰是决定产品能否从“好用的查询工具”升级为“不可或缺的财务大脑”的关键。

从技术层面看,Veltrix 最大的护城河不在于AI问答本身,而在于对混乱商业数据的接入和清洗能力。它强调支持CSV、PDF等杂乱格式,并打通多个官方实时API,这比通用大模型的理解力更贴近实际业务场景。但这也意味着其市场拓展高度依赖账务(QuickBooks/Xero)、电商(Shopify/Squarespace)及营销(HubSpot)生态的稳定性。一旦生态位被深耕该领域的垂直服务商(如专门做Shopify财务分析的App)蚕食,通用性反而可能成为弱点。

更现实的问题是定价与用户留存。虽然新手入门免费,但面对快节奏的小型企业,如果非要把简单查询包装成“AI副驾”的噱头,而实际报表生成速度和数据准确性无法显著优于手动操作,用户尝鲜后很可能会流失。Veltrix 的价值,最终要落在“能否让一个不懂SQL的老板,在30秒内做出一个招人决策”这种具体且高频率的闭环场景上。目前来看,它完成了第一步——让数据“可问”,但距离让数据“可决策”还有一段硬仗要打。

查看原始信息
Veltrix AI
Veltrix AI gives founders and finance teams instant clarity on cash flow, profitability, burn, and business performance. Connect QuickBooks, Xero, Shopify, Square, and HubSpot, then ask finance questions in plain English to get source-backed answers, anomalies, and recommended next steps. Replace spreadsheet chaos and static dashboards with real-time financial intelligence built to help you make faster, smarter business decisions.

Hey Product Hunt 👋

We built Veltrix AI because most business owners still don’t get answers from their financial data fast enough to act on them.

  • Your sales data lives in Shopify or Square. 

  • Your accounting sits in QuickBooks or Xero. 

  • CRM and marketing data lives somewhere else. 

By the time someone exports CSVs, builds dashboards, or sends monthly reports, the moment to make a decision has already passed.

Veltrix changes that.

Veltrix AI is an AI finance copilot that connects directly to the tools businesses already use QuickBooks, Xero, Shopify, Square, HubSpot, and more and turns raw financial and operational data into plain-English answers, insights, alerts, and recommended next steps.

Instead of digging through spreadsheets or building dashboards, you can ask questions like:

• “Can we afford to hire another salesperson?”
• “Why is cash flow tighter this month?”
• “Which customers are paying late?”
• “Are our ad campaigns actually profitable?”
• “What vendors increased spend unexpectedly?”

…and Veltrix answers instantly using your real business data.

What makes Veltrix different from traditional BI tools or AI data analysts?

Most analytics tools still expect you to:

  • Know SQL

  • Build dashboards

  • Create formulas

  • Understand financial modeling

  • Hire analysts to interpret reports

Veltrix is designed for operators, founders, and finance teams who need answers, not more dashboards.

Here’s what Veltrix does for you instead:

📊 Gives clear answers from your actual business finances

💡 Surfaces proactive insights before you know what to ask

⚠️ Flags unusual changes and problems

📈 Creates automated AI powered dashboards and reports

✅ Shows next steps you can act on

🔎 Shows what each answer is based on

Who is Veltrix for?

  • Founders & CEOs who want a live pulse on the business

  • Finance teams that need faster reporting and analysis

  • Agencies & ecommerce brands tracking profitability

  • SMB owners without dedicated data teams

  • Operators managing cash flow and growth decisions

We also built Veltrix with privacy and control in mind:

  • Read-only connections

  • Private workspaces

  • No selling customer data

  • Disconnect anytime

🎁 Exclusive for Product Hunt:

Use code HUNTVLTRX26 for 2 free months of the Starter plan.

Try Veltrix and ask your toughest business question. I’ll be in the comments all day.

28
回复

@veltrixms Wow, congratulations on the launch!

1
回复

@veltrixms I'm proud to be a part of Velrix! Thanks Guys!
Thanks @Product Hunt !

8
回复

I'm so excited to share this new version of the product! We worked so much, and it is such an amazing feeling seeing the product that is truly useful to owners, founders, bookkeepers, finance specialists and advisors. Lots of early feedback we've got was so rewarding. Now, I'm happy Product Hunt community gets to see the new Veltrix too:) The team is here in the comments all day - share your thoughts, comments, questions, feedback. Thank you all for the support!

14
回复

@bohdan_sitar Thank you, Bohdan! It's an amazing feeling indeed. Thanks to all early testers!

9
回复

I'm thrilled to finally get this version of the product out there! We've been working hard to bring a real value to people, especially those that are struggling with getting clarity on businesses they are running! I truly believe this will help a lot! Congrats to everyone!

12
回复

@nazar_parashchuk Thank you, Nazar! Congrats!

7
回复

I’m really happy to be part of this product!

Veltrix makes it incredibly easy to get fast and accurate answers from my integrations and uploaded files. What I especially like is how it helps make financial data easier to understand by providing clear explanations, not just numbers.

Beyond answering questions, it also surfaces valuable insights and highlights areas that may need attention or improvement. It’s a great way to stay on top of business performance without spending hours digging through reports and dashboards.

Congratulations to the entire team on the launch!

12
回复

@roman_zayac Thank you, Roman! So excited!

6
回复

Most finance copilots I've seen are good at surfacing numbers you already have and bad at the part that actually matters: flagging when a trend is about to become a problem rather than explaining one that already happened. Curious whether Veltrix is doing anything genuinely predictive on cash flow timing, or whether the "copilot" framing mostly means natural language queries over your existing data. Also wondering which data sources you connect to out of the box, specifically whether it handles messier inputs like a mix of Stripe, QuickBooks, and manual CSV exports without a lot of cleanup work on the user's end.

11
回复

@fberrez1 Hi Florent! Thank you for the support and such a good question! Yes, we've built Veltrix specifically as a copilot which proactively surfaces your data to flag what changed, what looks unusual, any anomalies, any duplicated charges, good/bad trends, and what may need checking. It will show you everything even before you ask a question!

You can connect to QuickBooks, Xero, HubSpot, Shopify, Square out of the box - all using official real-time integrations. We're rolling out more integrations in the next weeks, including Google and Microsoft Workspaces to talk to the data living there.

And lastly - Veltrix can handle messy data perfectly. You can drop any csv, xlsx or pdf - Veltrix will process everything and make it ready for analytics or any your questions. No need to spend time consolidating data - just drop whatever you have.

Give it a try, and let me know what you think. Excited to see your feedback!

8
回复

So excited to finally share this with you all!
As a developer, it feels great to see new version of Veltrix finally out here today. We worked hard to make it easy to connect your tools and get clear answers from your data, not just numbers. What I like most is how simple it makes things for people who are not data experts.
Can't wait to see how people use it!

10
回复

@oleksandr_drohomyretskyi2 Thank you, Oleksandr! Can't wait as well. All the feedback so far been amazing and critical, which is a blessing.

6
回复

The excitement today is so real!
We’ve been counting down the minutes to bring this new version of Veltrix to the Product Hunt community. We built this product because managing business finances shouldn't give anyone a headache, and this major update makes it smoother than ever. Having it out here in the open today is an amazing milestone for us. Drop any questions or feedback below, the whole team is fired up and ready to chat! 🙌

10
回复

@_mkcd_ Thank you, Mark! Absolutely, now we can finally say no to 'finance headache'. I'm so proud of Veltrix, a product that can truly help startups, businesses and anyone working with finances. Finally, we can focus on what matters, and not spending time trying to make sense out of numbers.

6
回复

How long does it take to process my business data before actually interacting with it?

10
回复

@a_bryant Hi Angela! It's instant for any of the integrations (QuickBooks, Xero, HubSpot, Shopify, and more). We're using official real-time integrations, so the data is ready for the use right away. Moreover, Veltrix will surface all the data first and show you a proactive overview, including any anomalies, duplicates and action items! If your data lives in files or Google Drive, then the data processing will take a minute or so. Please, give it a try and let us know what you think!

11
回复

Congrats on the launch! Love the idea. I work with business owners, and I'm wondering - can I manage multiple Shopify stores and QuickBooks accounts with Veltrix? Would love to speed up my monthly review.

8
回复

@nata_sv Thank you so much, Nata! Yes, you can. We have multi-workspace management and you can invite team members too. So, definitely - yes.

1
回复

Proud of my team. Watching every part of the company move in the same direction at the same time was the part I won't forget.

To the people who asked the hard questions in the comments — Xero specifics, multi-entity, what's missing — you're the reason the next version will be better.

Thank you to everyone who supported!

8
回复

@mykola_svystun Thank you, Mykola! Glad to have you on the team!

2
回复

Two years of work goes live today — Veltrix is on Product Hunt Here!

I work on integration development and support! The part I'm proudest of: you ask a question about your business in normal language, and it answers with the actual numbers — no exports, no dashboards.

8
回复

@maksym_shcherbakov1 Congrats! And it's also all in real-time. Anytime you ask, it's always based on your actual data right now. Amazing work!

7
回复

Hi! Congrats!
This looks useful for small teams. Can Veltrix highlight trends over time, like spotting a gradual drop in cash flow before it becomes critical?

7
回复

@alexkhlystova Thank you, Oleksandra! Yes, Veltrix proactively surfaces your data to alert highlight any critical changes in cash.

1
回复

Amazing product by an amazing team!!!

7
回复

@pascal_weinberger Thank you, Pascal! I appreciate the support!

5
回复

Dashboard fatigue is a very real thing...
love the concept of just getting straight answers. Quick question - when I connect QuickBooks, how far back does the data history go for Veltrix to analyze? And are you guys thinking about a mobile app or WhatsApp integration for quick questions on the go? Upvoted!

7
回复

@sara_spanger Hi Sara! Can't agree more. And we felt dashboard fatigue with our first launch as well. Dashboard can only tell you much, why not just go straight to the answers? Glad you like the concept:)

When you connect to QuickBooks, Veltrix will instantly show you what's happening right now. If you need any historic data or trends - just ask a question. It has a power of directly querying your QuickBooks data, so there is no limit on how far it can go.

And yes! We're already thinking on solutions for Veltrix on the go. Will be announced soon!

6
回复

Congrats on the launch!

6
回复

@oksana_ch Thank you, Oksana!

1
回复

Congrats! Interesting tool. How does it handle multiple currencies or international accounts? Will it consolidate cash flow across regions automatically?

6
回复

@dima_kulaksyz Thank you, Dima! It all depends on what setup you have in your sources. Veltrix will use and analyze everything you have connected consolidating multiple currencies and accounts. We do recommend to have one workspace dedicated to one location or business, and switch between others.

1
回复

This is useful. Does Veltrix flag discrepancies between your books and what your tools report?

6
回复

@dhiraj_patel5 Yes, exactly! Thank you for the support, Dhiraj!

4
回复

Love the combination of AI recommendations and actionable next steps, it feels like a real co-pilot for decision-making.

6
回复

@marianna_tymchuk Thank you, Marianna! Glad you like it!

3
回复
Nice, good project!
1
回复

Financial copilots seem most useful when they explain why a recommendation was made rather than just producing numbers.

Have you found users spending more time in forecasting and planning workflows, or using Vetrix primarily for real-time visibility into business performance?

1
回复

Congrats on the launch! How different your product is from asking Claude (with its library of connectors and MCPs) to do finance stuff?

1
回复

@nikitaeverywhere Thank you, Nikita! That is a great question. What I would say is different:
- We're using a set of different models, each for the things they do best.
- We're proactively surfacing your data even before you ask or know what to ask to show you what's going on, what needs your attention and what to do about it.
- We built it specifically for small businesses: it takes a minute to set up, you don't need to know what is MCP, and you just focus on what matters most - your financials.
- We have specific systems in place to accurately save, ingest, query and analyze any data source in real time.
- Lastly, you can connect all your sources in your private workspace, including the integrations Claude currently doesn't support.

Let me know if you have any other questions! Thank you for the support!

1
回复

Congrats on the launch!

I think cashflow analysis is crucial especially with current rush to overspend on AI tokens!

0
回复
So it actually builds dashboards but very lightweight and customised ?
0
回复
#5
Ideogram 4.0
Generate design-ready image with open weight, layout control
224
一句话介绍:Ideogram 4.0 是一款开源权重、支持边界框布局控制与多语言文字渲染的文生图模型,专为需要精准排版和品牌视觉的设计场景而生,解决了设计类图像生成中布局失控与文字错乱的核心痛点。
Design Tools Open Source Social Media GitHub
文生图模型 设计生成 布局控制 边界框 文字渲染 多语言 品牌视觉 自托管 API 开源权重
用户评论摘要:用户普遍关注排版与布局一致性,询问在多品牌视觉中能否保持统一布局;有用户指出视频暗示的可编辑文本功能未实装;有用户关心复杂多文本区域的边界框忠实度;其他反馈集中于长文本生成和皮肤处理效果出色,祝贺发布。
AI 锐评

Ideogram 4.0 的核心卖点在于“设计可用性”,而非“通用美观度”。它通过训练时就引入结构化JSON标注(边界框+元素描述),从根本上让模型理解空间结构,而不是像Stable Diffusion等开源模型那样靠文本推测布局。这直接命中了当前视觉AI在生产环境中的死穴:设计师不可能接受一个“生成的图看起来很漂亮,但文字是乱码、logo位置不对”的玩具。

从技术角度看,边界框布局控制、六进制颜色调色板条件化、原生2K输出、以及多语言文字渲染的加持,使其在品牌物料、海报、PPT封面、包装设计等细分场景中,具备替代闭源工具(如Midjourney+后期PS)的潜力。尤其是自托管的开源权重+商用许可,对于有数据隐私和定制需求的企业开发商而言,是一枚关键的“信任状”。

但需要泼点冷水:评论中用户提到的“可编辑文本未实装”是一个重大功能缺失,用户期待的是生成后能局部修改文字内容(类似Canva的编辑能力),而非只能通过JSON重新生成立即。这意味着它在设计工作流中仍是一个“初稿生成器”,而非“终稿编辑器”。此外,边界框的忠实度被标为“接近但非像素级”,这在高精度排版(如密集的多文字海报)时仍可能产生偏移,降低直发率。

总体而言,Ideogram 4.0 是一款务实且有竞争力的工程化产品,填补了开源模型在专业设计领域的空白,但要真正成为设计工具的标配,它还需要补齐编辑交互能力和更高精度的布局控制。

查看原始信息
Ideogram 4.0
Ideogram 4.0 is an open-weight text-to-image model trained from scratch, with bounding-box layout control, multilingual text rendering, and native 2K output. For developers and enterprises building on visual AI.

Ideogram 4.0 is an open-weight text-to-image model trained from scratch on structured JSON captions, built specifically for design-oriented output including typography, logos, posters, and brand visuals.

Proprietary models have held the lead on layout fidelity and accurate text rendering. Open alternatives have been usable for general photorealism but fall apart when a design needs copy to land in the right place, in the right font, at the right size. Ideogram solves this at the training level, pairing bounding-box coordinates with per-element descriptions so the model learns spatial structure rather than guessing at it.

Here is what that translates to in practice:

  • Explicit bounding-box layout control via JSON prompts, so every text region and object lands where the brief says it should

  • Multilingual text rendering across signage, logos, and multi-line typographic layouts, at native 2K resolution

  • Hex color palette conditioning for brand color control directly in the prompt

  • Self-hostable with fine-tuning support on proprietary data, and a commercial license that scales by deployment size

  • Hosted API access from $0.03/image with no subscription required

If you are an ML engineer evaluating open-weight image models for a production pipeline, or a creative technologist who needs design output that actually handles typography without manual cleanup, this is worth a serious look.

Download the weights on HuggingFace or try the model live at ideogram.ai.

P.S. I hunt the latest and greatest launches in tech, SaaS and AI, follow to be notified @rohanrecommends

6
回复

@rohanrecommends Appreciate the hunt!

0
回复

The typography and layout control part is the most interesting to me. Most image models still need a lot of cleanup when there is real text in the design. Can Ideogram keep the same layout consistent across several brand visuals?

5
回复

@ada_johnsen (I work at Ideogram) Yes! And btw: a good workflow is to generate something with natural language first to get some layouts you like (the natural language prompts get turned into structured json), then use the power of the structured json if you want to tweak copy/layout

1
回复

Been a fan of the free version for image generation, so it's great to see how much Ideogram is evolving. Layout precision and typography, which are still pain points for a lot of image models, are great addition. Good luck with the launch!

4
回复

@rohanrecommends Training on bounding-box coordinates rather than letting the model guess at layout is really nice. But the hex palette conditioning is a great feature too. Brand colour control in the prompt is something I'm seeing as a real requirement in today's branding. Question on the bounding box - how tight does the bounding-box adherence hold when you push it. Does a dense, multi-text-region layout stay faithful, or does it start drifting the way the proprietary models do past a certain complexity?

3
回复

@rohanrecommends  @tom_palmer_ux (I work at Ideogram) Bounding box is not pixel perfect adherence, but it is pretty close

1
回复

I was excited to try out editable text but it's not included in this release despite the video implying it is.

1
回复

Looks honestly amazing! Congrats on the launch. We're launching today as well. Good luck to us!

1
回复

Great benchmark performance! I'll share it with our development team. Congrats on the launch. We're also launching today, so good luck to all of us!

1
回复

Congratulations on your launch. Been using this model and this is really great in long text and human skin. Thank you for the early access!

0
回复
#6
Agent Mode on Arena
Get real-world tasks done with autonomous AI agents
166
一句话介绍:Agent Mode on Arena通过单一提示词驱动自主AI代理,在沙盒环境中执行浏览、研究、编码等多步骤真实任务,解决传统AI基准测试脱离实际、模型评估不透明、用户需频繁切换工具的痛点。
Productivity Artificial Intelligence
自主AI代理 多步骤工作流 AI模型基准测试 沙盒环境 排行榜 真实任务评估 编码代理 AI透明度 行为信号 数据
用户评论摘要:用户赞赏其提供实用的模型对比评估,取代营销炒作。关键问题包括:如何防止代理执行破坏性操作?反馈建议:用户发现代理在决策与规划环节易有“虚报完成”的缺陷,执行端到端任务时需人工收紧控制;期待GitHub集成、全栈编码、图片视频生成等工具扩展。
AI 锐评

Agent Mode的聪明之处在于,它把“基准测试从实验室搬到了战场”。传统AI评估是一场闭卷考试,而它改成了开卷做项目——这直接戳中了业内最大的虚假繁荣:模型在排行榜上屠榜,但干活时却像个实习生。通过让用户亲自跑真实任务并贡献数据,它既获得了比任何学术基准都更“脏”但更真实的反馈,又用排行榜的荣誉感形成了用户贡献的飞轮。

但它的命门也正在于此:沙盒环境屏蔽了破坏风险,却也限制了任务的“破坏性”,比如真正的代码审计、安全渗透测试无法被评估。此外,“行为信号”如“确认成功”和“bash恢复”是好的起点,但用户评论已明确指出“虚报完成”的棘手问题——这恰恰是代理走向实用最大的敌人。如果排行榜无法有效惩罚这种“优雅撒谎”,它就会沦为另一个观赏性体育赛事。

说白了,这个产品的价值不在于它展示模型有多强,而在于它把这个难堪的真相摆到了台面上。对于AI builder而言,它提供了比任何营销话术都可靠的决策依据。接下来,它需要证明自己不是又一个乌托邦,而是能真正容忍并量化代理“犯蠢”的残酷试炼场。

查看原始信息
Agent Mode on Arena
Most AI benchmarks test models in controlled environments. Agent Mode tests them on complex tasks to get more work done. Run autonomous agents that browse, research, code, use files, and complete multi-step workflows from a single prompt. Then watch each workflow unfold step by step. Every run contributes to the Agent Arena Leaderboard, ranking frontier models by real-world agentic performance.

One thing I appreciate about Arena is that it shifts the conversation from "which model is trending" to "which model actually performs best for my use case." With the pace of AI innovation today, having a reliable way to evaluate and compare models is incredibly valuable. This feels like a product that can help builders make smarter decisions instead of relying on assumptions or marketing.

Congratulations on the launch — excited to see how the platform evolves and serves the AI community! 🚀

14
回复

@1mirul Yes! It's all about how the models perform for actual real-world use cases. Appreciate the well wishes, excited to get this out to our community.

0
回复

Arena feels like a much-needed reality check for the AI space. Instead of guessing or trusting scattered benchmarks, it brings everything into one place where models can be evaluated side by side in a practical way. For anyone building with AI, this kind of clarity is extremely valuable. Excited to see how it grows and how the community contributes to making AI evaluation more transparent and useful over time.

11
回复

@monir_ Really appreciate the thoughtful comment! We're very excited to bring Agent Mode, and the Agent Arena leaderboard to the public to help better measure agentic AI!

0
回复

👋 Hey Product Hunt! We're excited to launch Agent Mode on Arena.


AI chat experiences are often limited to rigid, single-modality interactions that require switching tools or

additional prompting. Agent Mode changes that. You can now prompt once and the agent will plan, browse,

research, and code in a sandbox testing environment to complete real-world, multi-step tasks for you.


Every Agent Mode session also powers our new Agent Leaderboard, built entirely from behavioral signals (such as confirmed success, bash recovery, steerability, and more) collected from real users running real-world workflows. We’re excited to have our community contributing to the leaderboard, and provide a new standard for measuring AI advancement.


We'd love your feedback: What agentic tasks did you throw at it? What tools should we add next? Thanks for checking it out 🙏

11
回复

@elliott_gluck let's gooooooooo

0
回复

Really interesting. How does Arena prevent agents from doing something destructive during a benchmark run?

10
回复

@dhiraj_patel5 great question – the agent's actions are limited to a sandbox environment at the moment.

1
回复

the UI alone is pretty awesome. you guy really have that "taste", Elliott

I just tested it, and it's mind-blowing

3
回复

@nathan_tran2 Thank you Nathan, super exciting to hear you're already getting value and loving using the product!!

0
回复

@nathan_tran2 very kind of you, thank you!

0
回复
Really impressed by how Agent Mode cuts through the usual friction—prompt once and it actually carries the task end-to-end. The leaderboard angle is clever too, feels like a transparent way to measure real progress. Curious to see what new tools you’ll plug in next.
2
回复

@odeth_negapatan1 next up is: Github integration, full stack coding, slide creation, pdf creation, image edit, video gen, and wayyy more. Strap in 🤘

1
回复

One thing we've noticed with agent workflows is that execution is becoming less of a bottleneck than decision-making.

Are you seeing users struggle more with planning and prioritization, or with agents actually completing the tasks once they're started?

1
回复

@zaid_mallik1 we see a couple things
1) Most users start messages by handing over a whole job rather than asking for advice: the delegation posture skews heavily toward "build this deliverable" and "operate autonomously." However, after seeing the first response, they tighten the reins — pulling control back far more often than they hand over more.
2) We also find that when the opening ask bundles several explicit parts, agents usually cover all of them; the typical shortfall is leaving one incomplete. A rarer but more consequential shortfall is covert: the agent could have surfaced the incomplete work, but instead presents the result as complete. We call this "Bluffing".

More info here in our blog about the Agent Leaderboard if you're curious: https://arena.ai/blog/agent-arena-methodology/

2
回复

@elliott_gluck Congratulations. And happy product launch.

0
回复

@huisong_li Thank you @huisong_li , appreciate your support!!

0
回复
#7
Nemotron 3 Ultra by NVIDIA
Powers faster, efficient reasoning for long-running agents
152
一句话介绍:Nemotron 3 Ultra 是 NVIDIA 专为长时间运行的 AI 智能体(Agent)设计的开源大模型,通过混合架构解决长会话场景下推理成本高、上下文丢失的痛点,让智能体在执行复杂任务(如编程、深度研究)时更快速、更便宜。
Developer Tools Artificial Intelligence
AI智能体 大语言模型 开源模型 MoE混合专家 长上下文 推理加速 量化技术 550B参数 NVIDIA Agent框架
用户评论摘要:用户普遍认可其550B参数(55B活跃)、1M上下文和300 tok/s的性能,认为它是目前最强的美国开源模型。核心讨论点在于:长上下文窗口是否削弱了对外部检索和记忆层的依赖?模型自身改进与外围检索/记忆系统,哪个对长周期智能体任务提升更大?
AI 锐评

NVIDIA 这波操作很聪明,但也很“现实”。Nemotron 3 Ultra 并非试图在单轮推理的基准上硬刚 O1,而是精准切入了“智能体”这个现实落地的痛点——长链条、多工具、易崩溃。550B MoE 配 1M 上下文加 5 倍推理加速,看似豪华,实则是为了解决 Token 成本爆炸和上下文遗忘这两个智能体的“送命”问题。其核心价值不在于“更强”,而在于“更持久、更经济”,这比单纯刷榜更具工程意义。

然而,吹毛求疵地说,评论中用户提出的问题直击要害:大上下文窗口并不意味着解决了一切。当模型能“记住”1M token时,瓶颈反而从“记忆容量”转移到了“信号检索”——智能体如何在庞大的历史中准确找到关键信息?NVIDIA 此次虽开源了模型和训练秘方,但并未提供配套的高效检索机制。如果用户仍然需要依靠外部向量数据库或复杂的记忆架构来“喂食”这个长上下文模型,那么其宣称的优化效果可能大打折扣。此外,“开源”的诚意值得商榷:虽然权重和数据公开,但 550B 模型(即使只有 55B 活跃)的部署门槛依然极高,算力成本只是从“天价”降到了“非常贵”,对中小开发者并不友好。Nemotron 3 Ultra 是 NVIDIA 在智能体赛道上的一次精准卡位,但要让其成为生态标准,还需要解决“如何让它高效工作”这一比模型本身更难的问题。

查看原始信息
Nemotron 3 Ultra by NVIDIA
A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models.Ultra excels at complex tasks like coding and deep research. Long-running agents spend their time planning, using tools, recovering from failures, and deciding what to do next.

NVIDIA just shipped Nemotron 3 Ultra, a 550B open frontier model purpose-built for long-running AI agents.

Most frontier reasoning models are optimised for single-turn accuracy. Agentic tasks are different: agents plan, call tools, delegate to sub-agents, handle failures, and pass history back into the model across many turns. As sessions get longer, token costs compound and models start losing the thread.

Nemotron 3 Ultra addresses this with a hybrid Mamba-Transformer architecture that handles long-context sequences without losing recall, and NVFP4 quantisation that delivers 5x higher throughput per GPU compared to BF16 on Blackwell.

Here's what ships:

  • 550B total / 55B active parameters via LatentMoE so you get frontier reasoning without activating the full model on every token

  • Up to 1M token context window handles large codebases, long tool-call chains, and multi-document synthesis natively

  • Multi-token prediction layers reduces generation time on long outputs and multi-turn workflows

  • Post-trained for OpenClaw, Hermes Agent, and LangChain Deep Agents accurate across agent harnesses, not just chat benchmarks

  • Multi-Teacher On-Policy Distillation trained with dense feedback from 10+ domain-specific teacher models across code, math, and tool use

  • Fully open weights, synthetic training data, and post-training recipes all released under OpenMDW-1.1

P.S. I hunt the latest and greatest launches in tech, SaaS and AI, follow to be notified @rohanrecommends

6
回复

550B params (55B active), 1M context, 300 tok/sec. probably the strongest US open-weights model out there right now - and it's currently available for free on @Kilo Code

ouch.

2
回复

Big release. What’s interesting to me is less the “bigger context window” headline and more what it means for actual agent runs, where most of the work is planning, tool calls, backtracking, and keeping state over time.

I’m curious how you’re seeing people use Nemotron 3 Ultra alongside retrieval or external memory. With a 1M context window, does that layer become less important, or does it just shift toward deciding what should live in memory vs what gets passed straight into the run?

0
回复

A lot of frontier models are improving raw reasoning, but context management still feels like a separate bottleneck.

Have you seen longer-horizon agent workloads benefit more from the model improvements themselves, or from better retrieval and memory layers around them?

0
回复
#8
Moodloom
Ad-free Pinterest Alternative with AI content filtering
120
一句话介绍:Moodloom 是一款无广告的 Pinterest 替代品,通过AI内容过滤功能,帮助时尚、家居和艺术爱好者摆脱广告和劣质AI图片的干扰,专注于真实的视觉灵感探索。
Art Beauty & Fashion Design
无广告社交平台 Pinterest替代 AI内容过滤 视觉发现 时尚灵感 家居装饰 AI图片识别 数字艺术 书签导入 兴趣社区
用户评论摘要:用户普遍赞赏无广告和Pinterest导入功能,但核心疑问集中在商业模式(如何盈利)和扩展成本上。建议细化AI过滤等级(如按类别展示AI内容),有用户指出产品自身在兴趣选择页使用了AI图片。
AI 锐评

Moodloom切中的痛点真实且迫切——Pinterest用户早已对广告、赞助帖和日益泛滥的AI“垃圾图”感到厌倦。其核心价值并不在于“又一个视觉发现平台”,而在于“通过技术手段主动帮用户过滤噪音”,尤其是AI内容检测功能,这在当前AI生成内容井喷的语境下,直击用户信任危机。然而,120票和寥寥数条深度评论暴露了产品仍处于极早期阶段。创始人声称通过电商而非广告变现,逻辑自洽——用户对商品的真实图片需求天然排斥AI伪造品,但这也对供应链和内容生态提出了极高要求:如何确保非AI的优质创作者持续入驻?如何平衡AI过滤的误伤率(例如手绘、CG艺术被误判)?评论中用户发现产品自身在兴趣选择页用了AI图,这是致命伤:一个宣称“过滤AI”的初创平台若内部审核都不严,会迅速瓦解最核心的信任资产。此外,Pinterest的卖点之一是“发现”,而Moodloom的AI过滤本质是一种“筛选”,两者导向的用户行为不同——前者是信息量驱动的浏览,后者是精准度驱动的搜索。如果Moodloom只做“更干净的Pinterest”,而缺乏独特的算法推荐机制或社区互动生态,那么其替代性依然脆弱。说到底,AI过滤不是护城河,只是入场券;真正的战场在于:在去除了“噪音”后,产品能否提供比Pinterest更独特、更高效的“信号”。目前来看,Moodloom有犀利的矛,却尚未证明自己有能扛住规模的盾。

查看原始信息
Moodloom
Moodloom is an ad-free visual discovery platform for those fed up of Pinterest. Save and explore fashion, decor, and art - without the noise. What makes it different: Clean feed - no ads, no sponsored posts AI content filter - hide or show AI-generated images For fashion lovers, decorators & digital artists who wish Pinterest were better. We have built Save to Moodloom Extension as well using which you can Import all your Pinterest pins/boards to Moodloom.

I think we need more ad-free alternatives to other websites, too. The thing is – how do you want to make money from it (at least to cover costs)?

6
回复

@busmark_w_nika I loved the idea of this product, too. I'm tired of seeing AI slop on Pinterest when I'm looking for inspiration for a project, but the same question popped in my mind, too. Apart from the business making money. But the In-app shop sounds like good for the users of this product to monetize content.

0
回复

@busmark_w_nika we are planning to monetize through commerce, not ads so the incentives stay aligned with users

0
回复
Hey Product Hunt 👋 I'm Manas, the founder of Moodloom. I built Moodloom because Pinterest stopped feeling like a place for genuine inspiration. The feeds are full of ads, sponsored posts, and increasingly-AI-generated images mixed in with no way to tell the difference. As someone who loves collecting visual inspiration for fashion and design, I wanted a platform that actually respects your taste. No noise, no manipulation - just images you care about. So we built Moodloom differently. You can filter AI-generated content entirely, discover images without noise, and shop directly from things you save - all without a single ad in sight. We're still early. AI filtering is live but we're actively making it better. In-app shopping is on the roadmap and coming soon. There's a lot more we want to build - better personalisation, smarter discovery, and deeper category exploration. This is day one for us and Product Hunt feels like the right place to start. Would love your feedback - what resonates, what's missing, and what you'd want to see next 🙌
4
回复

@manas_khandelwal2 Idea seems cool! I have been using Pinterest for years and been affecting the search experience. However, What's your thought about scaling Moodloom? isn't it going to be expensive when it comes to handling a massive data?

0
回复

The Pinterest import feature was what convinced me to try it. Being able to bring existing boards over without starting from scratch removes a huge barrier for switching.

4
回复

@ishika_meena That was exactly the thinking behind it! Switching costs are the biggest reason people stay on Pinterest even when they're frustrated. Really glad it worked for you, hope you enjoy Moodloom!

0
回复

Really like the clean-feed direction and your web design.

One idea I thought could be more control over the AI content filter. For example, instead of only "hide AI-generated images," maybe users could choose different levels like "hide all AI", "show only cleary labeled AI," or "allow AI in certain categories."

3
回复

@evakk Thanks Evak! Love the idea of granular levels. The "show only clearly labeled AI" tier is tricky in practice though because it relies on creators self-reporting, and a lot of AI-generated content gets posted without any disclosure. That's actually why we went the detection route instead of trusting tags. Category-level filtering is something we've thought about and could definitely see on the roadmap. Appreciate the thoughtful feedback!

0
回复

One more thing, I just signed in on the website and I noticed that AI images were used when I was asked to select interests, have you checked into that?

1
回复

@akin_kevin Thanks for pointing this out, Kevin. When we set up onboarding, we missed that a few interest thumbnails were AI images. I've checked and only a handful are affected. We're swapping them out for non-AI ones.

0
回复
#9
LocalClicky
Control your Mac with your voice locally
119
一句话介绍:LocalClicky是一款完全离线的Mac菜单栏语音助手,通过本地运行语音识别和LLM模型,实现无需联网、不泄露数据的自然对话式电脑操控,解决了用户对隐私和云端依赖的核心痛点。
Open Source GitHub Tech Audio
语音控制 Mac工具 离线AI 隐私保护 开源 本地大模型 Whisper Ollama 菜单栏应用 语音交互
用户评论摘要:用户关注安全和确认机制,开发者承认尚无防护措施,仅依赖重试反馈。低配置Mac延迟较明显,M1机型约3-5秒。离线唤醒词是当前最大短板,依赖Google语音识别,正计划替换。本地架构对敏感场景(如养老)极具吸引力。
AI 锐评

LocalClicky的价值不在技术突破,而在“离线即隐私”这一极端承诺的真实落地。它用Whisper+Ollama+macOS say的本地堆栈,彻底切断了语音助手行业默认的数据外流通道,这对隐私敏感用户(开发者、医疗、法律)是刚需。但产品远未成熟:无任何确认机制直接操作Mac,这意味着一次误识别就能引发灾难;延迟在非顶级硬件上几乎摧毁“对话感”;最讽刺的是,唯一的云端依赖恰恰是那个最基础的“唤醒词”——Google Speech API。开发者承认“没有防护”,这种裸奔式的代理授权,在生产力场景中无异于引火烧身。MIT开源是个好起点,但真正的问题不在于模型跑得快不快,而在于用户敢不敢让它点击“删除”按钮。如果接下来不构建完善的权限分级、沙箱执行和关键操作确认机制,LocalClicky只能停留在极客玩具层面,而不是可信的生产力工具。

查看原始信息
LocalClicky
LocalClicky is a Mac menubar app that lets you have a real conversation with your computer - completely offline. Say "Computer" to start a session. It stays listening. You chain commands back to back. Say "goodbye" when you're done. Everything runs on your machine: voice transcription, LLM multi models, VAD, macOS say No API keys. No subscription. No data leaving your Mac. MIT licensed.
When LocalClicky decides it needs to “see the screen,” how do you keep that reliable and safe in practice—what guardrails exist to prevent mis-clicks or destructive actions, and how do you handle confirmation/retries in real workflows?
2
回复

@curiouskitty Good question, Tbh there's no guardrails as such being implemented yet. There's no confirmation step before clicking. We have retries for tool calls that uses shell commands. So it runs, checks the output and will try a different approach if it get non desired output. Retries several times then provide feedback if still not able to perform the action.

0
回复
Hey PH! I built LocalClicky because every voice assistant I tried made the same tradeoff: you get convenience, they get your data. Audio uploaded. Screenshots sent to a server. Commands logged. LocalClicky flips that. The whole stack is local, Whisper for transcription, Ollama (qwen3 + gemma4) for reasoning and vision, macOS say for responses. Nothing phones home. Not your voice, not your screen, not your commands. The session model is what makes it feel natural, you say "Computer" once and it stays with you (Just like Siri on steroids). Chain commands, ask follow up questions, say "bye" to end. VAD stops recording as soon as you stop talking, so there's no fixed timeout awkwardness. It's open source and early. The top contributor priority right now is replacing the Google Speech Recognition wake word with something fully offline. GitHub in the comments, happy to answer anything. Give it a try !!!
1
回复

The fully local stack is the part that resonates with me. I work on voice AI for aging-in-place, and privacy comes up in nearly every conversation with families, because the data involved (health, daily routines, who stopped by) is about as sensitive as it gets. Keeping Whisper and the LLM on-device the way you have here removes a whole category of worry. One thing I am curious about: on older or lower-spec Macs, what is the realistic gap from end of speech to spoken response with the local models, and which model sizes stay usable? That latency is usually where local voice either feels like a real conversation or starts to break it.

1
回复

@igorgurovich Oh, thats a solid use case.

To answer your question "on older or lower-spec Macs", the gap is noticeable, I haven't tested myself but on M1 with 16 gb ram with nothing else running, the latency should be under 3-5s for a normal command, but for complex commands like generating some health report, it can take a while. Yes latency has been a concern but I was earlier looking into possibility of distributed computing just like how torrent works, offloading compute, but couldn't make much progress there :)

1
回复

Love the local-first approach here. The session model also feels like the right UX for voice control, much closer to an actual conversation than one-shot commands.

Curious how you’re thinking about the safety layer before clicks or destructive actions, especially since everything can see/control the Mac locally. Is the plan to add confirmations for certain action types, or keep it lightweight and rely on retries/feedback?

0
回复

the whisper + ollama stack being fully local makes the google wake-word the last phone-home piece — funny that the smallest model in the chain is the hardest one to cut the cord on. what's actually blocking the offline swap?

0
回复
#10
FloatPic
Ultra-minimalist, borderless macOS native image viewer
113
一句话介绍:FloatPic是一款极简无边框的macOS原生图片查看器,让参考图片“悬浮”在工作区上方不抢占视觉空间,解决创作者、设计师、摄影师在整理参考素材时被厚重边框和窗口干扰的痛点。
Mac Design Tools Productivity
macOS图片查看器 极简设计 无边框悬浮窗 图像工具 EXIF查看 取色器 OCR识别 直方图 图片对比 快速浏览 30+格式
用户评论摘要:用户关注多桌面支持(已实现)与多窗口并排需求(当前为单图替换模式,未来考虑多面板)。专业用户对一键调取EXIF、直方图、取色器等工具表示认可,尤其色彩提取功能意外成为区分点,被设计师和开发者视为全天候参考工具。
AI 锐评

FloatPic抓住了“消失”这一设计哲学,但这更像是对macOS原生图片预览器“预览.app”的精致魔改。它的核心价值并非在“看图”本身——30+格式兼容性已是当今基础素养,而是在于把图片从“文件浏览”升维为“工作台悬浮参考”。评论区验证了这一点:用户不想要又一个看图软件,他们需要的是一个能吸附在代码、设计稿、照片库上方的“即时拾取工具集”。它聪明地避开了与LunaSea、Pixa等重型素材管理器的正面竞争,专注于“单张图片的全屏瞬间交互”。

但潜在的短板也很明显:单窗口限制会立刻劝退需要多图叠放对比的用户(如UI设计师对照多个界面稿),虽然他们自圆其说“做好一个窗口更可靠”,但这实质是对多窗口管理复杂度的技术妥协。另一个隐忧是脱离Finder的独立文件管理路径——如果只能通过文件选择器或拖拽打开,面对大量图片的快速切换效率远不如Bridge或Eagle那种带缩略图网格的管理模式。它目前更像是“极客的图片瑞士军刀”,而非普通用户的生产力工具。OCR与颜色提取虽惊艳,但若缺乏后续的文本编辑、颜色方案导出(如Sketch/Figma插件联动),则容易停留在“玩具级”功能。如果FloatPic坚持不做搜索、标签与库管理,其用户天花板便是那些需要高频参考、但不喜欢维护图片库的创意工作者——这个群体确实存在,但规模有限。

查看原始信息
FloatPic
FloatPic is the ultra-minimalist macOS native image viewer. Borderless floating window, native gestures, blazing-fast loading, supports 30+ image formats. Make the software disappear, let images float.

Hi Product Hunt! 👋

I built FloatPic because I was tired of bulky image viewers that clutter my screen with borders and windows. As a creator/developer, I just wanted my reference images to "float" cleanly on top of my workspace without getting in the way.

FloatPic is designed to disappear. It’s an ultra-minimalist, macOS-native image viewer that supports over 30+ formats, responds to native gestures instantly, and runs blazing fast.

I’d love to hear your thoughts! What features would make your workflow even smoother? Thank you for the support! 🚀

🎁 LAUNCH DAY SPECIAL:
Thank you for checking out FloatPic! Leave a comment below and DM me on X (Twitter) @tapfunapp to claim a 1-month Premium promo code as a thank-you gift! Let me know what you think! 👇

0
回复
You combined ultra-minimal UI with pro inspection features (EXIF groups, OCR, histogram, color picker/palette, compare modes). How did you decide that scope, and which feature surprised you as the biggest differentiator once real users tried it?
0
回复

@curiouskitty 

Thanks for the thoughtful question! 🐱

The philosophy was simple: the viewer should disappear, and you're just looking at your image — but when you need to dig deeper, one keystroke brings up any pro tool instantly. No menus, no panels, no chrome. Just the image.

Scope decisions were driven by our own daily workflows — EXIF and histogram for photography, OCR and color tools for design/development work. Every feature has a single-key shortcut: H for histogram, ⌘I for EXIF, T for OCR, C for color picker, P for palette, W for compare. They appear when you need them and vanish when you don't.

The surprise hit? Color palette extraction. We expected EXIF to be the main "pro" draw, but designers and devs consistently tell us they love hitting P on any image and getting a usable 6-color palette they can copy in HEX/RGB/HSL. It turned FloatPic from "nice viewer" into something people keep open all day as a reference tool.

0
回复

cuteeee

0
回复

@madalina_barbu Thank you for the kind words! Highly appreciate the support.

0
回复

The borderless, floating approach is interesting for reference work, where you want an image pinned on screen while you work in something else without it fighting for visual space. Curious whether it respects macOS Spaces or stays on all desktops, and whether you can pin multiple images at once without them stacking awkwardly. That second part is usually where "minimal" viewers quietly fall apart.

0
回复

@fberrez1 
Great observations — those are exactly the edge cases we focused on.

Spaces: FloatPic panels use canJoinAllSpaces, so each floating image appears on every desktop/Space automatically. Combined with the optional floating-on-top level (togglable in Settings), the image genuinely stays pinned above your work regardless of which Space you're on.

Multiple images: Currently FloatPic uses a single-panel approach — opening a new image replaces the one in the existing floating window, and you navigate between images in the same directory with swipe or arrow keys. This was an intentional tradeoff: for the "pinned reference image" workflow, one clean floating frame is usually what you actually want. That said, we're actively considering multi-panel support for scenarios where you need two images visible side-by-side (and the compare mode already handles that for images in the same folder).

Fair point on where minimal viewers fall apart — we'd rather do one window really well than half-support a feature that creates the stacking chaos you described. Appreciate the thoughtful feedback!

0
回复
#11
Agent Browser Shield
Block prompt inject & cut token costs for AI browser agents
103
一句话介绍:Agent Browser Shield 是一款开源的浏览器扩展,通过在AI代理与网页之间设置过滤层,精准拦截提示注入、屏蔽PII、移除暗黑模式并滤除页面噪音(如Cookie横幅),从而解决AI代理在浏览网页时易被恶意指令操控及无谓消耗Token成本的痛点。
Browser Extensions Open Source Artificial Intelligence GitHub
AI安全 提示注入防护 浏览器代理 Token优化 开源工具 PII脱敏 暗黑模式过滤 网页内容净化 浏览器扩展 AI代理工具
用户评论摘要:用户点赞Token噪音是代理失败的主因,比安全威胁更常见;建议增加对动态注入元素(如弹窗、聊天组件)的DOM变化监听。部分用户询问是否提供过滤差异追踪(diff)以调试误报;还有用户建议在输出控制台显示diff信息。另有评论认为,静态规则可能难应对攻击演进,呼吁自适应模式。
AI 锐评

Agent Browser Shield 切中了一个当前AI代理应用中的“房间里的大象”——当AI开始像人类一样“阅读”网页时,我们对其安全性和效率的想象还停留在理想状态。产品巧妙地将“对抗提示注入”与“降低Token消耗”打包为同一解决方案,这并非巧合,而是一个深刻的洞察:恶意指令伪装成正常内容,而正常内容中的垃圾信息(Cookie弹窗、页脚导航)本质上也是一种“噪音污染”。两者都指向一个核心问题——AI缺乏比人类更敏锐的内容辨别力。

该产品的真正价值不在于“拦截”,而在于“选择”。它本质上是一个为AI代理设计的“注意力管理器”,它强制让模型只看到经过筛选的、高信噪比的页面核心内容。这种“输入净化”思路,比在模型端增加安全微调或事后审计更主动、更低成本。尤其是“PII脱敏”和“暗黑模式移除”的加入,让产品从单一安全工具升级为“AI上网合规与效率一体机”。

然而,产品的关键挑战在于“规则静态性”与“攻击动态性”之间的永恒赛跑。当前基于规则的方法(如CSS选择器、文本模式匹配)对已知模式有效,但面对利用DOM动态注入或变种混淆的复杂攻击,很容易被绕过。更棘手的“上下文污染”——攻击者通过合法内容渐进式诱导模型——规则系统基本无能为力。产品提出的“自适应模式”和“外部API辅助分析”是正确方向,但如何平衡边缘计算的性能与云端分析的隐私,对轻量级浏览器扩展是个考验。

另外,用户对“diff差异追踪”的呼声非常高。这恰恰指出了当下AI可观测性的缺失:开发者需要知道模型到底在看什么。Agent Browser Shield如果能不仅过滤,还能提供“过滤前后的内容对比报告”,它将从一个“屏障”进化为“审计日志”,这对企业级部署中的合规审计和故障根因分析具有不可替代的价值。一句话:它是AI代理必备的“浏览器安全套”,但想长成生态系统,必须补上“可观测性”这块拼图。

查看原始信息
Agent Browser Shield
AI agents browsing the web have a problem: they read everything — cookie banners, hidden instructions, dark patterns — and can't tell real content from a trap. Agent Browser Shield sits between your agent and the web, stripping prompt injections, masking PII, removing dark patterns, and filtering page noise that burns tokens. Free, source-available, works with browser-use and Browserbase.
Hey PH! 👋 Britt from PixieBrix here. We've been building browser tooling for enterprise teams for years, and when we started seeing AI agents get deployed at scale to browse the web, we noticed a gap: agents have zero protection from the stuff (most) humans have learned to watch out for. Prompt injection is OWASP's #1 AI security threat — and it's trivially easy to embed hidden instructions in a webpage that your agent will follow without question. On top of that, most pages are full of junk (cookie banners, footers, chat widgets) that your agent reads and pays for in tokens. Agent Browser Shield is our answer: a free, source-available browser extension that strips all of that before the model sees the page. We built it in the open because this is an evolving problem — new dark patterns, new injection techniques — and no single team can keep up alone. Would love your feedback on what to build next. And if you're running browser agents in production, let us know what failure modes you've hit because we want to keep building and write better rules to make this even better. GitHub: https://github.com/pixiebrix/age...
2
回复

The token noise problem is honestly what gets me more than the injection side. Most agent failures I've seen aren't dramatic security breaches, they're just the model losing track of what matters after reading 3 cookie banners and a footer nav in a row.

Quick question though, does the filter run once on initial page load or does it watch for DOM mutations? Asking because a lot of the annoying stuff (cookie popups, chat widgets) gets injected dynamically after the page loads.

2
回复

@imoluuu great question! It watches on page load and for mutations on the page.

We’re adding an option for controlling whether it watches on inactive tabs (in case you’re using chat on pages that aren’t currently visible, but could sneakily swap out content when you’re not looking).

0
回复
If someone is already using a token-efficient agent browser (e.g., DOM-diff / structured extraction approaches) or a containment-focused platform approach, where does Agent Browser Shield still add unique value—and where do you explicitly *not* try to compete?
2
回复

@curiouskitty Great question! For us, token efficiency is a secondary benefit. Our primary focus is on prompt injection, dark patterns, and context pollution that can compromise the agent or make them fail at the task.

Because it's a normal browser extension, you can use it inside your token-efficient agent browser

1
回复

The prompt-injection angle is obviously important, but the token-noise part is what stood out to me. When using browser agents for research/social workflows, the annoying failure mode is often less "one malicious instruction" and more the agent wasting context on cookie banners, nav, footers, and hidden junk before it reaches the actual page content.

Curious if Agent Browser Shield exposes a diff or trace of what it stripped from the page. That would be really useful for debugging false positives, especially when a page has weird layout or important content that looks like boilerplate.

Congrats on the launch — this feels like one of those unglamorous layers that becomes necessary once browser agents move from demos to real workflows.

2
回复

@grace_lee26 absolutely - the security is important but the token saving is a super sweet benefit.

That’s a fantastic point about the diff. We don’t have that at the moment but that would be a great feature. Would you expect to see that in console logs or where would you want that output of the diff sent?

Thanks for the support!

0
回复

PII masking at the proxy layer is smart - agents leak more than most realize. Is the injection detection rules-based or does it adapt as attack patterns evolve? Static filters tend to fall behind pretty quickly in this space.

0
回复

@christian_knaut Currently rules-based. We're working on adaptive patterns, but to determine what should be run on device (which might be a low-resource VM) vs. supporting external injection detection APIs

0
回复
#12
VisionSync
Where strategy execution meets the people doing the work
101
一句话介绍:VisionSync将战略规划与执行人员实时连接,解决组织战略“制定后无人跟进、问题发现滞后”的痛点,让领导者能提前预警执行偏差,而非事后补救。
Human Resources Business Intelligence Change Management
战略执行 项目跟踪 实时状态 团队绩效 KPI 企业管理 协同工具 360度评估 教育科技 战略对齐
用户评论摘要:用户普遍认同战略制定与执行之间存在巨大鸿沟。创始人点出核心痛点:计划沦为无人问津的PDF。竞争者表达共鸣与祝贺,评论区未出现反面意见或具体功能建议,整体以欢呼和支持为主。
AI 锐评

VisionSync切中的确实是企业管理中的“最后一公里”——战略落地。从产品形态看,它并非颠覆性创新,而是将原本分散在项目管理、绩效管理、战略文档中的功能(状态更新、责任人绑定、实时看板)做了一次精准的整合。真正值得关注的是其获客场景:从医疗中心和大学系统切入,这类组织的典型特征是“决策层与执行层脱节严重,且跨部门协作依赖行政指令”,VisionSync提供的“实时可见性”在官僚层级复杂的组织中价值极高。然而,产品护城河并不深。一旦主流项目管理工具(如Asana、Jira)强化战略对齐与高层汇总看板,或者老牌绩效管理系统(如BetterWorks、Quantive)嵌入更轻量的执行追踪,VisionSync的差异化将快速收窄。其最大挑战在于:如何从“信息同步工具”进化成“执行分析引擎”——仅仅是告诉大家“谁在做什么、进度如何”还远远不够,必须利用数据识别出“为什么停滞、如何干预”,才算真正做透“使战略发生”这个命题。目前101个投票和有限的企业用户验证尚不足以论成败,但方向对了。

查看原始信息
VisionSync
Existing tools help you write a strategic plan. VisionSync makes sure it actually happens. It links every goal to the people executing it and shows the live status of each initiative, so leaders see what's slipping before it hits quarterly results, not after. Add ValueSync 360 to connect strategy to performance on the people side. Proven across all four University of Nebraska campuses, Maryville University, and the Greater Omaha Chamber of Commerce.

Congrats on the launch! The gap between strategy creation and strategy execution is much larger than most organizations realize.

4
回复

@marianna_tymchuk Thank you!

0
回复
Hi Product Hunt, I'm Taylor, founder of VisionSync. Here's the problem we kept running into: organizations spend months building a strategic plan, then it lands in a PDF nobody opens again. Most tools help you write the plan. Almost none help you actually execute it. So the first sign something is off tends to be the quarterly results, which is far too late to fix it. VisionSync connects the strategy at the top to the people doing the work. Everyone sees the plan, the initiatives under it, who owns each one, and its real-time status, at any time. When something starts to slip, you catch it early. ValueSync 360 extends that to the people side, tying strategy to how teams are actually performing. It started at the University of Nebraska Medical Center and is now live across all four campuses of the NU system, plus Maryville University and the City of Omaha. As Dr. Jane Meza at UNMC describes it: "All I have to do is open up VisionSync and I can see the status of any of our programs or units at any given time." That live visibility is the whole point. I'll be here all day answering questions. What I'd most love to hear: how does your team keep a strategic plan alive after the kickoff meeting? That's the exact gap we built this to close. Thanks for taking a look.
3
回复

Wow! I already see it being a useful part of team management. Congrats on the launch. We're also launching today, so good luck to all of us!

3
回复

@veltrixms Thanks, and good luck to you!

0
回复
#13
Microsoft MAI-Voice-2
Expressive TTS with voice cloning in 15 languages
100
一句话介绍:微软MAI-Voice-2是一款支持15种语言情感表达与语音克隆的TTS模型,通过精细情感控制和短样本克隆技术,解决了语音代理在通话前8秒内“听出是机器人”的体验痛点,且定价远低于OpenAI Realtime API。
Productivity Developer Tools Artificial Intelligence
语音克隆 情感合成 多语言TTS 语音代理 Azure AI 音频生成 人机交互 微软AI 文本转语音
用户评论摘要:用户高度认可其在短通话中的情感表现和跨语言语音一致性,尤其适用于健康护理、移民长辈陪伴等场景。但多位用户质疑长对话(如10分钟)中韵律是否会漂移、声音身份是否稳定,并询问WebRTC实时流的延迟表现及模型回答复杂问题的基准能力。
AI 锐评

微软这次打出的牌很聪明:不是去卷大模型对话能力,而是卡准了“语音质感”这个关键缝隙。OpenAI Realtime API虽然多模态能力强,但TTS部分的韵律表现一直是用户吐槽的重灾区——高昂的成本并未解决“机器人感”的前8秒死亡时刻。MAI-Voice-2用$22/M chars的价格直接对标ElevenLabs并低于GPT Realtime TTS层,显然是冲着降级替代来的。

但要注意,这款产品的真正价值不在于“AI牛不牛”,而在于Azure AI Foundry这个平台织的网:VSCode、Dynamics 365、Teams三件套的集成,意味着它根本不是给独立开发者兴奋一下的玩具,而是为微软云生态内的企业客户提供“语音即服务”的标准化模块。企业买它的理由,不是因为它比OpenAI好,而是因为它“够用且不脱离架构”。

隐患也很明确。评论里用户对长对话稳定性提出的质疑不是挑剔,而是致命伤。如果一个号称production-grade的TTS产品,在10分钟的通话中就会从“情感饱满”滑向“中性冷淡”,那它本质上还只是一个demo级的短视频配音工具,而不是客服系统的核心组件。微软能否给出超过2分钟session的压力测试数据,决定了它到底是“语音代理的答案”还是“又一个漂亮的PR素材”。至于WebRTC低延迟和复杂问答的基准缺失,则是目前产品成熟度拼图上两个明确的窟窿——填不上,就出不了云。

查看原始信息
Microsoft MAI-Voice-2
Microsoft's most expressive TTS model yet — voice cloning from short samples, fine-grained emotional control, and consistent voice identity across 15 languages. Now live in Azure AI Foundry at $22 per million characters, with integrations rolling out in VSCode, Dynamics 365 Contact Center, and Teams. For builders shipping voice agents who need production-grade prosody without the OpenAI Realtime API price tag.
I build voice agents for service businesses — mostly healthcare and home services — and the #1 unsolved problem in this space is prosody. The "is this a robot?" moment usually happens in the first 8 seconds of a call. MAI-Voice-2 is the first TTS I've A/B tested where my pilot users couldn't tell. The $22/M chars pricing lands below ElevenLabs and matches gpt-realtime's TTS layer. If you're shipping voice and wedded to OpenAI Realtime, worth running the side-by-side. Curious if Microsoft is planning sub-200ms first-token latency via WebRTC streaming next.
1
回复

The consistent voice identity across 15 languages is what stands out to me here. I work on a voice companion that calls aging parents every day, and a lot of our families are immigrants whose parents are most at ease in their first language. A warm, familiar voice that holds up in Tagalog or Mandarin is often the difference between a call someone looks forward to and one they let ring out. Question for the team: how stable is the cloned identity and emotional control over a full 10-minute conversation, or does the prosody drift toward neutral as the session runs longer?

1
回复

@igorgurovich I love that application for aging parents, Igor. In my experience building out these workflows, the longer the call, the harder it is to avoid that 'robot moment.' Most of my pilot testing so far has been focused on shorter, highly qualified meeting booking flows where the emotional control performs beautifully. I actually haven't pushed a single session to a continuous 10 minutes yet, so I'm incredibly curious to hear the makers' answer on how the prosody holds up at that length. Tagging the team to chime in!

0
回复

Incredible that these voice models are becoming indistinguishable from real human voices. I was wondering if there are any benchmarks or detailed testing that was explored on the complexity of quesitons that the models can answer? This gap has been a major challenge for me to adopt AI voice agents that take on the role of customer support without assistance, but curious on how this is evolving.

0
回复
#14
Lumo Studios
Build Decks that Speak for Themselves
98
一句话介绍:Lumo Studios 是一款AI驱动的幻灯片制作工具,帮助用户将想法快速生成为精美、结构完整的演示文稿,省去排版和设计烦恼,适合职场汇报、品牌展示等场景。
Design Tools Productivity Artificial Intelligence
AI幻灯片生成 智能演示工具 品牌设计自动化 AI内容生成 Canvas编辑器 品牌指南导入 演示效率工具 团队协作 Deck制作 AI锐评
用户评论摘要:用户关注AI生成后的编辑自由度,特别是Canvas编辑器对布局、视觉样式的控制能力。也询问品牌指南导入方式及AI如何平衡自动化与品牌一致性。团队回应强调支持上传PDF、网站URL自动提取品牌资产,并提供风格统一与手动微调的双重控制。
AI 锐评

Lumo Studios 试图在 AI 生成的“快”与设计可控的“准”之间找到平衡,这正是当前AI演示工具赛道的核心痛点。从评论反馈看,团队敏锐捕捉到了竞品“生成惊艳、修改痛苦”的共性缺陷,并以此作为差异化核心——通过 Canvas 编辑器提供手动调整能力,同时用“以文改图”的指令逻辑降低操作门槛。这种“AI先搭骨架,人再雕琢细节”的混合模式,确实比单纯生成更有实用价值。但值得警惕的是,编辑器本身的易用性、响应速度和流畅度,才是用户留存的关键,评论区已有反馈因网站动画导致按钮难以点击,这暴露出团队在交互细节上的打磨不足。此外,品牌指南导入功能是亮点,但如何保证从网站URL自动提取的色彩字体等资产足够精确,且能适配多种复杂品牌体系,仍需长期测试。总体而言,Lumo Studios 在“生成后编辑”环节的投入方向正确,但若要将“好用的AI工具”转化为“必不可少的效率工具”,必须在生成质量、编辑效率与品牌适配深度上实现持续迭代,不能止步于“比竞品少改一小时”。

查看原始信息
Lumo Studios
Great ideas deserve great slides. LUMO uses AI to turn your thinking into polished decks instantly, beautifully, and exactly the way you want them. LUMO turns your ideas into beautiful presentations in seconds. Generate slides from a prompt, refine them in the canvas editor, and share with your team - all in one place.

Hey Product Hunt 👋

Nicholas here. We built LUMO Studio because building slides is the part of work everyone quietly dreads hours lost wrestling with layouts, fonts, and alignment instead of focusing on the actual ideas.

LUMO flips that. You describe what you want, and it builds a clean, beautiful, fully-structured deck in minutes. No template-hunting, no design skills needed, no “why won't this box line up” at 1am.

What makes it different:

• It handles structure AND design, not just one

• The output actually looks intentional, not auto-generated

• You stay in control — tweak anything with a sentence

We'd genuinely love your honest feedback, and if it saves you even one late night, that's a win for us. We will continuously build to make it the best tool that everyone can use! AMA in the comments 🙏

2
回复

@nicholas_trajeco will definitely give this a shot, thanks

0
回复

@nicholas_trajeco P.s. I'm having a lot of trouble clicking any buttons on your website due to the animations, maybe I'm just a boomer ¯\_(ツ)_/¯

https://studio.lumotechnology.com/

0
回复

The canvas editor part is very important to me. A lot of AI deck tools create a decent first draft but editing the details afterward can still be painful. Does LUMO give much control over layout and visual style after the deck is generated or is it more focused on getting the first version done quickly?

2
回复

@busra_seker1 Yeah this was honestly the main thing we obsessed over, because we ran into the same problem with every other tool we tried. The draft looks fine and then you spend an hour fighting with it.

So with LUMO you're not stuck with whatever it gives you.

After it generates, you go into a proper canvas editor and can actually move stuff around, resize things, change the layout, swap images, mess with fonts and colors, fix the spacing, whatever you need.

You can also change the overall style and it'll update the whole deck so it stays consistent instead of you fixing every slide one by one.

Or if you don't feel like dragging things around, you can just type what you want changed and it does it.

Basically the first version comes together fast but you've still got full control after.

1
回复

Can you import brand guidelines or do you need to have a template already built out?

0
回复

@sasha_mcclendon 

No template needed at all, you can bring your brand in directly. There's a Style Guide section where you can either upload a brand guidelines PDF, drop in some inspiration images, or just paste your website URL and it'll pull everything from there.

It reads your site and builds out a full brand kit, so it grabs your colors with the actual hex codes, your fonts, logo and brand images, even picks up on your brand values and tone.

Then that becomes the style your decks are built on, so you're not starting from a blank template and trying to recreate your brand by hand.

So short answer, you can import guidelines OR just point it at your site, whichever's easier for you.

1
回复

Interesting approach. How do you balance automation with maintaining a unique brand identity across presentations?

0
回复

@marianna_tymchuk It's the thing we think about most. The automation works inside your brand, not instead of it, so it builds within your saved brand kit (colors, fonts, logo, tone) rather than inventing a new look each time.

The repetitive stuff gets automated, like layout and structure, but the identity stays yours, and since every deck pulls from the same brand kit they stay consistent with each other.

Then you can fine-tune anything in the canvas. Automate the busywork, keep the brand front and center.

1
回复
#15
Treadmill Pro
Control your treadmill from your iPhone, wirelessly
98
一句话介绍:Treadmill Pro 通过蓝牙连接iPhone,让用户摆脱跑带老旧面板的束缚,直接在手机上无线控制速度、坡度和计时,并同步数据至Apple Health,解决了“跑时调参数不便”和“运动数据碎片化”的痛点。
Health & Fitness Productivity User Experience
跑步机控制 蓝牙连接 健康数据同步 iPhone应用 Apple Watch心率 健身工具 智能设备 运动追踪 免费应用 iOS
用户评论摘要:制作者分享了自己因跑带面板难用而开发应用的初衷,评论提问UI在不同跑带上如何自适应,制作者回应目前聚焦共通的“速度/坡度”核心功能,保持界面简洁,未来计划加入滑动手势操控。
AI 锐评

Treadmill Pro 切入的是一个极其细分的“老问题”——大多数健身房或家用跑步机的控制面板确实反人类,物理按键响应迟钝、菜单逻辑混乱。与其等待厂商更新售价高昂的触控屏,不如直接让用户用手机接管。

从产品力看,它的价值不在于技术壁垒(蓝牙遥控原理并不复杂),而在于“找对了受众”:那些早已习惯了Apple Health闭环,且对跑带原装体验不满的中高频跑者。制作者清醒地放弃了兼容所有花哨功能的幻想,只在速度、坡度和计时上做深做透,这种“化繁为简”的克制,远比那些企图模拟跑带完整面板的臃肿App聪明。

但风险同样明显:蓝牙协议兼容性是命门。虽然声称支持5种协议99+设备,但健身房和家用跑带的蓝牙标准高度碎片化,一次系统更新或跑带固件变动就可能“断连”。此外,免费版几乎只能看,Pro解锁健康同步和Apple Watch的功能是正确但不够性感的变现路径——用户是否愿意为一个“遥控器”付费,取决于它能否稳定、无感地替换现有的物理操作。如果连接不稳定,用户一天就会卸载。总的来说,这是一个“小而美”但生长天花板分明的工具,属于特定人群的雪中送炭,而非大众的锦上添花。

查看原始信息
Treadmill Pro
Connect your treadmill via Bluetooth and control speed, incline, timer and stats right from your iPhone. Full Apple Health integration. Free on the App Store.
Hey Product Hunt! I built Treadmill Pro because my treadmill's control panel was clunky and I wanted to run it from my phone instead. It connects over Bluetooth, auto-detects your machine across 5 protocols and 99+ devices, and lets you control speed and incline, track every workout, and sync it all to Apple Health. There's a live Apple Watch heart-rate view too. Free to start, Pro unlocks Health sync, full history and Watch integration. Would love your feedback, happy to answer anything in the comments!
4
回复

Love it when people solve problems they face first hand like this! How does the UI change when connecting to different treadmills since different ones will have varying features?

1
回复

@scott_davidson_jr, hi! Most treadmills share the same core functionality: speed and incline. That's what I use personally, and that's what the app focuses on.

Speed shortcuts work great on physical treadmills with a large built-in display. On a smartphone screen, they tend to clutter the UI and pull your attention away from what matters. I kept the interface clean and minimal on purpose.

In future versions, we'll add speed and incline shortcuts via swipe gestures, so power users can access them without sacrificing the simplicity of the main screen.

Thanks!

0
回复
#16
Clarafy
Type messy and have it instantly polished
94
一句话介绍:Clarafy 是一款零干扰的“混乱翻译器”,让你在任意输入框(Gmail、Slack、ChatGPT等)中直接打字或语音输入杂乱思绪,一键热键瞬间替换成符合上下文的得体文字,省去复制粘贴到AI工具来回加工的繁琐流程。
Chrome Extensions Productivity Education
AI写作助手 文本润色 上下文感知 语音输入 效率工具 浏览器扩展 风格记忆 内容重写 生产力工具 桌面应用
用户评论摘要:用户关注上下文适配能力(如Slack保持随意、邮件正式),以及一键撤销(Ctrl+Z)带来的试用安全感。有人希望用于处理客户反馈的混乱信息,开发者已确认此场景有效。开发者强调工具会识别应用并自动调整风格,且支持学习用户个人语气。
AI 锐评

Clarafy切中了一个被忽视但高频的痛点:写作的“中间态”。现有工具要么提供满屏红线建议(强迫你分段决策),要么让你穿梭于多应用间(破坏心流)。它狠就狠在“零建议”——不给你选择恐惧症,直接给出最终结果,而“撤销”按钮则抵消了用户对失控的恐惧,这是极高明的信任设计。

真正的护城河是“App-Aware”。很多通用AI写作工具只会输出一种“AI腔”,但Clarafy知道在Slack里用梗图语言,在Gmail里用三板斧结构,在ChatGPT里用详细指令。这种基于UI上下文的自动调参,比单纯“调语气滑块”聪明得多。加上“口述→释放→成文”的语音交互,它几乎把“思维速度”与“输出质量”之间的延迟降到了最低。

不过,产品仍面临挑战。首先,“零建议”意味着用户对改写结果的控制权极低,遇到大段改写偏离原意时,撤销是唯一退路,信任修复成本高。其次,风格学习需要大量数据,初期用户可能觉得“通用模式”不够个性。最后,浏览器扩展依赖网页框架,对原生桌面端应用的兼容性(如本地记事本、IDE)才是真正考验。若能彻底打通跨应用的无缝改写,它有望成为比Grammarly更“激进”但更直击用户潜意识的写作层基础设施。

查看原始信息
Clarafy
Most tools force you to edit. Clarafy is a zero-suggestion chaos translator. Type or dictate a messy stream-of-consciousness, hit a hotkey, and it instantly rewrites perfect text in place. Standout Features: App-Aware: Formats contextually for Gmail, Slack, or ChatGPT. Hold-to-Dictate: Ramble out loud; release to inject polished prose. Tone Matching: Learns and mirrors your unique voice. No suggestions, no underlines. Just one-and-done clarity.

@liam_tidholm this looks super useful and the keyboard short cuts are nice! Quick question, does it adapt to context at all, like knowing a Slack message should stay casual but an email should read more formally, or is it one consistent "clean up" pass for now?

5
回复

@tom_palmer_ux Yep, context awareness is a big part of Clarafy.

It doesn’t apply the same rewrite everywhere. Clarafy looks at what you’re writing and adapts automatically. A Slack message stays casual and conversational, an email becomes more polished and professional with clear structure, and an AI prompt gets detailed for better results. So in short it know wether your in Gmail, slack, notion, or ChatGPT etc.

The goal isn’t to make everything sound the same, it’s to make your writing clearer while preserving the intent and tone that fit the context.

We also have tone learning called StyleMemory so over time Clarafy can better understand and match your own writing style rather than forcing a generic AI voice.

2
回复
The workflow for fixing your text right now is so annoying: type a rough draft, copy it, open ChatGPT, paste it, ask it to clean it up, copy it back, and paste it again. I built Clarafy to fix that. Instead of all that bouncing around, you just hold Ctrl + Space (or Alt + Space on Mac) while your cursor is in any text field. It automatically reads the whole box, does a quick processing animation, and swaps your text with a clean, clear version right in place. No manual highlighting or copy-pasting needed. If you don't like the edit, a regular Ctrl+Z brings your original draft right back. The Chrome extension is officially live in the web store today for an early public beta, and I'm currently working on a native Windows desktop app to bring this system-wide. Check it out, try to break it, and let me know what you think! I'm here to answer any questions or fix any bugs you find.
4
回复

This sounds useful for someone like me. My thoughts often come out messy, especially when I'm writing reports, I can spend a long time just polishing the wording and trying to make everything sound clear.

I can also imagine this being helpful for understanding customer feedback or support messages. Sometimes users explain their needs in a very long or messy way, and a tool like this could help "translate" the message into something clearer so we can understand the main point at a glance.

Good luck with the launch!

4
回复

@evakk Really appreciate that.

That’s actually one of the use cases that inspired Clarafy. Most people don’t struggle with what they want to say, they struggle with turning messy thoughts into clear communication.

The customer feedback example is interesting too. We’ve found that Clarafy can be useful for taking long, unstructured messages and surfacing the core idea more clearly, which saves time when you’re processing lots of information.

Thanks for the support and for taking the time to share your thoughts!

2
回复

Ctrl+Z bringing back the original is the detail that makes this feel safe to try. You can use it on something important without committing blind. That's a small design call that probably drives a lot of first-time trust. Congrats on the launch!

1
回复

@jared_salois Thanks! Yeah and we have a revert button after each polish that users can easily click👍

0
回复
#17
Recursi
Self improving vibe coding env with no API fees
93
一句话介绍:Recursi 是一个通过自动化复制粘贴流程、无缝对接网页版AI聊天机器人(如Claude、Gemini等),让开发者无需支付API费用即可高效进行“氛围编程”的自我进化式开发环境。
Developer Tools GitHub Vibe coding
氛围编程 无API费用 AI辅助开发 代码自动化 网页版ChatGPT 自我改进IDE 本地存储 儿童编程 项目模板
用户评论摘要:用户主要关心其“自我改进”机制的实现原理,以及能否跨会话保留项目上下文。开发者回应称,项目上下文可保存在浏览器或本地磁盘;代码更改由LLM统一处理,极少出现依赖冲突;产品同样适用于编程新手,但缺乏系统指引,主要靠示例视频引导。
AI 锐评

Recursi 的核心价值不在于它发明了某种革命性的AI代码生成算法,而在于它以一种极其务实、甚至有些“原教旨主义”的方式,解构并重构了“用AI编程”的传统流程。它敏锐地捕捉到了当前AI写代码的最大痛点:高昂的API费用和频繁的手动复制粘贴带来的心智负担。通过“寄生”在免费网页版聊天机器人上,它用自动化粘贴脚本巧妙地绕过了成本壁垒,这本身就是一种极具野心的逆向工程。

然而,这种“偷懒”的聪明也暴露了其本质的局限性:它并非一个独立的AI引擎,而是一个高度耦合的“管道工”工具。其“自我改进”更多是指开发者通过使用它来高效迭代自己的代码库,而非AI模型本身的进化。尽管开发者声称通过复杂流程解决了大规模项目中的依赖冲突,但这仍建立在用户对LLM行为有深刻理解的前提下。对于普通用户而言,学习的陡峭曲线可能依然存在。

作为一款“独狼”开发的产物,Recursi 在简洁性和探索精神上令人印象深刻,尤其对希望用极低成本玩转AI编程的发烧友和教程制作者极具吸引力。但它距离一个企业级、稳定可靠的工程工具还有显著距离。其最大的贡献,或许是证明了在当前AI生态下,通过巧妙的工程化封装,完全可以在不依赖昂贵API的情况下,榨干大语言模型的生产力潜能。这是一场漂亮的技术游击战,但注定不会是主流战场的解决方案。

查看原始信息
Recursi
An extremely powerful environment for vibe coding, allowing you to use web based chatbots (Claude/Gemini via aistudio/ChatGPT, etc) extremely efficiently while staying firmly in their terms of service. A whole lot of apps are provided as samples / templates, including a YouTube playlist app that not only allows a great YouTube experience without ads (also within terms of service!), but "Guitar Hero for piano" that works with a MIDI piano. So much more.
I've been refining this since ChatGPT first came out, just to streamline coding using the web interface. But at some point last year, it suddenly reached a "recursive self improvement" state where it made huge leaps forward. It was almost scary. I've been trying to add all the right elements to make it a great environment for kids, newbies to coding, teachers, creatives..... everyone. I think it is there. Check out the video on the site, it's long, but flip around in it if you need to. Lots of cool visuals, music stuff, etc.
1
回复

The copy-paste automation angle makes sense as a starting point, the friction of manually finding the right function and replacing it is real. What I'm curious about is how it handles conflicts when the LLM changes a function that other parts of the codebase depend on. Does Recursi detect those downstream breaks automatically or does that still fall on the user to catch?

1
回复

@imoluuu Recursi works in tandem with the LLM and almost never breaks things like that it.

It will change things in both files for, instance it will change a function and change where the function is called if that has to change too.

If you're working on a small app the LLM will tend to have all your code to the entire app, but if you're working on a large app -- for instance recursi itself is about 80,000 lines of code -- technically that can all fit into Gemini 3.5 although it's kind of pushing it.

But you also have a lot of tools for choosing which files to include and for it to know what is going on. Sometimes it does result in things not working or a crash but it's really good at sending that information back to the llm so it can figure out which file it needs to go and pull to fix it.

I mean honestly I could go on for a long time about this whole process but that problem is, if not 100% solved, darn close to it. It's extremely rare that it breaks something in that sense, and when it does it's quite easy to fix - its good at pulling not only files but error messages.

I'll go ahead and tell you about a process I use when working on really big projects which is to take the whole project build a prompt that has all the code to the entire thing, and drop that into Gemini. And I tell Gemini to analyze the problem that I'm talking about, tell me the full path to all the files that are involved even if peripherally, as well as just giving some analysis of how it probably can be fixed. But don't write any code.

Then I take what it gives me paste the entire output into a particular tool in recursi and it will build a new prompt that includes all the file paths saw in that output, that is it will grab all the code to all those files, or if I want to set it to only get the signatures or the documentation I can do that. And then I will use that as a start for a new thread where it's more focused and it doesn't have 100,000 lines of code but might have more like 10,000 that is pertinent to the specific problem. And I find that when you do that it tends to be a little smarter because it's not dedicating a lot of it's thought process to stuff that is completely irrelevant.

So there's lots of tricks like that but those really only come into play when you're working on huge apps if you're working on smaller ones no that problem is basically non-existent. If you look at the video where I tell it to take the 3D model and change it to have those amusing explosion effects or whatever, you'll see it affects multiple files and it goes and changes a whole bunch of methods in different files at the same time. So it knows what it's doing if you feed the context to it correctly and this doesn't take a lot of effort on the part of the user most of the tedium of it is handled by the app itself

1
回复

I'd like to know whether the self-improvement loop is mostly focused on code generation quality, or on helping the system build a better understanding of the project over time.

We've seen a lot of agent failures come from losing context between sessions rather than writing bad code in a single session.

1
回复

@zaid_mallik1 I'm not entirely sure what you're asking about the self-improvement loop. This is an app that I have worked on for almost a year, and it's about 80,000 lines of code. When it started out it was just a couple small files that helped manage cutting and pasting between the llm and files so that it would put the changes in the right place. Those also a script to help grab all the files in a directory and put them into a single markdown file so that I could paste it. So like maybe 500 lines of code total.

And when it got to maybe 2000 lines of code, suddenly my productivity started to get much higher because the system was just doing better at all these little tedious things. Maybe by the time it got to 20,000 lines of code, I was way more productive but still here and there there were little issues where I'd have to drop to the file system to fix something. But because my productivity was enhanced that much, fixing these issues with all the faster. So it took way longer to get from 2000 to 20,000 lines of code, then it took to get from 72,000 to 80,000 lines of code. Usually you'd almost expect it to be the other way around because once things get big it starts getting fragile. But again through the process it just keeps improving the thing in every way.

Everything about it gets better the more it increases its functionality and improves itself.

That said when I talk about self-improvement, I'm not talking about what most users are expected to use it for, they're probably not going to be improving recursi itself. Some of them might, I hope for that, but the primary reason they're getting benefit out of self-improvement is that the app is just getting really really sophisticated for something that one person could have put together.

And when I talk about lines of code, obviously I'm not claiming that more lines of code is necessarily better. Recently I went through a process of removing about 20 to 30,000 lines, as it became clear that so much old functionality the llm had just swept under the rug and papered over rather than removing as the app evolved. That is one of the toughest things I've had with llms that it does that. I don't know if the solution to that can be directly built into the app, but it's certainly something that can be taught to users..... You need to regularly go through and tell the llm to get rid of old functionality. And you need to not trust it when it says it's wiped it clean and it's now pristine you need to open it up in a brand new thread and ask it to find the same things and often it will find still a bunch that are left behind. I hope the models get better at this, but if not, it's certainly something that users can learn how to manage so they're apps don't just get filled up with a bunch of crust and then you start seeing old features coming back because the llm sees it in there and thinks it's gospel when it's really obsolete.

1
回复

I like that you're thinking abt beginners and kinds, not only experienced coders. That changes the product a lot. Does Recursi guide people toward a first small project or is it better for users who already know what they want to build?

1
回复

@busra_seker1 There's not a huge amount of guidance other than "watch a video showing how simple it is". I do one in the video where I grab the basic 3d app and have it make it so when you click on the sample geometry, it does some amusing explodey effects. It takes only a minute or so and is easy to copy.

I'd like to do a lot more things like that, and add more samples.

0
回复

The "no API fees" framing is interesting but the part I'd want to understand is what's actually happening under the hood. Are you running models locally, routing through your own hosted inference, or something else? That changes the tradeoff pretty significantly, especially for anything compute-heavy. Also curious what "self-improving" means concretely here, whether the environment is updating prompts and context based on your past sessions, or if it's something closer to fine-tuning on your codebase over time.

1
回复

@fberrez1 You can see it in the video what's happening.... it uses the web-based chatbots that are free or cheap. It uses copy and paste but it makes it super efficient so you're not having to go in and open individual files and find the right place and select the right text and then control V or whatever. Instead you just click a single button and it has Within the code coming back from the llm information as to which file and what function to replace. There are other things for instance you can add a dependency or delete a function or things like that but for the most basic stuff you just are replacing one function within one class within one file and the system just knows how to do it so you don't have to go and do it yourself you just click a couple buttons.

Look at 11:50 to 12:30 or so in the video where I do it..... it takes half a minute or so only because I am explaining it as I go.

1
回复

Does it retain what it learned about your codebase or does each session start fresh?

1
回复

@joy_shekhar It retains it. The first and simplest mode saves it into the "indexed db" in your browser. It is sort of like cookies, except it can store a whole lot of data. You can clear it from your browser history in similar ways to clearing cookies. But it will automatically be there when you come back.. You can tell it to ignore it and come back fresh. (it only changes the functions that change, so not the whole project. good for fairly minor tweaks)

The second thing you can do it open a directory on your disk, and you can "fork" a project into it, renaming it if you want (including changing the names of Javascript classes within). Then it stores it there. Each time you come back to the site, there is a button you can click to re-open your directory and it will just reopen the last subdirectory(s) you were in automatically. So it will be on your disk forever, saving the entire project (but not copying things like shared libraries)

(if you use the first way, you can "upgrade" to the second way, through a menu item. It will copy your projects to disk.)

The third way is to download the repo, which is appropriate if you are changing the whole app itself, or if you want it to store all shared libraries in case the site disappears or something. But it still uses the technology above (the browsers ability to read and write to an approved directory and all subdirectories) to read and save to disk.

1
回复
#18
English Jobs
English jobs in Europe. No local language needed
90
一句话介绍:English Jobs 是一个面向国际求职者、无需掌握当地语言即可快速筛选并直投15万+欧洲在招英语岗位的免费求职工具,核心解决“多语国家职位要求本地语言”导致的求职筛选耗时与挫败感。
Hiring Education Career
用户评论摘要:用户普遍认可产品精准解决了欧洲求职中的语言门槛痛点,尤其称赞直投公司官网的简洁流程。但反馈指出:仍有少量要求本地语言或“德语优先”的岗位混入;过滤器偶现bug、字段抓取不全;建议增设实习/应届生筛选、经验年限细化、用户反馈机制以持续优化信噪比。
AI 锐评

English Jobs 精准切中了一个被忽视但极度真实的痛点——在欧洲,语言门槛不仅是“不懂德语”,更是“每份职位描述看到第三段才发现需要德语”的反复消耗。产品价值不在于技术壁垒多高,而在于它做了“对的事且只做对的事”:去中介化,直链公司官网;去广告,零干扰体验;去伪存真,试图过滤模糊表述的隐性语言要求。

然而,核心挑战正在于“模糊表述的边界”。评论区反复提及的“German required 标题仍出现”绝非偶然——许多企业的JD本身就是双语混写,HR对“英语优先”的定义也充满弹性。这意味着纯规则或简单NLP无法根治假阳性,而靠用户反馈打标签又是成本高昂的苦活。如果这条线划得不够“狠”,产品就会滑向另一个LinkedIn数据子集,而非一个可信赖的信号放大器。

此外,产品目前定位“为不会本地语言的人找英语岗”,这个客群天然包括印度、东南亚、东欧、拉美的高学历冷启动求职者。他们对“实习/Entry Level”和“经验年限”的过滤需求极强,而这恰恰是目前数据抓取的弱项。如果团队能尽快建立一套“经验年限+语言信心分”的双轴评分体系,并与高频职位推送联动,就能从“好用的工具”变为“求职者的长期伴侣”。

一句话本质:它不是又一个招聘聚合器,而是用“语言门槛”这一单一切片,重新定义跨国求职的信息获取效率。但能否站稳,取决于它是否敢于“宁可少推、也要准推”——信任一旦打折,用户就会回到那个令人疲惫的搜索循环。

查看原始信息
English Jobs
english-jobs.com helps you find English-speaking jobs in Europe. You do not need to speak German. See 150,000+ live jobs in 10+ countries. Every job links to the real company page, so you apply in one click. It is free.
Hi Product Hunt! 👋 I am Kapil. I made english-jobs.com The problem In Germany, many jobs need German. In France, many jobs need French. In Spain, many jobs need Spanish. If you do not speak the local language, it is hard to find a good job. You open many jobs, but they say “German required” or “French required”. This is slow and tiring. I had this problem. I got frustrated. So I built english-jobs.com. Now you can find English jobs only. No ads. Simple page. 150k+ jobs, free tools, apply on the company site.
140
回复

@ikapilm Impressed with the concept and execution. The job listings are actually current and the Germany-first approach makes sense given the tech hub there. Would love to see similar expansion to other European countries. Definitely beats scrolling through LinkedIn for English-only roles.

0
回复

@ikapilm It's great to see you thinking about English-speaking people and helping them by providing resources to find job opportunities in Europe.

0
回复

@ikapilm Looks genuinely useful, easy to browse English speaking jobs across Europe, and the direct company application links are a nice touch.

0
回复

I think this might be very helpful to folks like me and thousands of others. In 2024, I was searching for a college in India, and I didn't have multiple good options, but I did have some in Germany. There was a major problem in Germany that you need to know German for most opportunities, but I didn't have the time to learn it. I was helpless then, but with something like English Jobs coming into the market now, I feel that the things I had to suffer through while searching for a college won't happen to me when I start searching for a good job. Thank you for creating this :)

4
回复

@ritik_anand3 Happy to know it helps :))

Would be curious about any feedback that we can get on the product as well

1
回复

Being a student in India looking at opportunities abroad, the language barrier is a real concern. I like that this filters at the source instead of showing irrelevant listings. 150K+ direct company links is useful. An internship or fresher filter would be a great addition.

3
回复

@shaurya_singh21 hey thanks for your feedback, there are internship/fresher filter if you filter from seniority on the platform. let me know if you need any help locating it

2
回复

As an Indian CS student, I've definitely had moments where I found an interesting role and then immediately realized it required a local language. Really like how this removes that friction upfront and keeps the job search focused. The product feels simple, but it's solving a very real problem for international applicants.

2
回复

As a college student in India, I'm always looking out for remote roles to upskill and get some real-world experience. I'm really glad I found this site because having actual opportunities aggregated in one place saves me from the usual frustration of digging through irrelevant listings.

That being said, it isn't completely perfect yet. Just wanted to offer a little constructive feedback:

  • Even though it's an english-jobs site, I still saw a few bilingual roles slip through. It would be better if it strictly filtered for just English speakers, or provides an option for users to specify the languages they know in their profile. That way, they could get specific targeted jobs.

  • For some of the listings, the category, job type, and seniority data wasn't fetched properly and looked bugged. The system you're using to fetch details from listings should have a certain fallback if it fails to fetch those details. This is an example from one of the listings: "Category enrichmentfailed"

  • It would be super helpful to add a specific "years of experience required" section in jobs, especially for freshers like me who need to know if we actually qualify.

Overall though, I really appreciate you building this. It's a really solid start and already incredibly useful!

1
回复

With over 150,000 live listings, how does your platform filter out roles that might still secretly require 'professional proficiency' in the local language?

1
回复

As an engineering student in India, finding opportunities abroad often feels difficult because many job portals don't clearly mention language requirements. I really like how English Jobs focuses specifically on English-speaking roles and links directly to company websites. This saves a lot of time and effort. Adding filters for internships and entry-level positions would make it even more useful for students and fresh graduates.

1
回复
As someone from India who occasionally looks at opportunities in Europe, this solves a very real frustration. A lot of job boards make you spend time reading through listings only to discover near the end that German,French, or another local language is required. I also checked the site and liked that it sends applications to the original company pages rather than reposting jobs through multiple layers. Overall this product is really good that solves a genuine pain point.
0
回复

Being a student in India looking for opportunities abroad, the language barrier is a real concern. This is a great source to find the right job with many options. 150K+ direct company links is useful.

0
回复
Congratulations on the launch! The filtering at source is what makes this different most boards show the listing first and bury the language requirement three paragraphs in. The direct company page links also remove the middleman friction that kills most aggregator apply flows.There is one question like do you have a feedback loop where users can flag mislabeled listings to retrain it over time ?.
0
回复

The "German required" buried at the end of an otherwise perfect JD is one of the most demoralizing things about job hunting in Europe as a non-EU applicant. Glad someone finally built a clean solution for this instead of just complaining about it.

0
回复

Congratulations on the launch! The idea solves a real problem for international job seekers. I'm curious, how do you verify whether a job truly requires only English and doesn't have hidden local-language requirements in the description?

0
回复

Ended up spending more time on the methodology page than the homepage and what stood out was how specific it is about the limitations. It even calls out where the English filter breaks today: postings in Spain, Italy, and France with English boilerplate but a local language requirement. Looking forward to the post-launch improvements like language confidence scoring, title normalization, and candidate feedback loops.

I actually ran into a few of those false positives on the jobs page. One was literally titled "AI Engineer... German Required." Seeing the same edge cases documented upfront made the methodology feel a lot more credible.

Congrats on the launch, Kapil!

0
回复

Finally a proper tool, which can help to find only English jobs! Else I was gonna start learning French or German

0
回复
Congrats on the launch! This solves the exact pain of going through jobs where the title says “English” but the description later hides a German/French/Spanish requirement. I like that you kept the flow simple and intuitive. One thing I’m curious about: how are you detecting the language requirements when they are phrased indirectly, like “client-facing German market role” or something like “local stakeholder communication”? Is it like a rule-based parsing, a classifier, or embeddings over the JD text to separate hard requirements from optional ones?
0
回复

@ikapilm What I liked most is that this solves a problem much earlier in the job search process than most platforms do. As someone who looks at opportunities abroad, I have often seen roles that look interesting at first but later turn out to require a local language. Filtering for English-friendly jobs upfront removes a lot of that frustration.

The country-level insights and company research pages were a nice surprise too—they make the platform feel more useful than a simple job board.

One thing I kept thinking about is that job-posting language and workplace language are not always the same. A role may be listed in English, but day-to-day communication can still happen mostly in a local language. It would be interesting to see how future versions help candidates understand that part of the experience as well.

0
回复

Congrats on the launch! @ikapilm Navigating global job boards can be incredibly frustrating when language requirements aren’t explicitly clear up front. Filter structures like this save job seekers hours of vetting.

Out of curiosity, are there plans to integrate a filter or badge for 'Visa Sponsorship Available' vs. 'Local English-Speaking Only'? For expats and international applicants, that’s usually the second biggest hurdle right after language alignment. Looking forward to watching this scale!

0
回复

this honestly solves a problem i've run into quite a few times. when looking for internships or jobs, a lot of time gets wasted figuring out whether a role is actually accessible for english speakers. having that clarity from the start makes the process so much easier. really useful idea, congrats on the launch

0
回复

The English-first approach actually solves a real problem, and I love that the apply links go straight to the company's real page instead of dead re-posts. Found a couple of roles I hadn't seen elsewhere too.

The honest feedback would be that the auto-detected "English" label let through a few listings that still wanted the local language, so a verified badge would go a long way. And some visa or sponsorship info would be huge, especially for non-EU folks like me.

0
回复

a very practical idea and handy for those working in Europe. will you also highlight companies that actively support relocation or visa sponsorship, and will companies be able to post directly, or is it curated from existing job boards? there should also be a verified job listing tag who pass a legitimacy check.

0
回复

What got me was the ATS indexing directly from company career pages. as a student most of the european roles i find on linkedin are either expired or just reposted from somewhere else. pulling fresh listings straight from the source means you're not wasting applications on ghost jobs which honestly changes how useful a job board actually is.

0
回复

As a student who will be searching for roles in Europe soon, finding English-speaking roles seems unnecessarily difficult on regular job boards. The fact that this filters that out upfront would save a lot of wasted time. Do you also list internships or is it mostly full-time roles? That would make this even more useful for people just starting out.

0
回复

Really clean way to find English-only roles across Europe w/o getting lost in translation. having the direct corporate links instead of third-party redirects makes things way easier. it'd be awesome to have a simple 'job alert' notification toggle for a specific country though, just so you don't have to manually check back every day

0
回复

@ikapilm This is so true. It is incredibly frustrating to open ten job listings just to see a language requirement at the very bottom. Such a great idea to fix this

0
回复

Really useful concept. One feature I'd love to see is a dedicated visa sponsorship filter. For many international applicants, language requirements and sponsorship availability are the two biggest barriers. Having both visible upfront would make the platform even more valuable.

0
回复

Interesting niche product, congrats on the launch!

Really like the focus on English-speaking roles in Europe. I'm curious about how you source the jobs — especially how you determine the English requirement. Do you include roles where English is mainly used in interviews but limited day-to-day, or only positions where the company and team are fully English-friendly?

0
回复

Being a student in india it is very difficult to do internships and other opportunities abroad, but through this we can filter out real opportunity instead of randon listing, thus very helpful

0
回复

I spent a few minutes exploring the platform and liked how easy it is to jump directly to the company's application page. That saves a lot of time compared to many job boards.

I'm curious how you handle expired roles and keep the listings fresh across so many countries. For students and early-career professionals looking beyond their home country, this seems like a useful resource.

0
回复

So true. Finding language blockers hidden deep in a job description is a massive time sink. Pair that with buggy, third-party application flows that break halfway through, and the job hunt becomes twice as frustrating as it needs to be. Direct links and better filter controls should be standard practice by now.

0
回复

I think this might be very helpful to folks like me and thousands of others. In 2024, I was searching for a college in India, and I didn't have multiple good options, but I did have some in Germany. There was a major problem in Germany that you need to know German for most opportunities, but I didn't have the time to learn it. I was helpless then, but with something like English Jobs coming into the market now, I feel that the things I had to suffer through while searching for a college won't happen to me when I start searching for a good job. Thank you for creating this :)

0
回复
#19
Cleo Atlas Legal API
The world's regulations, machine-readable.
42
一句话介绍:Cleo Atlas Legal API 将全球177个司法管辖区的25.6万+条法规转化为机器可读的API,帮助企业和开发者快速查询产品合规、金融、AI等领域的法律要求,解决法规“存在但不可查”的痛点。
API Legal Artificial Intelligence
法规API 合规查询 法律科技 产品合规 全球监管 机器可读法规 数据标准化 智能体 跨国合规 治理智能
用户评论摘要:用户关注多语言处理、法规冲突、变更追踪及数据颗粒度。创始人回应:支持统一Schema但保留原版,冲突自动暴露,版本可追溯至义务级。总体评价积极,但质疑数据源权威性与高频更新的实时性。
AI 锐评

Cleo Atlas踩中的是一个真实且昂贵的痛点——全球法规的分散性、非结构化与版本混乱,让合规成为品牌出海与AI落地的隐形杀手。但“机器可读”不等于“机器可懂”。其核心卖点(覆盖25.6万条法规、义务级解析)看似庞大,实际面临两大拷问:第一,数据源清洗的“最后一公里”仍是人肉劳工,仅靠MARIA引擎无法解决政府文档中大量的模糊表述与歧义,创始人承认“人类验证”是关键环节,这决定了其拓展速度难以指数级爆发。第二,合规决策权不会轻易交给一个第三方API,尤其当法规冲突时,用户需要的不是“自动暴露冲突”,而是基于判例或监管指引的优先级建议。目前看,Cleo更像一个高效的“法规搜索引擎”而非合规决策引擎。它能帮AI律所、合规SaaS省掉爬虫和人工解析成本,但距离成为“合规操作系统”还差一个从结构化到决策化的跃迁。对于尚未形成内部合规中台的大品牌,这是加速器;对于已养肥了内部法务团队的巨头,则是一个需要验证可信度的候选工具。本质上,它是在用做数据基础设施的思维打法律市场,这比做法律AI应用更苦更重,但护城河一旦建成,难以绕过。

查看原始信息
Cleo Atlas Legal API
Cleo Atlas turns 256,000+ regulations across 177 jurisdictions into one machine-readable API. Two atlases: Product (cosmetics, devices, food, toys, electronics) and Global (finance, AI, data, ESG). Query any law, anywhere. Built by Cleo Labs.
Hi Product Hunt, Naomie here, co-founder at Cleo Labs! Two years ago we started Cleo because we kept seeing the same painful scene: a brand spends nine months building a beautiful new product, then learns three weeks before launch that one ingredient is banned in the country they were betting on. Or that the EU AI Act actually applies to that little internal model they shipped last quarter. Every time, the gap was the same — the law existed, it was public, but it was unqueryable. So we built the layer underneath. Cleo Atlas is one API exposing two atlases: a Product Atlas with 46,031 regulations across 20 product categories and 50 jurisdictions, and a Global Atlas with 210,508 regulations across 177 jurisdictions — health, finance, environment, AI, data, labor, tax. 234M+ documents normalized into clean JSON by MARIA, our retrieval engine. Today it already runs inside the biggest brands. We are now opening it so anyone building a compliance product, a legal copilot, or an AI-governance layer can stop scraping ministry websites at 2 a.m. The explorer is free, no credit card: legaldata-public.cleolabs.co. Hit it, break it, ask it weird questions. I would love your honest feedback especially: which regulation would you query first, and what does our API miss for your use case? I will be in the thread all day.
3
回复

@naomie_halioua The regulatory side of global commerce is incredibly complex, so having a structured compliance map instead of manually tracking requirements across dozens of jurisdictions sounds really valuable. How do you handle situations where regulations change frequently or different authorities provide conflicting guidance? Awesome work on this!

0
回复

really interesting launch. how do you handle multilingual regulations and legal terminology? is everything normalized into a single schema?

1
回复

@barnaby_lloyd Unified schema, but original source preserved plus an interpretation layer, and ambiguous terms go to human validation rather than being flattened :)

0
回复

curious about the underlying data model. how granular is the normalization process? can developers access both raw source references and structured outputs?


1
回复

@james_carter35 We normalize at the obligation level, not the document level each regulation is broken into discrete requirements with their own metadata (jurisdiction, authority, product category, effective dates). That's what lets us answer "what applies to this product in this market."

0
回复

Well good luck rank is already yours even if you feature it for monthly 🙌🎉

1
回复
@xanderiang thank you for your continuous trust !!
0
回复

Regulatory data is notoriously messy. how does MARIA deal with ambiguous or poorly formatted government documents? are confidence scores exposed through the API?


0
回复

interesting vision. how does Atlas handle regulatory documents that are updated incrementally? do users see the exact amendments and changes?


0
回复
@gaius_loxley Great question 🙂 Yes Atlas tracks regulations as they evolve, not just snapshots. When a document is amended incrementally, you see what changed, when, and the diff between versions, so nothing slips by silently. That versioning is a big part of why teams trust it for ongoing compliance rather than one-off lookups.
0
回复

interesting product! can users build jurisdiction specific compliance dashboards on top of the API? or is Atlas primarily focused on data retrieval?


0
回复
@luz_bidelspach Thanks Luz! Both 🙂 Atlas is the structured open-data layer, the API lets you build your own jurisdiction-specific dashboards on top
0
回复

The human-in-the-loop validation piece is what makes this credible to me.
How long does a full compliance map take from URL submission to verified output?

0
回复
@abod_rehman thanks Abdul! The pipeline runs in minutes, the rest is human validation depending on jurisdiction complexity. Speed where it’s safe, human eyes where it counts. :)
0
回复

Congrats on the launch!
The Digital Product Passport deadline in 2030 is real and most brands are completely unprepared. Is that already part of what Cleo maps for EU markets?

0
回复
@boyuan_deng1 Thanks! And yes, DPP is squarely in scope. We map ESPR / Digital Product Passport requirements for the EU and you’re right, most brands are nowhere near rea
0
回复

very interesting approach to regulatory intelligence.how do you handle conflicting regulations across multiple jurisdictions? does the API surface those conflicts automatically?


0
回复
@daniel_harris11 Great question! Yes, it’s core: when a product hits diverging requirements across markets, the API surfaces the conflict expli
0
回复
🚀🚀🚀🚀🏆🏆
0
回复

@alexandre_bloch thanks for contributing a lot on this one!!!

0
回复
#20
Liance
Connect systems. Collect proof. Stay audit ready.
19
一句话介绍:Liance 帮助初创团队连接业务系统、自动收集合规证据并追踪控制状态,在审计准备场景中,解决因手动操作繁琐、现有工具门槛高导致的合规进度滞后痛点。
SaaS Developer Tools Security
合规自动化 SOC 2 准备 初创企业工具 审计证据收集 控制追踪 安全审查 自动化工作流 合规管理 SaaS 信任管理
用户评论摘要:用户点赞其源于真实痛点(非追热点),设计精美。有反馈指出合规流程中的规划与交接常被忽视,保持上下文连贯是难点。一条疑似刷票回帖获无视,核心讨论集中在产品实用性和团队真实需求。
AI 锐评

Liance 切中的确实是一个“沉默但致命”的痛点:小团队不是不想做合规,而是被传统安全审计的繁琐和昂贵吓退。产品思路清晰——把“手动截图+飞书表格”式的审计准备,转向系统连接与证据自动采集,本质上是将合规从“一次性灾难”降维成“持续低负担”。创始人有工具类连续创业背景,设计感和用户洞察在线,这从评论区对视觉的推崇可见一斑。但必须指出:SOC 2 只是合规冰山一角,且关键在于对“控制点”的理解深度,而非界面有多美。19 票的冷启动数据也反映出这类产品面临的尴尬——最需要它的初创团队可能还没意识到痛;而意识到痛的团队,一旦越过某个规模又会倾向更成熟的 Vanta 或 Drata。Liance 的真正价值不在于替代,而在于做它们的“前置仓”:在团队年营收不到 50 万美金、客户还没问 SOC 2 时,帮它们用最低成本把“雏形合规”跑通。如果它能实现与下一代 AI 审计标准的自动对齐,并在证据失效前主动预警,那就能从“服务商”升级成“护航者”。但目前来看,产品还处于“把 Excel 搬上 Web”的早期阶段,核心壁垒尚未建立——不要低估合规的法律属性,也不要把“好看”和“好用”混为一谈。

查看原始信息
Liance
Liance helps startups stay audit ready by connecting systems, collecting evidence, tracking controls, and keeping compliance work moving between reviews.

Hi everyone,

I'm Shashank, founder of Liance.

I started working on Liance after building Superplan.md and speaking with a larger company that was interested in using it. The conversations were going well, but one of the requirements was SOC 2 readiness.

The funny part is that I knew almost nothing about SOC 2 at the time.

That sent me down the compliance rabbit hole.

As I started researching the process and looking at existing solutions, I realized two things. First, a lot of compliance work still happens through spreadsheets, screenshots, documents, and manual follow-ups. Second, many solutions felt out of reach for smaller startups and early-stage products.

I wanted to build something that helps even small teams understand where they stand, connect their systems, collect proof, and improve their readiness before compliance becomes a blocker.

That's how Liance started.

I'd love to hear how your team handles security reviews, compliance requirements, or customer trust questionnaires.

10
回复

@ishashankmi 

"Hey, just upvoted your launch. Would appreciate your support back.

https://www.producthunt.com/posts/deliveryman-ai"

0
回复

@ishashankmi Love the product.

What stands out is that both Liance and Superplan came from a real founder pain point rather than a trend-chasing idea. You built something useful, hit a wall yourself, and then turned that obstacle into a product opportunity.

Also, both websites are ridiculously beautiful. The attention to detail in the design and product experience makes it obvious that a lot of care went into them.

The compliance space desperately needs products that make security and trust accessible to smaller teams instead of feeling like an enterprise-only game. Liance seems to be tackling a problem many founders don't think about until a big customer asks for it.

Huge kudos to the creators of both Liance and Superplan for solving meaningful problems and shipping them with such a high design bar. 🚀

0
回复

We are excited to get this out finally & we would love any feedback about the product!

3
回复
1
回复

This is definite need, planning and handing off those plans is what is work today essentially!

1
回复

@siddhartha_saxena2 Agreed. The planning part gets a lot of attention, but handoff and keeping context intact across people and tasks is often the harder problem. :)

1
回复