{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Hi HN,
Three months ago, I took a job at Baseten to help craft and document an application builder that lets data scientists build full-stack, production-ready applications around their ML models without worrying about containers, Flask, or React. From my first day, everyone was focused on what would happen today: opening up our public beta. I\u2019m super excited to see what you build with Baseten.
If you want to take Baseten for a full-speed test drive, follow along with this tutorial, where you can build and deploy an application in 20 minutes: https://docs.baseten.co/getting-started
While Baseten is built for data scientists and machine learning engineers, something I\u2019m particularly excited about that doesn\u2019t come up often when we talk about Baseten is how it also makes building with ML available to people like me with a general software engineering background but no real experience with ML. With our library of pre-trained models, you can build and deploy an application around models for tasks like sentiment analysis, image classification, and speech transcription. By building applications around pre-trained models, I\u2019ve gained a deeper understanding of the use cases, capabilities, and limitations of machine learning.
If you want to play around with some models and applications without signing up for an account yet, check out our gallery (https://baseten.co/gallery) and try the demo apps.
P.S. We are also hiring; I found Baseten from HN."},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Show HN: Baseten \u2013 Build ML-powered applications"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/"}},"_tags":["story","author_philipkiely","story_31169193","show_hn"],"author":"philipkiely","children":[31169213,31169382,31169384,31177381,31177475,31178534],"created_at":"2022-04-26T16:05:37Z","created_at_i":1650989137,"num_comments":11,"objectID":"31169193","points":112,"story_id":31169193,"story_text":"Hi HN,
Three months ago, I took a job at Baseten to help craft and document an application builder that lets data scientists build full-stack, production-ready applications around their ML models without worrying about containers, Flask, or React. From my first day, everyone was focused on what would happen today: opening up our public beta. I\u2019m super excited to see what you build with Baseten.
If you want to take Baseten for a full-speed test drive, follow along with this tutorial, where you can build and deploy an application in 20 minutes: https://docs.baseten.co/getting-started
While Baseten is built for data scientists and machine learning engineers, something I\u2019m particularly excited about that doesn\u2019t come up often when we talk about Baseten is how it also makes building with ML available to people like me with a general software engineering background but no real experience with ML. With our library of pre-trained models, you can build and deploy an application around models for tasks like sentiment analysis, image classification, and speech transcription. By building applications around pre-trained models, I\u2019ve gained a deeper understanding of the use cases, capabilities, and limitations of machine learning.
If you want to play around with some models and applications without signing up for an account yet, check out our gallery (https://baseten.co/gallery) and try the demo apps.
P.S. We are also hiring; I found Baseten from HN.","title":"Show HN: Baseten \u2013 Build ML-powered applications","updated_at":"2024-09-20T11:00:59Z","url":"https://www.baseten.co/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ingve"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Base Ten for Almost Everything"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://randomascii.wordpress.com/2016/02/13/base-ten-for-almost-everything/"}},"_tags":["story","author_ingve","story_11096088"],"author":"ingve","children":[11096433,11096685,11097141,11097160,11097999,11099353,11099483,11099623,11099643,11099762,11099855,11100010,11100432],"created_at":"2016-02-13T22:54:44Z","created_at_i":1455404084,"num_comments":82,"objectID":"11096088","points":51,"story_id":11096088,"title":"Base Ten for Almost Everything","updated_at":"2025-09-08T02:16:02Z","url":"https://randomascii.wordpress.com/2016/02/13/base-ten-for-almost-everything/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"Hooke"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Renaissance Science: the base ten number system"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://thonyc.wordpress.com/2021/05/05/renaissance-science-ix/"}},"_tags":["story","author_Hooke","story_27179804"],"author":"Hooke","children":[27194584,27194679,27196876],"created_at":"2021-05-17T03:26:56Z","created_at_i":1621222016,"num_comments":11,"objectID":"27179804","points":29,"story_id":27179804,"title":"Renaissance Science: the base ten number system","updated_at":"2024-09-20T08:36:57Z","url":"https://thonyc.wordpress.com/2021/05/05/renaissance-science-ix/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"sahillavingia"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"BaseTen: The fastest way to build ML-powered applications"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://baseten.co"}},"_tags":["story","author_sahillavingia","story_27224638"],"author":"sahillavingia","children":[27225224,27225448],"created_at":"2021-05-20T17:57:02Z","created_at_i":1621533422,"num_comments":4,"objectID":"27224638","points":20,"story_id":27224638,"title":"BaseTen: The fastest way to build ML-powered applications","updated_at":"2024-09-20T08:35:11Z","url":"https://baseten.co"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mikejulietbravo"},"story_text":{"matchLevel":"none","matchedWords":[],"value":"Launched today - happy to answer any and all questions!"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Show HN: Baseten Chains \u2013 Framework and SDK for Multi-Model AI Products"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/introducing-baseten-chains/"}},"_tags":["story","author_mikejulietbravo","story_40813005","show_hn"],"author":"mikejulietbravo","children":[40813045,40813434,40813436],"created_at":"2024-06-27T17:39:45Z","created_at_i":1719509985,"num_comments":5,"objectID":"40813005","points":9,"story_id":40813005,"story_text":"Launched today - happy to answer any and all questions!","title":"Show HN: Baseten Chains \u2013 Framework and SDK for Multi-Model AI Products","updated_at":"2024-09-20T17:19:39Z","url":"https://www.baseten.co/blog/introducing-baseten-chains/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mich5632"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"High performance client for Baseten.co"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://github.com/basetenlabs/truss/tree/main/baseten-performance-client"}},"_tags":["story","author_mich5632","story_44270214"],"author":"mich5632","children":[44270215],"created_at":"2025-06-13T16:58:21Z","created_at_i":1749833901,"num_comments":1,"objectID":"44270214","points":7,"story_id":44270214,"title":"High performance client for Baseten.co","updated_at":"2025-06-14T03:13:10Z","url":"https://github.com/basetenlabs/truss/tree/main/baseten-performance-client"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"kodablah"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Baseten raised a $1.5B Series F and achieved a $13B valuation"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/announcing-our-series-f/"}},"_tags":["story","author_kodablah","story_48632248"],"author":"kodablah","created_at":"2026-06-22T16:18:34Z","created_at_i":1782145114,"num_comments":0,"objectID":"48632248","points":5,"story_id":48632248,"title":"Baseten raised a $1.5B Series F and achieved a $13B valuation","updated_at":"2026-06-22T18:59:42Z","url":"https://www.baseten.co/blog/announcing-our-series-f/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"How BaseTen is using \u201cdocs as code\u201d"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://blog.baseten.co/docs-as-code/"}},"_tags":["story","author_philipkiely","story_30617631"],"author":"philipkiely","created_at":"2022-03-09T17:44:25Z","created_at_i":1646847865,"num_comments":0,"objectID":"30617631","points":5,"story_id":30617631,"title":"How BaseTen is using \u201cdocs as code\u201d","updated_at":"2024-09-20T10:39:18Z","url":"https://blog.baseten.co/docs-as-code/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Baseten raises $150M Series D at $2.15B"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://fortune.com/2025/09/05/exclusive-baseten-ai-inference-unicorn-raises-150-million-at-2-15-billion-valuation/"}},"_tags":["story","author_philipkiely","story_45139326"],"author":"philipkiely","children":[45140575],"created_at":"2025-09-05T14:57:30Z","created_at_i":1757084250,"num_comments":1,"objectID":"45139326","points":2,"story_id":45139326,"title":"Baseten raises $150M Series D at $2.15B","updated_at":"2026-03-05T22:39:20Z","url":"https://fortune.com/2025/09/05/exclusive-baseten-ai-inference-unicorn-raises-150-million-at-2-15-billion-valuation/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mikejulietbravo"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Open Source Inference Engine Baseten Raises $40M from IVP, Spark and Greylock"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/announcing-our-series-b/"}},"_tags":["story","author_mikejulietbravo","story_39710361"],"author":"mikejulietbravo","children":[39710371],"created_at":"2024-03-14T23:42:47Z","created_at_i":1710459767,"num_comments":1,"objectID":"39710361","points":2,"story_id":39710361,"title":"Open Source Inference Engine Baseten Raises $40M from IVP, Spark and Greylock","updated_at":"2024-09-20T16:34:13Z","url":"https://www.baseten.co/blog/announcing-our-series-b/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"todsacerdoti"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Target 1: Baseten"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.silares.com/targets/target-1-baseten"}},"_tags":["story","author_todsacerdoti","story_46751640"],"author":"todsacerdoti","created_at":"2026-01-25T07:28:30Z","created_at_i":1769326110,"num_comments":0,"objectID":"46751640","points":2,"story_id":46751640,"title":"Target 1: Baseten","updated_at":"2026-03-05T23:25:09Z","url":"https://www.silares.com/targets/target-1-baseten"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ollayf"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Serverless Infrastructure for AI apps \u2013 3x perf of baseten, 1/5 the cost"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www.hyperpodai.com"}},"_tags":["story","author_ollayf","story_44937085"],"author":"ollayf","children":[44937086],"created_at":"2025-08-18T03:24:16Z","created_at_i":1755487456,"num_comments":0,"objectID":"44937085","points":2,"story_id":44937085,"title":"Serverless Infrastructure for AI apps \u2013 3x perf of baseten, 1/5 the cost","updated_at":"2026-03-05T22:30:26Z","url":"https://www.hyperpodai.com"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"elmazout"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Baseten, inference platform, $75M Series C, $825M valuation"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://twitter.com/basetenco/status/1892259130540179863"}},"_tags":["story","author_elmazout","story_43109768"],"author":"elmazout","created_at":"2025-02-20T00:46:54Z","created_at_i":1740012414,"num_comments":0,"objectID":"43109768","points":2,"story_id":43109768,"title":"Baseten, inference platform, $75M Series C, $825M valuation","updated_at":"2025-02-20T00:54:26Z","url":"https://twitter.com/basetenco/status/1892259130540179863"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"CarolineW"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Q: In base ten 1=0.999\u2026, but what about in other bases? What about in base 1?"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"http://www.askamathematician.com/2017/01/q-in-base-ten-10-999-but-what-about-in-other-bases-what-about-in-base-1/"}},"_tags":["story","author_CarolineW","story_13387561"],"author":"CarolineW","created_at":"2017-01-13T00:53:26Z","created_at_i":1484268806,"num_comments":0,"objectID":"13387561","points":2,"story_id":13387561,"title":"Q: In base ten 1=0.999\u2026, but what about in other bases? What about in base 1?","updated_at":"2024-09-20T00:17:41Z","url":"http://www.askamathematician.com/2017/01/q-in-base-ten-10-999-but-what-about-in-other-bases-what-about-in-base-1/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Nvidia Invests $150M in AI Inference Startup Baseten"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.wsj.com/tech/ai/nvidia-invests-150-million-in-ai-inference-startup-baseten-fe7ede72"}},"_tags":["story","author_philipkiely","story_46736087"],"author":"philipkiely","children":[46736178],"created_at":"2026-01-23T18:42:34Z","created_at_i":1769193754,"num_comments":1,"objectID":"46736087","points":1,"story_id":46736087,"title":"Nvidia Invests $150M in AI Inference Startup Baseten","updated_at":"2026-03-05T23:24:17Z","url":"https://www.wsj.com/tech/ai/nvidia-invests-150-million-in-ai-inference-startup-baseten-fe7ede72"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"agcat"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Inferless Joins Baseten"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/announcing-the-acquihire-of-inferless-by-baseten/"}},"_tags":["story","author_agcat","story_47039179"],"author":"agcat","created_at":"2026-02-16T19:28:11Z","created_at_i":1771270091,"num_comments":0,"objectID":"47039179","points":1,"story_id":47039179,"title":"Inferless Joins Baseten","updated_at":"2026-03-05T23:33:57Z","url":"https://www.baseten.co/blog/announcing-the-acquihire-of-inferless-by-baseten/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tuhins"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Introducing BaseTen \u2014 build machine-learning powered applications"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog"}},"_tags":["story","author_tuhins","story_27252811"],"author":"tuhins","created_at":"2021-05-23T06:16:12Z","created_at_i":1621750572,"num_comments":0,"objectID":"27252811","points":1,"story_id":27252811,"title":"Introducing BaseTen \u2014 build machine-learning powered applications","updated_at":"2024-09-20T08:37:47Z","url":"https://www.baseten.co/blog"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"aaronrelph"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"ChatLLaMA is an experimental chatbot interface for interacting with variants of Facebook's LLaMA. Currently, we support the 7 billion parameter variant that was fine-tuned on the Alpaca dataset. This early versions isn't as conversational as we'd like, but over the next week or so, we're planning on adding support for the 30 billion parameter variant, another variant fine-tuned on LAION's OpenAssistant dataset and more as we explore what this model is capable of.
If you want deploy your own instance is the model powering the chatbot and build something similar we've open sourced the Truss here: https://github.com/basetenlabs/alpaca-7b-truss
We'd love to hear any feedback you have. You can reach me on Twitter @aaronrelph or Abu (the engineer behind this) @aqaderb.
Disclaimer: We both work at Baseten. This was a weekend project. Not trying to shill anything; just want to build and share cool stuff."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: ChatLLaMA \u2013 A ChatGPT style chatbot for Facebook's LLaMA"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://chatllama.baseten.co/"}},"_tags":["story","author_aaronrelph","story_35258553","show_hn"],"author":"aaronrelph","children":[35258627,35258715,35258768,35258806,35258826,35258835,35258841,35258887,35258897,35258926,35258927,35258963,35259313,35259324,35259496,35259515,35259529,35259534,35259556,35259596,35259819,35259820,35259833,35259852,35260091,35260114,35260259,35260330,35260701,35261084,35261520,35261767,35261991,35262631,35262778,35262950,35263166,35263187,35263286,35263688,35264131,35264560,35264566,35265899,35265971,35266091,35266476,35266695,35266780,35266922,35267213,35267803,35268625,35269222,35270801,35282911,35407077,35449391,35449398],"created_at":"2023-03-22T09:07:28Z","created_at_i":1679476048,"num_comments":215,"objectID":"35258553","points":402,"story_id":35258553,"story_text":"ChatLLaMA is an experimental chatbot interface for interacting with variants of Facebook's LLaMA. Currently, we support the 7 billion parameter variant that was fine-tuned on the Alpaca dataset. This early versions isn't as conversational as we'd like, but over the next week or so, we're planning on adding support for the 30 billion parameter variant, another variant fine-tuned on LAION's OpenAssistant dataset and more as we explore what this model is capable of.
If you want deploy your own instance is the model powering the chatbot and build something similar we've open sourced the Truss here: https://github.com/basetenlabs/alpaca-7b-truss
We'd love to hear any feedback you have. You can reach me on Twitter @aaronrelph or Abu (the engineer behind this) @aqaderb.
Disclaimer: We both work at Baseten. This was a weekend project. Not trying to shill anything; just want to build and share cool stuff.","title":"Show HN: ChatLLaMA \u2013 A ChatGPT style chatbot for Facebook's LLaMA","updated_at":"2026-08-19T22:08:40Z","url":"https://chatllama.baseten.co/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"davidtsong"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Introducing embeds.ai: an embedding playground to compare how embedding models work on a real world use case (retrieval augmented generation for Wikipedia articles + Elad Gil's High growth handbook)
A few weeks ago, Shreyan and I were looking for an embedding model to use for RAG. We eventually came across the MTEB leaderboard, but we struggled to understand the benchmark scores.
We wanted a tool to test various embedding models with example queries on real-world datasets. After unsuccessfully looking for such a \u201cplayground\u201d, we decided to just build one ourselves!
We embedded HuggingFace\u2019s Simple Wikipedia dataset using @OpenAI, @Cohere, and 2 open-source models via @Baseten. We then stored the embeddings in @Supabase using pgvector. Finally, we built a web app using NextJS and deployed it on @Vercel.
Now we\u2019re hosting the playground for anyone to use for free, as well as open-sourcing our work so people can try evaluating other models, datasets, or indexes.
Learn more here in our full blog post here: https://shreyanjain.substack.com/p/announcing-embedding-batt...
And the repo is here: https://github.com/EGCap/playground
If you have other suggestions / pain points from working with embedding models, vector DBs, or RAG, or if you would like to collaborate on any of the above or unrelated projects, please reach out!\n@shreyanj98 @davidtsong on Twitter"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Playground for comparing embedding models on Wikipedia+book retrieval"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www.embeds.ai/"}},"_tags":["story","author_davidtsong","story_38075196","show_hn"],"author":"davidtsong","children":[38077927,38077959,38078023,38078027,38079552,38079746],"created_at":"2023-10-30T20:19:41Z","created_at_i":1698697181,"num_comments":11,"objectID":"38075196","points":5,"story_id":38075196,"story_text":"Introducing embeds.ai: an embedding playground to compare how embedding models work on a real world use case (retrieval augmented generation for Wikipedia articles + Elad Gil's High growth handbook)
A few weeks ago, Shreyan and I were looking for an embedding model to use for RAG. We eventually came across the MTEB leaderboard, but we struggled to understand the benchmark scores.
We wanted a tool to test various embedding models with example queries on real-world datasets. After unsuccessfully looking for such a \u201cplayground\u201d, we decided to just build one ourselves!
We embedded HuggingFace\u2019s Simple Wikipedia dataset using @OpenAI, @Cohere, and 2 open-source models via @Baseten. We then stored the embeddings in @Supabase using pgvector. Finally, we built a web app using NextJS and deployed it on @Vercel.
Now we\u2019re hosting the playground for anyone to use for free, as well as open-sourcing our work so people can try evaluating other models, datasets, or indexes.
Learn more here in our full blog post here: https://shreyanjain.substack.com/p/announcing-embedding-batt...
And the repo is here: https://github.com/EGCap/playground
If you have other suggestions / pain points from working with embedding models, vector DBs, or RAG, or if you would like to collaborate on any of the above or unrelated projects, please reach out!\n@shreyanj98 @davidtsong on Twitter","title":"Show HN: Playground for comparing embedding models on Wikipedia+book retrieval","updated_at":"2025-04-22T11:35:18Z","url":"https://www.embeds.ai/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"shdalex"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Hi HN! I\u2019ve been building startfa.st, a curated directory of AI, developer, and product tools.\nLink: https://startfa.st
Like many people building in AI/dev, I found myself drowning in new tools every day, everything from agents, deployment frameworks, auth platforms, workflow engines, model APIs, automation tools, design tools, security stacks, etc. There\u2019s constant novelty, but it\u2019s very hard to find signal. Most lists on the internet recycle the same names.
So I built startfa.st to solve my own discovery problem.
What startfa.st is
A fast, searchable, hand-curated index of tools across categories including:
Development Tools
AI Agents
Automation Tools
Security Tools
Integrations & APIs
No-Code
Video / Design / Content Tools
Learning Tools
Finance / Analytics
And more (18+ categories total)
Each product has:\n a short manual description\n 4\u20136 tags for filtering\n category placement\n highlights (popular, rising, underrated)\n clean link out to the tool
There are 500+ tools so far, and I add new ones daily.
Why I built it
A few reasons:
Every day 10 new AI tools launch, and 7 of them die\nEarly discovery is useful, but most places surface the same 20\u201330 tools.
Twitter/X lists are noisy\nThey\u2019re good for hype, bad for structured discovery.
DevTools & AI tooling are becoming deeply fragmented\nFor example, discovering tools like Inngest, Arcjet, Clerk, Temporal, Baseten, Trigger.dev, CrewAI, etc. requires knowing where to look.
OpenAI/Google/Anthropic overshadow everything\nI wanted a home where great tools still get visibility even if they aren\u2019t the top model vendors.
I needed it for my own workflow\nI use this daily to prototype, compare stacks, test tools, and build things.
So this is partly a personal tool that grew into something larger.
What\u2019s technically interesting
The entire collection is hand-curated (no scraping, no auto-import).
Every entry is manually categorized, no model hallucinations.
The UI is intentionally lightweight and fast (I want it to load instantly).
Tags + categories are normalized to let you filter by capability (\u201cLLM\u201d, \u201cauth\u201d, \u201cworkflows\u201d, \u201cvector search\u201d, \u201cRAG\u201d, \u201cdeployment\u201d, etc).
I\u2019m working on full-text search (across tags, categories, and descriptions).
I\u2019m considering a structured API if enough people want it.
Even though it\u2019s a simple concept, the hard part is editing and curation.
What\u2019s unique / different
No hype, no affiliate links, no auto-generated garbage.
Everything is manually vetted. If a tool is bad, buggy, or spammy, I skip it.
I highlight rising or underrated tools, not just the obvious majors.
It\u2019s extremely fast and minimal, no bloat.
Tools are categorized by what builders actually care about (auth, agents, workflows, LLM infra, deployment, etc.).
What\u2019s upcoming
Full-text search
Maker pages (like IndieHackers but cleaner)
Collections (e.g. \u201cAI video stack\u201d, \u201cTools for indie hackers\u201d, \u201cDevOps AI\u201d)
Trending / upvotes
Ability to filter by tech (Python, JS, Go, cloud, etc.)
A weekly digest of new tools
Public API
\u201cI use this\u201d badges (opt-in)
If any of these matter to you, I\u2019d love to know which ones to prioritize.
Feedback I\u2019m looking for from HN
Is the categorization useful? What\u2019s missing?
Should entries include pricing, screenshots, or feature lists?
Would a public API be valuable?
Are there tools I\u2019m overlooking that deserve inclusion?
Any UI/UX simplifications you\u2019d recommend?
I\u2019m happy to answer anything in the comments.
Link
Thanks for reading, and thanks in advance for any feedback.
Happy to iterate quickly based on what HN suggests."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Startfa.st \u2013 A curated, fast directory of AI, dev, and product tools"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://startfa.st"}},"_tags":["story","author_shdalex","story_45964138","show_hn"],"author":"shdalex","created_at":"2025-11-18T12:00:20Z","created_at_i":1763467220,"num_comments":0,"objectID":"45964138","points":5,"story_id":45964138,"story_text":"Hi HN! I\u2019ve been building startfa.st, a curated directory of AI, developer, and product tools.\nLink: https://startfa.st
Like many people building in AI/dev, I found myself drowning in new tools every day, everything from agents, deployment frameworks, auth platforms, workflow engines, model APIs, automation tools, design tools, security stacks, etc. There\u2019s constant novelty, but it\u2019s very hard to find signal. Most lists on the internet recycle the same names.
So I built startfa.st to solve my own discovery problem.
What startfa.st is
A fast, searchable, hand-curated index of tools across categories including:
Development Tools
AI Agents
Automation Tools
Security Tools
Integrations & APIs
No-Code
Video / Design / Content Tools
Learning Tools
Finance / Analytics
And more (18+ categories total)
Each product has:\n a short manual description\n 4\u20136 tags for filtering\n category placement\n highlights (popular, rising, underrated)\n clean link out to the tool
There are 500+ tools so far, and I add new ones daily.
Why I built it
A few reasons:
Every day 10 new AI tools launch, and 7 of them die\nEarly discovery is useful, but most places surface the same 20\u201330 tools.
Twitter/X lists are noisy\nThey\u2019re good for hype, bad for structured discovery.
DevTools & AI tooling are becoming deeply fragmented\nFor example, discovering tools like Inngest, Arcjet, Clerk, Temporal, Baseten, Trigger.dev, CrewAI, etc. requires knowing where to look.
OpenAI/Google/Anthropic overshadow everything\nI wanted a home where great tools still get visibility even if they aren\u2019t the top model vendors.
I needed it for my own workflow\nI use this daily to prototype, compare stacks, test tools, and build things.
So this is partly a personal tool that grew into something larger.
What\u2019s technically interesting
The entire collection is hand-curated (no scraping, no auto-import).
Every entry is manually categorized, no model hallucinations.
The UI is intentionally lightweight and fast (I want it to load instantly).
Tags + categories are normalized to let you filter by capability (\u201cLLM\u201d, \u201cauth\u201d, \u201cworkflows\u201d, \u201cvector search\u201d, \u201cRAG\u201d, \u201cdeployment\u201d, etc).
I\u2019m working on full-text search (across tags, categories, and descriptions).
I\u2019m considering a structured API if enough people want it.
Even though it\u2019s a simple concept, the hard part is editing and curation.
What\u2019s unique / different
No hype, no affiliate links, no auto-generated garbage.
Everything is manually vetted. If a tool is bad, buggy, or spammy, I skip it.
I highlight rising or underrated tools, not just the obvious majors.
It\u2019s extremely fast and minimal, no bloat.
Tools are categorized by what builders actually care about (auth, agents, workflows, LLM infra, deployment, etc.).
What\u2019s upcoming
Full-text search
Maker pages (like IndieHackers but cleaner)
Collections (e.g. \u201cAI video stack\u201d, \u201cTools for indie hackers\u201d, \u201cDevOps AI\u201d)
Trending / upvotes
Ability to filter by tech (Python, JS, Go, cloud, etc.)
A weekly digest of new tools
Public API
\u201cI use this\u201d badges (opt-in)
If any of these matter to you, I\u2019d love to know which ones to prioritize.
Feedback I\u2019m looking for from HN
Is the categorization useful? What\u2019s missing?
Should entries include pricing, screenshots, or feature lists?
Would a public API be valuable?
Are there tools I\u2019m overlooking that deserve inclusion?
Any UI/UX simplifications you\u2019d recommend?
I\u2019m happy to answer anything in the comments.
Link
Thanks for reading, and thanks in advance for any feedback.
Happy to iterate quickly based on what HN suggests.","title":"Show HN: Startfa.st \u2013 A curated, fast directory of AI, dev, and product tools","updated_at":"2026-03-05T23:01:58Z","url":"https://startfa.st"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"Taikhoom2010"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"OpenAI is a deeply mismanaged company. Most recently, a blog post that feels like it has got to have dozens of PR violations, responding to the Apple lawsuit.
More broadly, OpenAI\u2019s main problem is that it does not have a real competitive advantage. No real models do; the real differentiation is price; they will become commodities. The need to move upwards in the market in enterprise software is difficult since the current stack should remain the same, since sticking with Salesforce plus its new AI features is much simpler than the cost of switching to an AI-first startup version, which won\u2019t be around 3 to 5 years from now. So consumer hardware and trying to become the next Apple could work, except you would likely have to focus on people not in the Apple ecosystem because of the double-sided lock-in Apple has.
The opposite dynamic works in consumer and mobile, really. Apple with AI is a much worse experience than something AI-first built from the ground up. Especially considering the consumer market does not have the switching costs enterprise does. So clearly OpenAI has some kind of opportunity to be a disruptor, just not to Apple, likely a large chunk of Android users, based on what their new device is.
So while at first it may look like OpenAI is doing too much, it is a desperate attempt to create some kind of real, defensible business, and show some proof of such before they IPO. If they do not do so, public market investors will cut their throats, and they will see a valuation drop like no other, regardless of where we are in the capital cycle. If they delay IPO plans, it will become clear they are in shambles. They are far behind in the enterprise game, which is short-term anyway.
Their best bet is hardware; credit to Altman and co for realizing this, although Apple is a roadblock, which I hope, for the sake of competition, is removed. Hardware is interesting, and perhaps somewhat disruptable towards Apple since OpenAI would be vertically integrated in a way Apple cannot since they do not have their own models. The difficulty is whether the burn OpenAI is spending on multiple fronts can be sustained long enough to see some promise. No doubt the consumer app is big, in terms of users, yet there is not enough activity to create a large ad business. The reason for this is Google and AI overviews, which are a far better user experience. So perhaps consumers can be an expensive customer acquisition cost to jumpstart the hardware business in some way? Regardless, OpenAI does have the right strategy in terms of attempting to create something defensible and long-term.
Now, Anthropic having a more coherent strategy in the short term of focusing on enterprise from the start is better; it is still not defensible. Perhaps because of the fact that the cost of paying Anthropic is far too great and outweighs any form of customer captivity. Plus, an enterprise gets maximum leverage by training a model, ideally a cheap one, on its own data so it can get insights tailored to the enterprise specifically, something Anthropic cannot deliver. Anthropic\u2019s fall will likely be in line with the broader capital cycle.
All in all, both companies as it stands today are massively overvalued, even Anthropic with its sky-high revenue, which is not real considering the capital cycle and short-term interest/desperation of enterprises, especially when there is a less expensive, far more valuable way to implement AI through using open-source models and training them on your data. Application layer companies such as BaseTen and OpenRouter should benefit from building on top of these new commodities. OpenAI, as of now, has the only long-term viable strategy, which has a lot of execution risk, yet excites me about the future of OpenAI."},"title":{"matchLevel":"none","matchedWords":[],"value":"Do You Think OpenAI Is Apple Circa the 1980s?"}},"_tags":["story","author_Taikhoom2010","story_49178393","ask_hn"],"author":"Taikhoom2010","children":[49179330],"created_at":"2026-08-05T03:50:17Z","created_at_i":1785901817,"num_comments":1,"objectID":"49178393","points":2,"story_id":49178393,"story_text":"OpenAI is a deeply mismanaged company. Most recently, a blog post that feels like it has got to have dozens of PR violations, responding to the Apple lawsuit.
More broadly, OpenAI\u2019s main problem is that it does not have a real competitive advantage. No real models do; the real differentiation is price; they will become commodities. The need to move upwards in the market in enterprise software is difficult since the current stack should remain the same, since sticking with Salesforce plus its new AI features is much simpler than the cost of switching to an AI-first startup version, which won\u2019t be around 3 to 5 years from now. So consumer hardware and trying to become the next Apple could work, except you would likely have to focus on people not in the Apple ecosystem because of the double-sided lock-in Apple has.
The opposite dynamic works in consumer and mobile, really. Apple with AI is a much worse experience than something AI-first built from the ground up. Especially considering the consumer market does not have the switching costs enterprise does. So clearly OpenAI has some kind of opportunity to be a disruptor, just not to Apple, likely a large chunk of Android users, based on what their new device is.
So while at first it may look like OpenAI is doing too much, it is a desperate attempt to create some kind of real, defensible business, and show some proof of such before they IPO. If they do not do so, public market investors will cut their throats, and they will see a valuation drop like no other, regardless of where we are in the capital cycle. If they delay IPO plans, it will become clear they are in shambles. They are far behind in the enterprise game, which is short-term anyway.
Their best bet is hardware; credit to Altman and co for realizing this, although Apple is a roadblock, which I hope, for the sake of competition, is removed. Hardware is interesting, and perhaps somewhat disruptable towards Apple since OpenAI would be vertically integrated in a way Apple cannot since they do not have their own models. The difficulty is whether the burn OpenAI is spending on multiple fronts can be sustained long enough to see some promise. No doubt the consumer app is big, in terms of users, yet there is not enough activity to create a large ad business. The reason for this is Google and AI overviews, which are a far better user experience. So perhaps consumers can be an expensive customer acquisition cost to jumpstart the hardware business in some way? Regardless, OpenAI does have the right strategy in terms of attempting to create something defensible and long-term.
Now, Anthropic having a more coherent strategy in the short term of focusing on enterprise from the start is better; it is still not defensible. Perhaps because of the fact that the cost of paying Anthropic is far too great and outweighs any form of customer captivity. Plus, an enterprise gets maximum leverage by training a model, ideally a cheap one, on its own data so it can get insights tailored to the enterprise specifically, something Anthropic cannot deliver. Anthropic\u2019s fall will likely be in line with the broader capital cycle.
All in all, both companies as it stands today are massively overvalued, even Anthropic with its sky-high revenue, which is not real considering the capital cycle and short-term interest/desperation of enterprises, especially when there is a less expensive, far more valuable way to implement AI through using open-source models and training them on your data. Application layer companies such as BaseTen and OpenRouter should benefit from building on top of these new commodities. OpenAI, as of now, has the only long-term viable strategy, which has a lot of execution risk, yet excites me about the future of OpenAI.","title":"Do You Think OpenAI Is Apple Circa the 1980s?","updated_at":"2026-08-05T17:56:36Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mathi0750"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"We're building off past Qwen-Image-Edit Fine-Tune Fridays and pushing the limits for 1.2 million images. We'll showcase a real world example where we cut an inference budget for a customer by nearly $50k on just one of their runs using Baseten for compute. Here's the luma if you want to join the live stream:\nhttps://luma.com/fine-tuning-friday-10"},"title":{"matchLevel":"none","matchedWords":[],"value":"Optimizing Qwen-Image-Edit to Generate 1.2M Images"}},"_tags":["story","author_mathi0750","story_45673975","ask_hn"],"author":"mathi0750","created_at":"2025-10-22T19:27:09Z","created_at_i":1761161229,"num_comments":0,"objectID":"45673975","points":2,"story_id":45673975,"story_text":"We're building off past Qwen-Image-Edit Fine-Tune Fridays and pushing the limits for 1.2 million images. We'll showcase a real world example where we cut an inference budget for a customer by nearly $50k on just one of their runs using Baseten for compute. Here's the luma if you want to join the live stream:\nhttps://luma.com/fine-tuning-friday-10","title":"Optimizing Qwen-Image-Edit to Generate 1.2M Images","updated_at":"2026-03-05T22:52:58Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/sota-performance-for-gpt-oss-120b-on-nvidia-gpus/"}},"_tags":["story","author_philipkiely","story_44819968"],"author":"philipkiely","children":[44820608,44820656,44820747,44820778,44820925,44821066,44821329,44821372,44821465,44821466,44822195,44822202,44822522,44822814,44822835,44823840,44824410,44824676,44824828,44826583,44826873,44833050],"created_at":"2025-08-07T02:28:47Z","created_at_i":1754533727,"num_comments":175,"objectID":"44819968","points":247,"story_id":44819968,"title":"Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs","updated_at":"2026-07-04T05:47:39Z","url":"https://www.baseten.co/blog/sota-performance-for-gpt-oss-120b-on-nvidia-gpus/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"The efficient frontier of LLM inference"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/the-efficient-frontier-of-llm-inference/"}},"_tags":["story","author_philipkiely","story_49529898"],"author":"philipkiely","children":[49530191,49530312,49530416,49530461,49530493,49530498,49530673,49530878,49531519,49531708,49531734,49531833,49531915,49531957,49531974,49532533,49532679,49532746,49533043,49533089,49533359,49533547,49533767,49533787,49534515,49535371],"created_at":"2026-09-01T23:48:05Z","created_at_i":1788306485,"num_comments":46,"objectID":"49529898","points":154,"story_id":49529898,"title":"The efficient frontier of LLM inference","updated_at":"2026-09-07T04:53:12Z","url":"https://www.baseten.co/blog/the-efficient-frontier-of-llm-inference/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"varunshenoy"},"title":{"matchLevel":"none","matchedWords":[],"value":"A guide to open-source LLM inference and performance"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/llm-transformer-inference-guide/"}},"_tags":["story","author_varunshenoy","story_38354254"],"author":"varunshenoy","children":[38354537,38356134,38356594,38356916,38357194,38358443,38358571],"created_at":"2023-11-20T20:33:16Z","created_at_i":1700512396,"num_comments":14,"objectID":"38354254","points":113,"story_id":38354254,"title":"A guide to open-source LLM inference and performance","updated_at":"2024-09-20T15:41:13Z","url":"https://www.baseten.co/blog/llm-transformer-inference-guide/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Truss \u2013 Serve any ML model without boilerplate code"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://github.com/basetenlabs/truss"}},"_tags":["story","author_philipkiely","story_32277894","show_hn"],"author":"philipkiely","children":[32279113,32279670,32280263,32280525,32280661,32283008],"created_at":"2022-07-29T15:06:35Z","created_at_i":1659107195,"num_comments":9,"objectID":"32277894","points":68,"story_id":32277894,"title":"Show HN: Truss \u2013 Serve any ML model without boilerplate code","updated_at":"2024-09-20T11:43:30Z","url":"https://github.com/basetenlabs/truss"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tuhins"},"title":{"matchLevel":"none","matchedWords":[],"value":"DALL-E Mini \u2013 Generate images from a text prompt"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://app.baseten.co/apps/RqgR9PV/operator_views/VBnA4qp"}},"_tags":["story","author_tuhins","story_31699841"],"author":"tuhins","children":[31700991,31701243,31701417,31701487,31702100,31702206,31702606,31702999,31705296,31707732,31709522],"created_at":"2022-06-10T22:01:34Z","created_at_i":1654898494,"num_comments":22,"objectID":"31699841","points":52,"story_id":31699841,"title":"DALL-E Mini \u2013 Generate images from a text prompt","updated_at":"2024-09-20T11:23:46Z","url":"https://app.baseten.co/apps/RqgR9PV/operator_views/VBnA4qp"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"varunshenoy"},"title":{"matchLevel":"none","matchedWords":[],"value":"How we got Stable Diffusion XL inference to under 2 seconds"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/sdxl-inference-in-under-2-seconds-the-ultimate-guide-to-stable-diffusion-optimization#"}},"_tags":["story","author_varunshenoy","story_37343158"],"author":"varunshenoy","children":[37347805,37347859,37348866,37352526],"created_at":"2023-08-31T20:20:27Z","created_at_i":1693513227,"num_comments":5,"objectID":"37343158","points":51,"story_id":37343158,"title":"How we got Stable Diffusion XL inference to under 2 seconds","updated_at":"2024-09-20T15:01:03Z","url":"https://www.baseten.co/blog/sdxl-inference-in-under-2-seconds-the-ultimate-guide-to-stable-diffusion-optimization#"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Free Stable Diffusion 2.0 hosted interface"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://app.baseten.co/apps/VBlnMVP/operator_views/nBrd8zP"}},"_tags":["story","author_philipkiely","story_33736749","show_hn"],"author":"philipkiely","children":[33737355,33741450],"created_at":"2022-11-24T22:05:06Z","created_at_i":1669327506,"num_comments":2,"objectID":"33736749","points":25,"story_id":33736749,"title":"Show HN: Free Stable Diffusion 2.0 hosted interface","updated_at":"2024-09-20T12:36:10Z","url":"https://app.baseten.co/apps/VBlnMVP/operator_views/nBrd8zP"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"aqader"},"story_text":{"matchLevel":"none","matchedWords":[],"value":"Hey HN, excited to release Blueprint into public beta. Blueprint is a general fine-tuning and model serving API built for developers. Fine-tuning models is an extremely powerful way to improve performance on a specific task without needing to collect prohibitively large amounts of data. With Blueprint you can kick off fine-tuning jobs for various open source models like Stable Diffusion and soon Flan-T5 using a Python SDK.
We'll also deploy your fine-tuned models onto serverless GPUs so that you aren't paying for idle GPU time. We scale the models up when you need to serve requests and have put a ton of engineering work into faster cold starts. We'll also autoscale replicas of your model if your model is receiving a lot of traffic.
Give it a shot \u2014 every new account gets a few hours of GPU credits. For support and feedback, join our Discord here: https://discord.gg/9pcXqWgB3g"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Fine-tune generative models in 1 line of code"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://blueprint.baseten.co/"}},"_tags":["story","author_aqader","story_34985166","show_hn"],"author":"aqader","created_at":"2023-03-01T17:24:47Z","created_at_i":1677691487,"num_comments":0,"objectID":"34985166","points":16,"story_id":34985166,"story_text":"Hey HN, excited to release Blueprint into public beta. Blueprint is a general fine-tuning and model serving API built for developers. Fine-tuning models is an extremely powerful way to improve performance on a specific task without needing to collect prohibitively large amounts of data. With Blueprint you can kick off fine-tuning jobs for various open source models like Stable Diffusion and soon Flan-T5 using a Python SDK.
We'll also deploy your fine-tuned models onto serverless GPUs so that you aren't paying for idle GPU time. We scale the models up when you need to serve requests and have put a ton of engineering work into faster cold starts. We'll also autoscale replicas of your model if your model is receiving a lot of traffic.
Give it a shot \u2014 every new account gets a few hours of GPU credits. For support and feedback, join our Discord here: https://discord.gg/9pcXqWgB3g","title":"Show HN: Fine-tune generative models in 1 line of code","updated_at":"2024-09-20T13:31:02Z","url":"https://blueprint.baseten.co/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"The Math Behind TurboQuant"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/i-spent-31-hours-on-the-math-behind-turboquant-so-you-dont-have-to/"}},"_tags":["story","author_philipkiely","story_47537690"],"author":"philipkiely","children":[47537735,47537760],"created_at":"2026-03-27T00:35:17Z","created_at_i":1774571717,"num_comments":3,"objectID":"47537690","points":8,"story_id":47537690,"title":"The Math Behind TurboQuant","updated_at":"2026-04-24T20:22:00Z","url":"https://www.baseten.co/blog/i-spent-31-hours-on-the-math-behind-turboquant-so-you-dont-have-to/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Hosted Stable Diffusion Demo"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://app.baseten.co/apps/VqK2vYP/operator_views/pqvba2q"}},"_tags":["story","author_philipkiely","story_32584186"],"author":"philipkiely","created_at":"2022-08-24T19:03:07Z","created_at_i":1661367787,"num_comments":0,"objectID":"32584186","points":7,"story_id":32584186,"title":"Hosted Stable Diffusion Demo","updated_at":"2024-09-20T11:55:08Z","url":"https://app.baseten.co/apps/VqK2vYP/operator_views/pqvba2q"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"We built the fastest API for GLM-5.2 (280 TPS)"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/how-we-built-the-worlds-fastest-api-for-glm-52/"}},"_tags":["story","author_philipkiely","story_48638427"],"author":"philipkiely","children":[48639137],"created_at":"2026-06-23T00:17:42Z","created_at_i":1782173862,"num_comments":0,"objectID":"48638427","points":6,"story_id":48638427,"title":"We built the fastest API for GLM-5.2 (280 TPS)","updated_at":"2026-06-23T01:49:42Z","url":"https://www.baseten.co/blog/how-we-built-the-worlds-fastest-api-for-glm-52/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Serving four million Riffusion requests in two days"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/serving-four-million-riffusion-requests-in-two-days"}},"_tags":["story","author_philipkiely","story_34083864"],"author":"philipkiely","created_at":"2022-12-21T17:44:21Z","created_at_i":1671644661,"num_comments":0,"objectID":"34083864","points":5,"story_id":34083864,"title":"Serving four million Riffusion requests in two days","updated_at":"2024-09-20T12:55:20Z","url":"https://www.baseten.co/blog/serving-four-million-riffusion-requests-in-two-days"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Try it yourself: Speech to text with Whisper"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://app.baseten.co/apps/b0dgKEB/operator_views/mP7AXaB"}},"_tags":["story","author_philipkiely","story_33050376"],"author":"philipkiely","created_at":"2022-10-01T21:38:38Z","created_at_i":1664660318,"num_comments":0,"objectID":"33050376","points":5,"story_id":33050376,"title":"Try it yourself: Speech to text with Whisper","updated_at":"2024-09-20T12:13:14Z","url":"https://app.baseten.co/apps/b0dgKEB/operator_views/mP7AXaB"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"Philpax"},"title":{"matchLevel":"none","matchedWords":[],"value":"How we built the fastest API for GLM-5.2"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/how-we-built-the-worlds-fastest-api-for-glm-52/"}},"_tags":["story","author_Philpax","story_48666063"],"author":"Philpax","created_at":"2026-06-24T21:53:41Z","created_at_i":1782338021,"num_comments":0,"objectID":"48666063","points":4,"story_id":48666063,"title":"How we built the fastest API for GLM-5.2","updated_at":"2026-06-25T00:41:19Z","url":"https://www.baseten.co/blog/how-we-built-the-worlds-fastest-api-for-glm-52/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"DarenWatson"},"title":{"matchLevel":"none","matchedWords":[],"value":"Inference Engineering: A free book on the systems behind AI inference"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/inference-engineering/"}},"_tags":["story","author_DarenWatson","story_49364934"],"author":"DarenWatson","children":[49369323],"created_at":"2026-08-19T18:00:59Z","created_at_i":1787162459,"num_comments":1,"objectID":"49364934","points":3,"story_id":49364934,"title":"Inference Engineering: A free book on the systems behind AI inference","updated_at":"2026-08-20T14:11:29Z","url":"https://www.baseten.co/inference-engineering/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mikejulietbravo"},"title":{"matchLevel":"none","matchedWords":[],"value":"How to get GLM 5.2 to 280 tokens per second"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/how-we-built-the-worlds-fastest-api-for-glm-52/"}},"_tags":["story","author_mikejulietbravo","story_48644981"],"author":"mikejulietbravo","children":[48644982],"created_at":"2026-06-23T13:50:36Z","created_at_i":1782222636,"num_comments":1,"objectID":"48644981","points":3,"story_id":48644981,"title":"How to get GLM 5.2 to 280 tokens per second","updated_at":"2026-06-23T14:42:48Z","url":"https://www.baseten.co/blog/how-we-built-the-worlds-fastest-api-for-glm-52/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tuhins"},"title":{"matchLevel":"none","matchedWords":[],"value":"SDXL inference in under 2 seconds"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/sdxl-inference-in-under-2-seconds-the-ultimate-guide-to-stable-diffusion-optimization"}},"_tags":["story","author_tuhins","story_37342605"],"author":"tuhins","children":[37342859],"created_at":"2023-08-31T19:37:41Z","created_at_i":1693510661,"num_comments":1,"objectID":"37342605","points":3,"story_id":37342605,"title":"SDXL inference in under 2 seconds","updated_at":"2024-09-20T15:00:57Z","url":"https://www.baseten.co/blog/sdxl-inference-in-under-2-seconds-the-ultimate-guide-to-stable-diffusion-optimization"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"How We Built the Fastest Kimi K2.5 on Artificial Analysis"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/how-we-built-the-fastest-kimi-k2-5-on-artificial-analysis/"}},"_tags":["story","author_philipkiely","story_46976035"],"author":"philipkiely","created_at":"2026-02-11T15:20:54Z","created_at_i":1770823254,"num_comments":0,"objectID":"46976035","points":3,"story_id":46976035,"title":"How We Built the Fastest Kimi K2.5 on Artificial Analysis","updated_at":"2026-03-05T23:34:48Z","url":"https://www.baseten.co/blog/how-we-built-the-fastest-kimi-k2-5-on-artificial-analysis/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Deploying Stable Diffusion in Production Using Truss"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/deploying-stable-diffusion"}},"_tags":["story","author_philipkiely","story_32683667"],"author":"philipkiely","created_at":"2022-09-01T21:41:19Z","created_at_i":1662068479,"num_comments":0,"objectID":"32683667","points":3,"story_id":32683667,"title":"Deploying Stable Diffusion in Production Using Truss","updated_at":"2024-09-20T11:57:35Z","url":"https://www.baseten.co/blog/deploying-stable-diffusion"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tikkun"},"title":{"matchLevel":"none","matchedWords":[],"value":"Faster Mixtral inference with TensorRT-LLM and quantization"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/faster-mixtral-inference-with-tensorrt-llm-and-quantization/"}},"_tags":["story","author_tikkun","story_38786797"],"author":"tikkun","children":[38786818],"created_at":"2023-12-27T21:13:06Z","created_at_i":1703711586,"num_comments":1,"objectID":"38786797","points":2,"story_id":38786797,"title":"Faster Mixtral inference with TensorRT-LLM and quantization","updated_at":"2024-09-20T16:05:02Z","url":"https://www.baseten.co/blog/faster-mixtral-inference-with-tensorrt-llm-and-quantization/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"codeAligned"},"title":{"matchLevel":"none","matchedWords":[],"value":"Another ML infra startup focused on developers"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://sarahguo.com/blog/baseten"}},"_tags":["story","author_codeAligned","story_37513127"],"author":"codeAligned","children":[37513233],"created_at":"2023-09-14T18:51:05Z","created_at_i":1694717465,"num_comments":1,"objectID":"37513127","points":2,"story_id":37513127,"title":"Another ML infra startup focused on developers","updated_at":"2024-09-20T15:07:55Z","url":"https://sarahguo.com/blog/baseten"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ajhai"},"title":{"matchLevel":"none","matchedWords":[],"value":"Inference Engineering by Philip Kiely \u2013 Digital Download"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/inference-engineering/"}},"_tags":["story","author_ajhai","story_49354292"],"author":"ajhai","created_at":"2026-08-18T23:21:44Z","created_at_i":1787095304,"num_comments":0,"objectID":"49354292","points":2,"story_id":49354292,"title":"Inference Engineering by Philip Kiely \u2013 Digital Download","updated_at":"2026-08-19T01:32:23Z","url":"https://www.baseten.co/inference-engineering/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"We built a day-0 API for Kimi K3"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/how-to-build-a-day-zero-api-for-kimi-k3/"}},"_tags":["story","author_philipkiely","story_49070886"],"author":"philipkiely","created_at":"2026-07-27T15:16:48Z","created_at_i":1785165408,"num_comments":0,"objectID":"49070886","points":2,"story_id":49070886,"title":"We built a day-0 API for Kimi K3","updated_at":"2026-07-27T18:14:18Z","url":"https://www.baseten.co/blog/how-to-build-a-day-zero-api-for-kimi-k3/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"We built the new fastest API for GLM-5.2"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/how-we-built-the-new-fastest-api-for-glm-52/"}},"_tags":["story","author_philipkiely","story_49054194"],"author":"philipkiely","created_at":"2026-07-26T02:51:43Z","created_at_i":1785034303,"num_comments":0,"objectID":"49054194","points":2,"story_id":49054194,"title":"We built the new fastest API for GLM-5.2","updated_at":"2026-07-26T03:06:42Z","url":"https://www.baseten.co/blog/how-we-built-the-new-fastest-api-for-glm-52/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"simonpure"},"title":{"matchLevel":"none","matchedWords":[],"value":"Inference Engineering"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.com/inference-engineering/"}},"_tags":["story","author_simonpure","story_47136864"],"author":"simonpure","created_at":"2026-02-24T13:26:35Z","created_at_i":1771939595,"num_comments":0,"objectID":"47136864","points":2,"story_id":47136864,"title":"Inference Engineering","updated_at":"2026-03-05T23:36:18Z","url":"https://www.baseten.com/inference-engineering/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"story_text":{"matchLevel":"none","matchedWords":[],"value":"There is a ton of demand for inference, but there are relatively few engineers working in the space. This leaves novel, interesting, and deeply technical challenges left to solve at every level of the stack.
To make it easier for more engineers to learn about inference, I wrote a book that provides a survey of the dozens of technologies that work together to make inference possible, along with an introduction to the primary techniques for inference optimization as well as commentary on how those techniques apply across various modalities.
This book is completely free to download digitally, and I'll have print copies with me at various conferences + available to purchase once Amazon decides to approve my account.
I hope you find Inference Engineering useful! Am around to answer any questions."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Inference Engineering"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.com/inference-engineering/"}},"_tags":["story","author_philipkiely","story_47129787","show_hn"],"author":"philipkiely","created_at":"2026-02-23T22:16:30Z","created_at_i":1771884990,"num_comments":0,"objectID":"47129787","points":2,"story_id":47129787,"story_text":"There is a ton of demand for inference, but there are relatively few engineers working in the space. This leaves novel, interesting, and deeply technical challenges left to solve at every level of the stack.
To make it easier for more engineers to learn about inference, I wrote a book that provides a survey of the dozens of technologies that work together to make inference possible, along with an introduction to the primary techniques for inference optimization as well as commentary on how those techniques apply across various modalities.
This book is completely free to download digitally, and I'll have print copies with me at various conferences + available to purchase once Amazon decides to approve my account.
I hope you find Inference Engineering useful! Am around to answer any questions.","title":"Show HN: Inference Engineering","updated_at":"2026-03-05T23:35:56Z","url":"https://www.baseten.com/inference-engineering/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"swyx"},"title":{"matchLevel":"none","matchedWords":[],"value":"Everything you need to run Mission Critical Inference (with DeepSeek v3, SGLang)"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.latent.space/p/baseten"}},"_tags":["story","author_swyx","story_42753516"],"author":"swyx","created_at":"2025-01-19T04:13:21Z","created_at_i":1737260001,"num_comments":0,"objectID":"42753516","points":2,"story_id":42753516,"title":"Everything you need to run Mission Critical Inference (with DeepSeek v3, SGLang)","updated_at":"2025-01-19T04:29:54Z","url":"https://www.latent.space/p/baseten"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"How to double tokens per second for Llama 3 with Medusa"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/how-to-double-tokens-per-second-for-llama-3-with-medusa/"}},"_tags":["story","author_philipkiely","story_41300242"],"author":"philipkiely","created_at":"2024-08-20T14:09:50Z","created_at_i":1724162990,"num_comments":0,"objectID":"41300242","points":2,"story_id":41300242,"title":"How to double tokens per second for Llama 3 with Medusa","updated_at":"2024-09-20T17:34:20Z","url":"https://www.baseten.co/blog/how-to-double-tokens-per-second-for-llama-3-with-medusa/"}],"hitsPerPage":50,"nbHits":2355,"nbPages":20,"page":0,"params":"query=Baseten&tags=story&hitsPerPage=50&advancedSyntax=true&analyticsTags=backend","processingTimeMS":13,"processingTimingsMS":{"_request":{"roundTrip":27},"afterFetch":{"format":{"highlighting":1,"total":1},"merge":{"mergeLoop":{"prepareNextHit":2,"total":2},"total":2},"total":2},"fetch":{"query":6,"scanning":3,"total":10},"total":13},"query":"Baseten","serverTimeMS":15}