{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Hi HN,

Three months ago, I took a job at Baseten to help craft and document an application builder that lets data scientists build full-stack, production-ready applications around their ML models without worrying about containers, Flask, or React. From my first day, everyone was focused on what would happen today: opening up our public beta. I\u2019m super excited to see what you build with Baseten.

If you want to take Baseten for a full-speed test drive, follow along with this tutorial, where you can build and deploy an application in 20 minutes: https://docs.baseten.co/getting-started

While Baseten is built for data scientists and machine learning engineers, something I\u2019m particularly excited about that doesn\u2019t come up often when we talk about Baseten is how it also makes building with ML available to people like me with a general software engineering background but no real experience with ML. With our library of pre-trained models, you can build and deploy an application around models for tasks like sentiment analysis, image classification, and speech transcription. By building applications around pre-trained models, I\u2019ve gained a deeper understanding of the use cases, capabilities, and limitations of machine learning.

If you want to play around with some models and applications without signing up for an account yet, check out our gallery (https://baseten.co/gallery) and try the demo apps.

P.S. We are also hiring; I found Baseten from HN."},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Show HN: Baseten \u2013 Build ML-powered applications"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/"}},"_tags":["story","author_philipkiely","story_31169193","show_hn"],"author":"philipkiely","children":[31169213,31169382,31169384,31177381,31177475,31178534],"created_at":"2022-04-26T16:05:37Z","created_at_i":1650989137,"num_comments":11,"objectID":"31169193","points":112,"story_id":31169193,"story_text":"Hi HN,

Three months ago, I took a job at Baseten to help craft and document an application builder that lets data scientists build full-stack, production-ready applications around their ML models without worrying about containers, Flask, or React. From my first day, everyone was focused on what would happen today: opening up our public beta. I\u2019m super excited to see what you build with Baseten.

If you want to take Baseten for a full-speed test drive, follow along with this tutorial, where you can build and deploy an application in 20 minutes: https://docs.baseten.co/getting-started

While Baseten is built for data scientists and machine learning engineers, something I\u2019m particularly excited about that doesn\u2019t come up often when we talk about Baseten is how it also makes building with ML available to people like me with a general software engineering background but no real experience with ML. With our library of pre-trained models, you can build and deploy an application around models for tasks like sentiment analysis, image classification, and speech transcription. By building applications around pre-trained models, I\u2019ve gained a deeper understanding of the use cases, capabilities, and limitations of machine learning.

If you want to play around with some models and applications without signing up for an account yet, check out our gallery (https://baseten.co/gallery) and try the demo apps.

P.S. We are also hiring; I found Baseten from HN.","title":"Show HN: Baseten \u2013 Build ML-powered applications","updated_at":"2024-09-20T11:00:59Z","url":"https://www.baseten.co/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ingve"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Base Ten for Almost Everything"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://randomascii.wordpress.com/2016/02/13/base-ten-for-almost-everything/"}},"_tags":["story","author_ingve","story_11096088"],"author":"ingve","children":[11096433,11096685,11097141,11097160,11097999,11099353,11099483,11099623,11099643,11099762,11099855,11100010,11100432],"created_at":"2016-02-13T22:54:44Z","created_at_i":1455404084,"num_comments":82,"objectID":"11096088","points":51,"story_id":11096088,"title":"Base Ten for Almost Everything","updated_at":"2025-09-08T02:16:02Z","url":"https://randomascii.wordpress.com/2016/02/13/base-ten-for-almost-everything/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"Hooke"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Renaissance Science: the base ten number system"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://thonyc.wordpress.com/2021/05/05/renaissance-science-ix/"}},"_tags":["story","author_Hooke","story_27179804"],"author":"Hooke","children":[27194584,27194679,27196876],"created_at":"2021-05-17T03:26:56Z","created_at_i":1621222016,"num_comments":11,"objectID":"27179804","points":29,"story_id":27179804,"title":"Renaissance Science: the base ten number system","updated_at":"2024-09-20T08:36:57Z","url":"https://thonyc.wordpress.com/2021/05/05/renaissance-science-ix/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"sahillavingia"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"BaseTen: The fastest way to build ML-powered applications"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://baseten.co"}},"_tags":["story","author_sahillavingia","story_27224638"],"author":"sahillavingia","children":[27225224,27225448],"created_at":"2021-05-20T17:57:02Z","created_at_i":1621533422,"num_comments":4,"objectID":"27224638","points":20,"story_id":27224638,"title":"BaseTen: The fastest way to build ML-powered applications","updated_at":"2024-09-20T08:35:11Z","url":"https://baseten.co"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mikejulietbravo"},"story_text":{"matchLevel":"none","matchedWords":[],"value":"Launched today - happy to answer any and all questions!"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Show HN: Baseten Chains \u2013 Framework and SDK for Multi-Model AI Products"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/introducing-baseten-chains/"}},"_tags":["story","author_mikejulietbravo","story_40813005","show_hn"],"author":"mikejulietbravo","children":[40813045,40813434,40813436],"created_at":"2024-06-27T17:39:45Z","created_at_i":1719509985,"num_comments":5,"objectID":"40813005","points":9,"story_id":40813005,"story_text":"Launched today - happy to answer any and all questions!","title":"Show HN: Baseten Chains \u2013 Framework and SDK for Multi-Model AI Products","updated_at":"2024-09-20T17:19:39Z","url":"https://www.baseten.co/blog/introducing-baseten-chains/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mich5632"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"High performance client for Baseten.co"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://github.com/basetenlabs/truss/tree/main/baseten-performance-client"}},"_tags":["story","author_mich5632","story_44270214"],"author":"mich5632","children":[44270215],"created_at":"2025-06-13T16:58:21Z","created_at_i":1749833901,"num_comments":1,"objectID":"44270214","points":7,"story_id":44270214,"title":"High performance client for Baseten.co","updated_at":"2025-06-14T03:13:10Z","url":"https://github.com/basetenlabs/truss/tree/main/baseten-performance-client"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"kodablah"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Baseten raised a $1.5B Series F and achieved a $13B valuation"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/announcing-our-series-f/"}},"_tags":["story","author_kodablah","story_48632248"],"author":"kodablah","created_at":"2026-06-22T16:18:34Z","created_at_i":1782145114,"num_comments":0,"objectID":"48632248","points":5,"story_id":48632248,"title":"Baseten raised a $1.5B Series F and achieved a $13B valuation","updated_at":"2026-06-22T18:59:42Z","url":"https://www.baseten.co/blog/announcing-our-series-f/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"How BaseTen is using \u201cdocs as code\u201d"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://blog.baseten.co/docs-as-code/"}},"_tags":["story","author_philipkiely","story_30617631"],"author":"philipkiely","created_at":"2022-03-09T17:44:25Z","created_at_i":1646847865,"num_comments":0,"objectID":"30617631","points":5,"story_id":30617631,"title":"How BaseTen is using \u201cdocs as code\u201d","updated_at":"2024-09-20T10:39:18Z","url":"https://blog.baseten.co/docs-as-code/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Baseten raises $150M Series D at $2.15B"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://fortune.com/2025/09/05/exclusive-baseten-ai-inference-unicorn-raises-150-million-at-2-15-billion-valuation/"}},"_tags":["story","author_philipkiely","story_45139326"],"author":"philipkiely","children":[45140575],"created_at":"2025-09-05T14:57:30Z","created_at_i":1757084250,"num_comments":1,"objectID":"45139326","points":2,"story_id":45139326,"title":"Baseten raises $150M Series D at $2.15B","updated_at":"2026-03-05T22:39:20Z","url":"https://fortune.com/2025/09/05/exclusive-baseten-ai-inference-unicorn-raises-150-million-at-2-15-billion-valuation/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mikejulietbravo"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Open Source Inference Engine Baseten Raises $40M from IVP, Spark and Greylock"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/announcing-our-series-b/"}},"_tags":["story","author_mikejulietbravo","story_39710361"],"author":"mikejulietbravo","children":[39710371],"created_at":"2024-03-14T23:42:47Z","created_at_i":1710459767,"num_comments":1,"objectID":"39710361","points":2,"story_id":39710361,"title":"Open Source Inference Engine Baseten Raises $40M from IVP, Spark and Greylock","updated_at":"2024-09-20T16:34:13Z","url":"https://www.baseten.co/blog/announcing-our-series-b/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"todsacerdoti"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Target 1: Baseten"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.silares.com/targets/target-1-baseten"}},"_tags":["story","author_todsacerdoti","story_46751640"],"author":"todsacerdoti","created_at":"2026-01-25T07:28:30Z","created_at_i":1769326110,"num_comments":0,"objectID":"46751640","points":2,"story_id":46751640,"title":"Target 1: Baseten","updated_at":"2026-03-05T23:25:09Z","url":"https://www.silares.com/targets/target-1-baseten"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ollayf"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Serverless Infrastructure for AI apps \u2013 3x perf of baseten, 1/5 the cost"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www.hyperpodai.com"}},"_tags":["story","author_ollayf","story_44937085"],"author":"ollayf","children":[44937086],"created_at":"2025-08-18T03:24:16Z","created_at_i":1755487456,"num_comments":0,"objectID":"44937085","points":2,"story_id":44937085,"title":"Serverless Infrastructure for AI apps \u2013 3x perf of baseten, 1/5 the cost","updated_at":"2026-03-05T22:30:26Z","url":"https://www.hyperpodai.com"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"elmazout"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Baseten, inference platform, $75M Series C, $825M valuation"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://twitter.com/basetenco/status/1892259130540179863"}},"_tags":["story","author_elmazout","story_43109768"],"author":"elmazout","created_at":"2025-02-20T00:46:54Z","created_at_i":1740012414,"num_comments":0,"objectID":"43109768","points":2,"story_id":43109768,"title":"Baseten, inference platform, $75M Series C, $825M valuation","updated_at":"2025-02-20T00:54:26Z","url":"https://twitter.com/basetenco/status/1892259130540179863"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"CarolineW"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Q: In base ten 1=0.999\u2026, but what about in other bases? What about in base 1?"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"http://www.askamathematician.com/2017/01/q-in-base-ten-10-999-but-what-about-in-other-bases-what-about-in-base-1/"}},"_tags":["story","author_CarolineW","story_13387561"],"author":"CarolineW","created_at":"2017-01-13T00:53:26Z","created_at_i":1484268806,"num_comments":0,"objectID":"13387561","points":2,"story_id":13387561,"title":"Q: In base ten 1=0.999\u2026, but what about in other bases? What about in base 1?","updated_at":"2024-09-20T00:17:41Z","url":"http://www.askamathematician.com/2017/01/q-in-base-ten-10-999-but-what-about-in-other-bases-what-about-in-base-1/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Nvidia Invests $150M in AI Inference Startup Baseten"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.wsj.com/tech/ai/nvidia-invests-150-million-in-ai-inference-startup-baseten-fe7ede72"}},"_tags":["story","author_philipkiely","story_46736087"],"author":"philipkiely","children":[46736178],"created_at":"2026-01-23T18:42:34Z","created_at_i":1769193754,"num_comments":1,"objectID":"46736087","points":1,"story_id":46736087,"title":"Nvidia Invests $150M in AI Inference Startup Baseten","updated_at":"2026-03-05T23:24:17Z","url":"https://www.wsj.com/tech/ai/nvidia-invests-150-million-in-ai-inference-startup-baseten-fe7ede72"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"agcat"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Inferless Joins Baseten"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/announcing-the-acquihire-of-inferless-by-baseten/"}},"_tags":["story","author_agcat","story_47039179"],"author":"agcat","created_at":"2026-02-16T19:28:11Z","created_at_i":1771270091,"num_comments":0,"objectID":"47039179","points":1,"story_id":47039179,"title":"Inferless Joins Baseten","updated_at":"2026-03-05T23:33:57Z","url":"https://www.baseten.co/blog/announcing-the-acquihire-of-inferless-by-baseten/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tuhins"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Introducing BaseTen \u2014 build machine-learning powered applications"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog"}},"_tags":["story","author_tuhins","story_27252811"],"author":"tuhins","created_at":"2021-05-23T06:16:12Z","created_at_i":1621750572,"num_comments":0,"objectID":"27252811","points":1,"story_id":27252811,"title":"Introducing BaseTen \u2014 build machine-learning powered applications","updated_at":"2024-09-20T08:37:47Z","url":"https://www.baseten.co/blog"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"aaronrelph"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"ChatLLaMA is an experimental chatbot interface for interacting with variants of Facebook's LLaMA. Currently, we support the 7 billion parameter variant that was fine-tuned on the Alpaca dataset. This early versions isn't as conversational as we'd like, but over the next week or so, we're planning on adding support for the 30 billion parameter variant, another variant fine-tuned on LAION's OpenAssistant dataset and more as we explore what this model is capable of.

If you want deploy your own instance is the model powering the chatbot and build something similar we've open sourced the Truss here: https://github.com/basetenlabs/alpaca-7b-truss

We'd love to hear any feedback you have. You can reach me on Twitter @aaronrelph or Abu (the engineer behind this) @aqaderb.

Disclaimer: We both work at Baseten. This was a weekend project. Not trying to shill anything; just want to build and share cool stuff."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: ChatLLaMA \u2013 A ChatGPT style chatbot for Facebook's LLaMA"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://chatllama.baseten.co/"}},"_tags":["story","author_aaronrelph","story_35258553","show_hn"],"author":"aaronrelph","children":[35258627,35258715,35258768,35258806,35258826,35258835,35258841,35258887,35258897,35258926,35258927,35258963,35259313,35259324,35259496,35259515,35259529,35259534,35259556,35259596,35259819,35259820,35259833,35259852,35260091,35260114,35260259,35260330,35260701,35261084,35261520,35261767,35261991,35262631,35262778,35262950,35263166,35263187,35263286,35263688,35264131,35264560,35264566,35265899,35265971,35266091,35266476,35266695,35266780,35266922,35267213,35267803,35268625,35269222,35270801,35282911,35407077,35449391,35449398],"created_at":"2023-03-22T09:07:28Z","created_at_i":1679476048,"num_comments":215,"objectID":"35258553","points":402,"story_id":35258553,"story_text":"ChatLLaMA is an experimental chatbot interface for interacting with variants of Facebook's LLaMA. Currently, we support the 7 billion parameter variant that was fine-tuned on the Alpaca dataset. This early versions isn't as conversational as we'd like, but over the next week or so, we're planning on adding support for the 30 billion parameter variant, another variant fine-tuned on LAION's OpenAssistant dataset and more as we explore what this model is capable of.

If you want deploy your own instance is the model powering the chatbot and build something similar we've open sourced the Truss here: https://github.com/basetenlabs/alpaca-7b-truss

We'd love to hear any feedback you have. You can reach me on Twitter @aaronrelph or Abu (the engineer behind this) @aqaderb.

Disclaimer: We both work at Baseten. This was a weekend project. Not trying to shill anything; just want to build and share cool stuff.","title":"Show HN: ChatLLaMA \u2013 A ChatGPT style chatbot for Facebook's LLaMA","updated_at":"2025-03-07T10:04:54Z","url":"https://chatllama.baseten.co/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"davidtsong"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Introducing embeds.ai: an embedding playground to compare how embedding models work on a real world use case (retrieval augmented generation for Wikipedia articles + Elad Gil's High growth handbook)

A few weeks ago, Shreyan and I were looking for an embedding model to use for RAG. We eventually came across the MTEB leaderboard, but we struggled to understand the benchmark scores.

We wanted a tool to test various embedding models with example queries on real-world datasets. After unsuccessfully looking for such a \u201cplayground\u201d, we decided to just build one ourselves!

We embedded HuggingFace\u2019s Simple Wikipedia dataset using @OpenAI, @Cohere, and 2 open-source models via @Baseten. We then stored the embeddings in @Supabase using pgvector. Finally, we built a web app using NextJS and deployed it on @Vercel.

Now we\u2019re hosting the playground for anyone to use for free, as well as open-sourcing our work so people can try evaluating other models, datasets, or indexes.

Learn more here in our full blog post here: https://shreyanjain.substack.com/p/announcing-embedding-batt...

And the repo is here: https://github.com/EGCap/playground

If you have other suggestions / pain points from working with embedding models, vector DBs, or RAG, or if you would like to collaborate on any of the above or unrelated projects, please reach out!\n@shreyanj98 @davidtsong on Twitter"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Playground for comparing embedding models on Wikipedia+book retrieval"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www.embeds.ai/"}},"_tags":["story","author_davidtsong","story_38075196","show_hn"],"author":"davidtsong","children":[38077927,38077959,38078023,38078027,38079552,38079746],"created_at":"2023-10-30T20:19:41Z","created_at_i":1698697181,"num_comments":11,"objectID":"38075196","points":5,"story_id":38075196,"story_text":"Introducing embeds.ai: an embedding playground to compare how embedding models work on a real world use case (retrieval augmented generation for Wikipedia articles + Elad Gil's High growth handbook)

A few weeks ago, Shreyan and I were looking for an embedding model to use for RAG. We eventually came across the MTEB leaderboard, but we struggled to understand the benchmark scores.

We wanted a tool to test various embedding models with example queries on real-world datasets. After unsuccessfully looking for such a \u201cplayground\u201d, we decided to just build one ourselves!

We embedded HuggingFace\u2019s Simple Wikipedia dataset using @OpenAI, @Cohere, and 2 open-source models via @Baseten. We then stored the embeddings in @Supabase using pgvector. Finally, we built a web app using NextJS and deployed it on @Vercel.

Now we\u2019re hosting the playground for anyone to use for free, as well as open-sourcing our work so people can try evaluating other models, datasets, or indexes.

Learn more here in our full blog post here: https://shreyanjain.substack.com/p/announcing-embedding-batt...

And the repo is here: https://github.com/EGCap/playground

If you have other suggestions / pain points from working with embedding models, vector DBs, or RAG, or if you would like to collaborate on any of the above or unrelated projects, please reach out!\n@shreyanj98 @davidtsong on Twitter","title":"Show HN: Playground for comparing embedding models on Wikipedia+book retrieval","updated_at":"2025-04-22T11:35:18Z","url":"https://www.embeds.ai/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"shdalex"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"Hi HN! I\u2019ve been building startfa.st, a curated directory of AI, developer, and product tools.\nLink: https://startfa.st

Like many people building in AI/dev, I found myself drowning in new tools every day, everything from agents, deployment frameworks, auth platforms, workflow engines, model APIs, automation tools, design tools, security stacks, etc. There\u2019s constant novelty, but it\u2019s very hard to find signal. Most lists on the internet recycle the same names.

So I built startfa.st to solve my own discovery problem.

What startfa.st is

A fast, searchable, hand-curated index of tools across categories including:

Development Tools

AI Agents

Automation Tools

Security Tools

Integrations & APIs

No-Code

Video / Design / Content Tools

Learning Tools

Finance / Analytics

And more (18+ categories total)

Each product has:\n a short manual description\n 4\u20136 tags for filtering\n category placement\n highlights (popular, rising, underrated)\n clean link out to the tool

There are 500+ tools so far, and I add new ones daily.

Why I built it

A few reasons:

Every day 10 new AI tools launch, and 7 of them die\nEarly discovery is useful, but most places surface the same 20\u201330 tools.

Twitter/X lists are noisy\nThey\u2019re good for hype, bad for structured discovery.

DevTools & AI tooling are becoming deeply fragmented\nFor example, discovering tools like Inngest, Arcjet, Clerk, Temporal, Baseten, Trigger.dev, CrewAI, etc. requires knowing where to look.

OpenAI/Google/Anthropic overshadow everything\nI wanted a home where great tools still get visibility even if they aren\u2019t the top model vendors.

I needed it for my own workflow\nI use this daily to prototype, compare stacks, test tools, and build things.

So this is partly a personal tool that grew into something larger.

What\u2019s technically interesting

The entire collection is hand-curated (no scraping, no auto-import).

Every entry is manually categorized, no model hallucinations.

The UI is intentionally lightweight and fast (I want it to load instantly).

Tags + categories are normalized to let you filter by capability (\u201cLLM\u201d, \u201cauth\u201d, \u201cworkflows\u201d, \u201cvector search\u201d, \u201cRAG\u201d, \u201cdeployment\u201d, etc).

I\u2019m working on full-text search (across tags, categories, and descriptions).

I\u2019m considering a structured API if enough people want it.

Even though it\u2019s a simple concept, the hard part is editing and curation.

What\u2019s unique / different

No hype, no affiliate links, no auto-generated garbage.

Everything is manually vetted. If a tool is bad, buggy, or spammy, I skip it.

I highlight rising or underrated tools, not just the obvious majors.

It\u2019s extremely fast and minimal, no bloat.

Tools are categorized by what builders actually care about (auth, agents, workflows, LLM infra, deployment, etc.).

What\u2019s upcoming

Full-text search

Maker pages (like IndieHackers but cleaner)

Collections (e.g. \u201cAI video stack\u201d, \u201cTools for indie hackers\u201d, \u201cDevOps AI\u201d)

Trending / upvotes

Ability to filter by tech (Python, JS, Go, cloud, etc.)

A weekly digest of new tools

Public API

\u201cI use this\u201d badges (opt-in)

If any of these matter to you, I\u2019d love to know which ones to prioritize.

Feedback I\u2019m looking for from HN

Is the categorization useful? What\u2019s missing?

Should entries include pricing, screenshots, or feature lists?

Would a public API be valuable?

Are there tools I\u2019m overlooking that deserve inclusion?

Any UI/UX simplifications you\u2019d recommend?

I\u2019m happy to answer anything in the comments.

Link

https://startfa.st

Thanks for reading, and thanks in advance for any feedback.

Happy to iterate quickly based on what HN suggests."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Startfa.st \u2013 A curated, fast directory of AI, dev, and product tools"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://startfa.st"}},"_tags":["story","author_shdalex","story_45964138","show_hn"],"author":"shdalex","created_at":"2025-11-18T12:00:20Z","created_at_i":1763467220,"num_comments":0,"objectID":"45964138","points":5,"story_id":45964138,"story_text":"Hi HN! I\u2019ve been building startfa.st, a curated directory of AI, developer, and product tools.\nLink: https://startfa.st

Like many people building in AI/dev, I found myself drowning in new tools every day, everything from agents, deployment frameworks, auth platforms, workflow engines, model APIs, automation tools, design tools, security stacks, etc. There\u2019s constant novelty, but it\u2019s very hard to find signal. Most lists on the internet recycle the same names.

So I built startfa.st to solve my own discovery problem.

What startfa.st is

A fast, searchable, hand-curated index of tools across categories including:

Development Tools

AI Agents

Automation Tools

Security Tools

Integrations & APIs

No-Code

Video / Design / Content Tools

Learning Tools

Finance / Analytics

And more (18+ categories total)

Each product has:\n a short manual description\n 4\u20136 tags for filtering\n category placement\n highlights (popular, rising, underrated)\n clean link out to the tool

There are 500+ tools so far, and I add new ones daily.

Why I built it

A few reasons:

Every day 10 new AI tools launch, and 7 of them die\nEarly discovery is useful, but most places surface the same 20\u201330 tools.

Twitter/X lists are noisy\nThey\u2019re good for hype, bad for structured discovery.

DevTools & AI tooling are becoming deeply fragmented\nFor example, discovering tools like Inngest, Arcjet, Clerk, Temporal, Baseten, Trigger.dev, CrewAI, etc. requires knowing where to look.

OpenAI/Google/Anthropic overshadow everything\nI wanted a home where great tools still get visibility even if they aren\u2019t the top model vendors.

I needed it for my own workflow\nI use this daily to prototype, compare stacks, test tools, and build things.

So this is partly a personal tool that grew into something larger.

What\u2019s technically interesting

The entire collection is hand-curated (no scraping, no auto-import).

Every entry is manually categorized, no model hallucinations.

The UI is intentionally lightweight and fast (I want it to load instantly).

Tags + categories are normalized to let you filter by capability (\u201cLLM\u201d, \u201cauth\u201d, \u201cworkflows\u201d, \u201cvector search\u201d, \u201cRAG\u201d, \u201cdeployment\u201d, etc).

I\u2019m working on full-text search (across tags, categories, and descriptions).

I\u2019m considering a structured API if enough people want it.

Even though it\u2019s a simple concept, the hard part is editing and curation.

What\u2019s unique / different

No hype, no affiliate links, no auto-generated garbage.

Everything is manually vetted. If a tool is bad, buggy, or spammy, I skip it.

I highlight rising or underrated tools, not just the obvious majors.

It\u2019s extremely fast and minimal, no bloat.

Tools are categorized by what builders actually care about (auth, agents, workflows, LLM infra, deployment, etc.).

What\u2019s upcoming

Full-text search

Maker pages (like IndieHackers but cleaner)

Collections (e.g. \u201cAI video stack\u201d, \u201cTools for indie hackers\u201d, \u201cDevOps AI\u201d)

Trending / upvotes

Ability to filter by tech (Python, JS, Go, cloud, etc.)

A weekly digest of new tools

Public API

\u201cI use this\u201d badges (opt-in)

If any of these matter to you, I\u2019d love to know which ones to prioritize.

Feedback I\u2019m looking for from HN

Is the categorization useful? What\u2019s missing?

Should entries include pricing, screenshots, or feature lists?

Would a public API be valuable?

Are there tools I\u2019m overlooking that deserve inclusion?

Any UI/UX simplifications you\u2019d recommend?

I\u2019m happy to answer anything in the comments.

Link

https://startfa.st

Thanks for reading, and thanks in advance for any feedback.

Happy to iterate quickly based on what HN suggests.","title":"Show HN: Startfa.st \u2013 A curated, fast directory of AI, dev, and product tools","updated_at":"2026-03-05T23:01:58Z","url":"https://startfa.st"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mathi0750"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"We're building off past Qwen-Image-Edit Fine-Tune Fridays and pushing the limits for 1.2 million images. We'll showcase a real world example where we cut an inference budget for a customer by nearly $50k on just one of their runs using Baseten for compute. Here's the luma if you want to join the live stream:\nhttps://luma.com/fine-tuning-friday-10"},"title":{"matchLevel":"none","matchedWords":[],"value":"Optimizing Qwen-Image-Edit to Generate 1.2M Images"}},"_tags":["story","author_mathi0750","story_45673975","ask_hn"],"author":"mathi0750","created_at":"2025-10-22T19:27:09Z","created_at_i":1761161229,"num_comments":0,"objectID":"45673975","points":2,"story_id":45673975,"story_text":"We're building off past Qwen-Image-Edit Fine-Tune Fridays and pushing the limits for 1.2 million images. We'll showcase a real world example where we cut an inference budget for a customer by nearly $50k on just one of their runs using Baseten for compute. Here's the luma if you want to join the live stream:\nhttps://luma.com/fine-tuning-friday-10","title":"Optimizing Qwen-Image-Edit to Generate 1.2M Images","updated_at":"2026-03-05T22:52:58Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/sota-performance-for-gpt-oss-120b-on-nvidia-gpus/"}},"_tags":["story","author_philipkiely","story_44819968"],"author":"philipkiely","children":[44820608,44820656,44820747,44820778,44820925,44821066,44821329,44821372,44821465,44821466,44822195,44822202,44822522,44822814,44822835,44823840,44824410,44824676,44824828,44826583,44826873,44833050],"created_at":"2025-08-07T02:28:47Z","created_at_i":1754533727,"num_comments":175,"objectID":"44819968","points":247,"story_id":44819968,"title":"Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs","updated_at":"2026-07-04T05:47:39Z","url":"https://www.baseten.co/blog/sota-performance-for-gpt-oss-120b-on-nvidia-gpus/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"varunshenoy"},"title":{"matchLevel":"none","matchedWords":[],"value":"A guide to open-source LLM inference and performance"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/llm-transformer-inference-guide/"}},"_tags":["story","author_varunshenoy","story_38354254"],"author":"varunshenoy","children":[38354537,38356134,38356594,38356916,38357194,38358443,38358571],"created_at":"2023-11-20T20:33:16Z","created_at_i":1700512396,"num_comments":14,"objectID":"38354254","points":113,"story_id":38354254,"title":"A guide to open-source LLM inference and performance","updated_at":"2024-09-20T15:41:13Z","url":"https://www.baseten.co/blog/llm-transformer-inference-guide/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Truss \u2013 Serve any ML model without boilerplate code"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://github.com/basetenlabs/truss"}},"_tags":["story","author_philipkiely","story_32277894","show_hn"],"author":"philipkiely","children":[32279113,32279670,32280263,32280525,32280661,32283008],"created_at":"2022-07-29T15:06:35Z","created_at_i":1659107195,"num_comments":9,"objectID":"32277894","points":68,"story_id":32277894,"title":"Show HN: Truss \u2013 Serve any ML model without boilerplate code","updated_at":"2024-09-20T11:43:30Z","url":"https://github.com/basetenlabs/truss"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tuhins"},"title":{"matchLevel":"none","matchedWords":[],"value":"DALL-E Mini \u2013 Generate images from a text prompt"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://app.baseten.co/apps/RqgR9PV/operator_views/VBnA4qp"}},"_tags":["story","author_tuhins","story_31699841"],"author":"tuhins","children":[31700991,31701243,31701417,31701487,31702100,31702206,31702606,31702999,31705296,31707732,31709522],"created_at":"2022-06-10T22:01:34Z","created_at_i":1654898494,"num_comments":22,"objectID":"31699841","points":52,"story_id":31699841,"title":"DALL-E Mini \u2013 Generate images from a text prompt","updated_at":"2024-09-20T11:23:46Z","url":"https://app.baseten.co/apps/RqgR9PV/operator_views/VBnA4qp"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"varunshenoy"},"title":{"matchLevel":"none","matchedWords":[],"value":"How we got Stable Diffusion XL inference to under 2 seconds"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/sdxl-inference-in-under-2-seconds-the-ultimate-guide-to-stable-diffusion-optimization#"}},"_tags":["story","author_varunshenoy","story_37343158"],"author":"varunshenoy","children":[37347805,37347859,37348866,37352526],"created_at":"2023-08-31T20:20:27Z","created_at_i":1693513227,"num_comments":5,"objectID":"37343158","points":51,"story_id":37343158,"title":"How we got Stable Diffusion XL inference to under 2 seconds","updated_at":"2024-09-20T15:01:03Z","url":"https://www.baseten.co/blog/sdxl-inference-in-under-2-seconds-the-ultimate-guide-to-stable-diffusion-optimization#"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Free Stable Diffusion 2.0 hosted interface"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://app.baseten.co/apps/VBlnMVP/operator_views/nBrd8zP"}},"_tags":["story","author_philipkiely","story_33736749","show_hn"],"author":"philipkiely","children":[33737355,33741450],"created_at":"2022-11-24T22:05:06Z","created_at_i":1669327506,"num_comments":2,"objectID":"33736749","points":25,"story_id":33736749,"title":"Show HN: Free Stable Diffusion 2.0 hosted interface","updated_at":"2024-09-20T12:36:10Z","url":"https://app.baseten.co/apps/VBlnMVP/operator_views/nBrd8zP"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"aqader"},"story_text":{"matchLevel":"none","matchedWords":[],"value":"Hey HN, excited to release Blueprint into public beta. Blueprint is a general fine-tuning and model serving API built for developers. Fine-tuning models is an extremely powerful way to improve performance on a specific task without needing to collect prohibitively large amounts of data. With Blueprint you can kick off fine-tuning jobs for various open source models like Stable Diffusion and soon Flan-T5 using a Python SDK.

We'll also deploy your fine-tuned models onto serverless GPUs so that you aren't paying for idle GPU time. We scale the models up when you need to serve requests and have put a ton of engineering work into faster cold starts. We'll also autoscale replicas of your model if your model is receiving a lot of traffic.

Give it a shot \u2014 every new account gets a few hours of GPU credits. For support and feedback, join our Discord here: https://discord.gg/9pcXqWgB3g"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Fine-tune generative models in 1 line of code"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://blueprint.baseten.co/"}},"_tags":["story","author_aqader","story_34985166","show_hn"],"author":"aqader","created_at":"2023-03-01T17:24:47Z","created_at_i":1677691487,"num_comments":0,"objectID":"34985166","points":16,"story_id":34985166,"story_text":"Hey HN, excited to release Blueprint into public beta. Blueprint is a general fine-tuning and model serving API built for developers. Fine-tuning models is an extremely powerful way to improve performance on a specific task without needing to collect prohibitively large amounts of data. With Blueprint you can kick off fine-tuning jobs for various open source models like Stable Diffusion and soon Flan-T5 using a Python SDK.

We'll also deploy your fine-tuned models onto serverless GPUs so that you aren't paying for idle GPU time. We scale the models up when you need to serve requests and have put a ton of engineering work into faster cold starts. We'll also autoscale replicas of your model if your model is receiving a lot of traffic.

Give it a shot \u2014 every new account gets a few hours of GPU credits. For support and feedback, join our Discord here: https://discord.gg/9pcXqWgB3g","title":"Show HN: Fine-tune generative models in 1 line of code","updated_at":"2024-09-20T13:31:02Z","url":"https://blueprint.baseten.co/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"The Math Behind TurboQuant"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/i-spent-31-hours-on-the-math-behind-turboquant-so-you-dont-have-to/"}},"_tags":["story","author_philipkiely","story_47537690"],"author":"philipkiely","children":[47537735,47537760],"created_at":"2026-03-27T00:35:17Z","created_at_i":1774571717,"num_comments":3,"objectID":"47537690","points":8,"story_id":47537690,"title":"The Math Behind TurboQuant","updated_at":"2026-04-24T20:22:00Z","url":"https://www.baseten.co/blog/i-spent-31-hours-on-the-math-behind-turboquant-so-you-dont-have-to/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Hosted Stable Diffusion Demo"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://app.baseten.co/apps/VqK2vYP/operator_views/pqvba2q"}},"_tags":["story","author_philipkiely","story_32584186"],"author":"philipkiely","created_at":"2022-08-24T19:03:07Z","created_at_i":1661367787,"num_comments":0,"objectID":"32584186","points":7,"story_id":32584186,"title":"Hosted Stable Diffusion Demo","updated_at":"2024-09-20T11:55:08Z","url":"https://app.baseten.co/apps/VqK2vYP/operator_views/pqvba2q"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"We built the fastest API for GLM-5.2 (280 TPS)"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/how-we-built-the-worlds-fastest-api-for-glm-52/"}},"_tags":["story","author_philipkiely","story_48638427"],"author":"philipkiely","children":[48639137],"created_at":"2026-06-23T00:17:42Z","created_at_i":1782173862,"num_comments":0,"objectID":"48638427","points":6,"story_id":48638427,"title":"We built the fastest API for GLM-5.2 (280 TPS)","updated_at":"2026-06-23T01:49:42Z","url":"https://www.baseten.co/blog/how-we-built-the-worlds-fastest-api-for-glm-52/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Serving four million Riffusion requests in two days"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/serving-four-million-riffusion-requests-in-two-days"}},"_tags":["story","author_philipkiely","story_34083864"],"author":"philipkiely","created_at":"2022-12-21T17:44:21Z","created_at_i":1671644661,"num_comments":0,"objectID":"34083864","points":5,"story_id":34083864,"title":"Serving four million Riffusion requests in two days","updated_at":"2024-09-20T12:55:20Z","url":"https://www.baseten.co/blog/serving-four-million-riffusion-requests-in-two-days"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Try it yourself: Speech to text with Whisper"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://app.baseten.co/apps/b0dgKEB/operator_views/mP7AXaB"}},"_tags":["story","author_philipkiely","story_33050376"],"author":"philipkiely","created_at":"2022-10-01T21:38:38Z","created_at_i":1664660318,"num_comments":0,"objectID":"33050376","points":5,"story_id":33050376,"title":"Try it yourself: Speech to text with Whisper","updated_at":"2024-09-20T12:13:14Z","url":"https://app.baseten.co/apps/b0dgKEB/operator_views/mP7AXaB"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"Philpax"},"title":{"matchLevel":"none","matchedWords":[],"value":"How we built the fastest API for GLM-5.2"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/how-we-built-the-worlds-fastest-api-for-glm-52/"}},"_tags":["story","author_Philpax","story_48666063"],"author":"Philpax","created_at":"2026-06-24T21:53:41Z","created_at_i":1782338021,"num_comments":0,"objectID":"48666063","points":4,"story_id":48666063,"title":"How we built the fastest API for GLM-5.2","updated_at":"2026-06-25T00:41:19Z","url":"https://www.baseten.co/blog/how-we-built-the-worlds-fastest-api-for-glm-52/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mikejulietbravo"},"title":{"matchLevel":"none","matchedWords":[],"value":"How to get GLM 5.2 to 280 tokens per second"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/how-we-built-the-worlds-fastest-api-for-glm-52/"}},"_tags":["story","author_mikejulietbravo","story_48644981"],"author":"mikejulietbravo","children":[48644982],"created_at":"2026-06-23T13:50:36Z","created_at_i":1782222636,"num_comments":1,"objectID":"48644981","points":3,"story_id":48644981,"title":"How to get GLM 5.2 to 280 tokens per second","updated_at":"2026-06-23T14:42:48Z","url":"https://www.baseten.co/blog/how-we-built-the-worlds-fastest-api-for-glm-52/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tuhins"},"title":{"matchLevel":"none","matchedWords":[],"value":"SDXL inference in under 2 seconds"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/sdxl-inference-in-under-2-seconds-the-ultimate-guide-to-stable-diffusion-optimization"}},"_tags":["story","author_tuhins","story_37342605"],"author":"tuhins","children":[37342859],"created_at":"2023-08-31T19:37:41Z","created_at_i":1693510661,"num_comments":1,"objectID":"37342605","points":3,"story_id":37342605,"title":"SDXL inference in under 2 seconds","updated_at":"2024-09-20T15:00:57Z","url":"https://www.baseten.co/blog/sdxl-inference-in-under-2-seconds-the-ultimate-guide-to-stable-diffusion-optimization"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"How We Built the Fastest Kimi K2.5 on Artificial Analysis"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/how-we-built-the-fastest-kimi-k2-5-on-artificial-analysis/"}},"_tags":["story","author_philipkiely","story_46976035"],"author":"philipkiely","created_at":"2026-02-11T15:20:54Z","created_at_i":1770823254,"num_comments":0,"objectID":"46976035","points":3,"story_id":46976035,"title":"How We Built the Fastest Kimi K2.5 on Artificial Analysis","updated_at":"2026-03-05T23:34:48Z","url":"https://www.baseten.co/blog/how-we-built-the-fastest-kimi-k2-5-on-artificial-analysis/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Deploying Stable Diffusion in Production Using Truss"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/deploying-stable-diffusion"}},"_tags":["story","author_philipkiely","story_32683667"],"author":"philipkiely","created_at":"2022-09-01T21:41:19Z","created_at_i":1662068479,"num_comments":0,"objectID":"32683667","points":3,"story_id":32683667,"title":"Deploying Stable Diffusion in Production Using Truss","updated_at":"2024-09-20T11:57:35Z","url":"https://www.baseten.co/blog/deploying-stable-diffusion"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tikkun"},"title":{"matchLevel":"none","matchedWords":[],"value":"Faster Mixtral inference with TensorRT-LLM and quantization"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/faster-mixtral-inference-with-tensorrt-llm-and-quantization/"}},"_tags":["story","author_tikkun","story_38786797"],"author":"tikkun","children":[38786818],"created_at":"2023-12-27T21:13:06Z","created_at_i":1703711586,"num_comments":1,"objectID":"38786797","points":2,"story_id":38786797,"title":"Faster Mixtral inference with TensorRT-LLM and quantization","updated_at":"2024-09-20T16:05:02Z","url":"https://www.baseten.co/blog/faster-mixtral-inference-with-tensorrt-llm-and-quantization/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"codeAligned"},"title":{"matchLevel":"none","matchedWords":[],"value":"Another ML infra startup focused on developers"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://sarahguo.com/blog/baseten"}},"_tags":["story","author_codeAligned","story_37513127"],"author":"codeAligned","children":[37513233],"created_at":"2023-09-14T18:51:05Z","created_at_i":1694717465,"num_comments":1,"objectID":"37513127","points":2,"story_id":37513127,"title":"Another ML infra startup focused on developers","updated_at":"2024-09-20T15:07:55Z","url":"https://sarahguo.com/blog/baseten"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"simonpure"},"title":{"matchLevel":"none","matchedWords":[],"value":"Inference Engineering"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.com/inference-engineering/"}},"_tags":["story","author_simonpure","story_47136864"],"author":"simonpure","created_at":"2026-02-24T13:26:35Z","created_at_i":1771939595,"num_comments":0,"objectID":"47136864","points":2,"story_id":47136864,"title":"Inference Engineering","updated_at":"2026-03-05T23:36:18Z","url":"https://www.baseten.com/inference-engineering/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"story_text":{"matchLevel":"none","matchedWords":[],"value":"There is a ton of demand for inference, but there are relatively few engineers working in the space. This leaves novel, interesting, and deeply technical challenges left to solve at every level of the stack.

To make it easier for more engineers to learn about inference, I wrote a book that provides a survey of the dozens of technologies that work together to make inference possible, along with an introduction to the primary techniques for inference optimization as well as commentary on how those techniques apply across various modalities.

This book is completely free to download digitally, and I'll have print copies with me at various conferences + available to purchase once Amazon decides to approve my account.

I hope you find Inference Engineering useful! Am around to answer any questions."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Inference Engineering"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.com/inference-engineering/"}},"_tags":["story","author_philipkiely","story_47129787","show_hn"],"author":"philipkiely","created_at":"2026-02-23T22:16:30Z","created_at_i":1771884990,"num_comments":0,"objectID":"47129787","points":2,"story_id":47129787,"story_text":"There is a ton of demand for inference, but there are relatively few engineers working in the space. This leaves novel, interesting, and deeply technical challenges left to solve at every level of the stack.

To make it easier for more engineers to learn about inference, I wrote a book that provides a survey of the dozens of technologies that work together to make inference possible, along with an introduction to the primary techniques for inference optimization as well as commentary on how those techniques apply across various modalities.

This book is completely free to download digitally, and I'll have print copies with me at various conferences + available to purchase once Amazon decides to approve my account.

I hope you find Inference Engineering useful! Am around to answer any questions.","title":"Show HN: Inference Engineering","updated_at":"2026-03-05T23:35:56Z","url":"https://www.baseten.com/inference-engineering/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"swyx"},"title":{"matchLevel":"none","matchedWords":[],"value":"Everything you need to run Mission Critical Inference (with DeepSeek v3, SGLang)"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.latent.space/p/baseten"}},"_tags":["story","author_swyx","story_42753516"],"author":"swyx","created_at":"2025-01-19T04:13:21Z","created_at_i":1737260001,"num_comments":0,"objectID":"42753516","points":2,"story_id":42753516,"title":"Everything you need to run Mission Critical Inference (with DeepSeek v3, SGLang)","updated_at":"2025-01-19T04:29:54Z","url":"https://www.latent.space/p/baseten"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"How to double tokens per second for Llama 3 with Medusa"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/how-to-double-tokens-per-second-for-llama-3-with-medusa/"}},"_tags":["story","author_philipkiely","story_41300242"],"author":"philipkiely","created_at":"2024-08-20T14:09:50Z","created_at_i":1724162990,"num_comments":0,"objectID":"41300242","points":2,"story_id":41300242,"title":"How to double tokens per second for Llama 3 with Medusa","updated_at":"2024-09-20T17:34:20Z","url":"https://www.baseten.co/blog/how-to-double-tokens-per-second-for-llama-3-with-medusa/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mikejulietbravo"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Automatically Build Nvidia TRT-LLM Engines"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/automatic-llm-optimization-with-tensorrt-llm-engine-builder/"}},"_tags":["story","author_mikejulietbravo","story_41131541","show_hn"],"author":"mikejulietbravo","created_at":"2024-08-01T17:39:06Z","created_at_i":1722533946,"num_comments":0,"objectID":"41131541","points":2,"story_id":41131541,"title":"Show HN: Automatically Build Nvidia TRT-LLM Engines","updated_at":"2024-09-20T17:32:32Z","url":"https://www.baseten.co/blog/automatic-llm-optimization-with-tensorrt-llm-engine-builder/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"FP8: Efficient model inference with 8-bit floating point numbers"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/blog/fp8-efficient-model-inference-with-8-bit-floating-point-numbers/"}},"_tags":["story","author_philipkiely","story_39642457"],"author":"philipkiely","created_at":"2024-03-08T16:12:20Z","created_at_i":1709914340,"num_comments":0,"objectID":"39642457","points":2,"story_id":39642457,"title":"FP8: Efficient model inference with 8-bit floating point numbers","updated_at":"2024-09-20T16:36:56Z","url":"https://www.baseten.co/blog/fp8-efficient-model-inference-with-8-bit-floating-point-numbers/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tuhins"},"title":{"matchLevel":"none","matchedWords":[],"value":"Build a Lensa-like avatar-generation app in a weekend"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://github.com/basetenlabs/avatar-generation-app-tutorial"}},"_tags":["story","author_tuhins","story_35042991"],"author":"tuhins","created_at":"2023-03-06T16:32:56Z","created_at_i":1678120376,"num_comments":0,"objectID":"35042991","points":2,"story_id":35042991,"title":"Build a Lensa-like avatar-generation app in a weekend","updated_at":"2024-09-20T13:26:02Z","url":"https://github.com/basetenlabs/avatar-generation-app-tutorial"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"philipkiely"},"title":{"matchLevel":"none","matchedWords":[],"value":"Code generation interactive demo (Salesforce Codegen mono 2B)"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://app.baseten.co/apps/2BYwaPG/operator_views/2qRdxX0"}},"_tags":["story","author_philipkiely","story_31953748"],"author":"philipkiely","created_at":"2022-07-01T22:19:48Z","created_at_i":1656713988,"num_comments":0,"objectID":"31953748","points":2,"story_id":31953748,"title":"Code generation interactive demo (Salesforce Codegen mono 2B)","updated_at":"2024-09-20T11:30:02Z","url":"https://app.baseten.co/apps/2BYwaPG/operator_views/2qRdxX0"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tuhins"},"title":{"matchLevel":"none","matchedWords":[],"value":"Working at an early-stage company as an early-stage engineer"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://blog.baseten.co/new-grad-part-1"}},"_tags":["story","author_tuhins","story_29387234"],"author":"tuhins","created_at":"2021-11-30T00:27:30Z","created_at_i":1638232050,"num_comments":0,"objectID":"29387234","points":2,"story_id":29387234,"title":"Working at an early-stage company as an early-stage engineer","updated_at":"2024-09-20T09:55:50Z","url":"https://blog.baseten.co/new-grad-part-1"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"jxmorris12"},"title":{"matchLevel":"none","matchedWords":[],"value":"Continual learning and the post monolith AI era"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["baseten"],"value":"https://www.baseten.co/resources/research/continual-learning/#introduction"}},"_tags":["story","author_jxmorris12","story_46919092"],"author":"jxmorris12","created_at":"2026-02-06T22:36:01Z","created_at_i":1770417361,"num_comments":0,"objectID":"46919092","points":1,"story_id":46919092,"title":"Continual learning and the post monolith AI era","updated_at":"2026-03-05T23:31:40Z","url":"https://www.baseten.co/resources/research/continual-learning/#introduction"}],"hitsPerPage":50,"nbHits":610,"nbPages":13,"page":0,"params":"query=Baseten&tags=story&hitsPerPage=50&advancedSyntax=true&analyticsTags=backend","processingTimeMS":13,"processingTimingsMS":{"_request":{"queue":10,"roundTrip":22},"afterFetch":{"format":{"highlighting":2,"total":2},"merge":{"mergeLoop":{"prepareNextHit":2,"total":2},"total":3},"total":3},"fetch":{"query":4,"scanning":4,"total":9},"total":13},"query":"Baseten","serverTimeMS":27}