{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"karimf"},"story_text":{"matchLevel":"none","matchedWords":[],"value":"Related: https://news.ycombinator.com/item?id=47653752"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Show HN: Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/fikrikarim/parlor"}},"_tags":["story","author_karimf","story_47652007","show_hn"],"author":"karimf","children":[47657363,47657413,47658211,47658279,47658501,47658648,47659094,47659177,47659785,47660100,47660340,47661021,47661687,47661861,47662393,47662495,47665274,47666583,47666903,47668166,47668838,47671411,47746679],"created_at":"2026-04-05T17:53:19Z","created_at_i":1775411599,"num_comments":38,"objectID":"47652007","points":298,"story_id":47652007,"story_text":"Related: https://news.ycombinator.com/item?id=47653752","title":"Show HN: Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B","updated_at":"2026-06-03T00:35:39Z","url":"https://github.com/fikrikarim/parlor"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"teamchong"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Show HN: Prompt-to-Excalidraw demo with Gemma 4 E2B in the browser (3.1GB)"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://teamchong.github.io/turboquant-wasm/draw.html"}},"_tags":["story","author_teamchong","story_47823460","show_hn"],"author":"teamchong","children":[47823853,47823998,47824114,47824999,47825162,47825402,47825880,47826456,47826542,47826920,47826932,47827244,47827715,47828086,47828721,47829409,47830298,47832800,47832852,47835815,47840897],"created_at":"2026-04-19T11:17:27Z","created_at_i":1776597447,"num_comments":62,"objectID":"47823460","points":163,"story_id":47823460,"title":"Show HN: Prompt-to-Excalidraw demo with Gemma 4 E2B in the browser (3.1GB)","updated_at":"2026-05-01T10:32:09Z","url":"https://teamchong.github.io/turboquant-wasm/draw.html"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"wittydeveloper"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Replicating Cursor's Agent Mode with E2B and AgentKit"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"https://e2b.dev/blog/replicating-cursors-agent-mode-with-e2b-and-agentkit"}},"_tags":["story","author_wittydeveloper","story_43184240"],"author":"wittydeveloper","children":[43184264,43184638,43184841,43185253,43185330,43186087,43193532],"created_at":"2025-02-26T15:00:59Z","created_at_i":1740582059,"num_comments":11,"objectID":"43184240","points":36,"story_id":43184240,"title":"Replicating Cursor's Agent Mode with E2B and AgentKit","updated_at":"2026-01-27T13:35:12Z","url":"https://e2b.dev/blog/replicating-cursors-agent-mode-with-e2b-and-agentkit"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"vegnus"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Show HN: Local LLM code-generation with Gemma 4 e2B via JSON AST to Clojure"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/quadracollision/llmisp"}},"_tags":["story","author_vegnus","story_48196707","show_hn"],"author":"vegnus","created_at":"2026-05-19T17:53:20Z","created_at_i":1779213200,"num_comments":0,"objectID":"48196707","points":22,"story_id":48196707,"title":"Show HN: Local LLM code-generation with Gemma 4 e2B via JSON AST to Clojure","updated_at":"2026-05-23T03:00:11Z","url":"https://github.com/quadracollision/llmisp"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"victormustar"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Gemma 4 E2B running in-browser at 255 tok/s"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://huggingface.co/spaces/webml-community/gemma-4-webgpu-kernels"}},"_tags":["story","author_victormustar","story_48577195"],"author":"victormustar","created_at":"2026-06-17T21:30:25Z","created_at_i":1781731825,"num_comments":0,"objectID":"48577195","points":19,"story_id":48577195,"title":"Gemma 4 E2B running in-browser at 255 tok/s","updated_at":"2026-07-06T18:07:21Z","url":"https://huggingface.co/spaces/webml-community/gemma-4-webgpu-kernels"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ryansen"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Gemma 4 E2B inference in 700 lines of C"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/ryanssenn/gemma4.c"}},"_tags":["story","author_ryansen","story_49468286"],"author":"ryansen","created_at":"2026-08-27T17:30:42Z","created_at_i":1787851842,"num_comments":0,"objectID":"49468286","points":17,"story_id":49468286,"title":"Gemma 4 E2B inference in 700 lines of C","updated_at":"2026-08-28T07:22:57Z","url":"https://github.com/ryanssenn/gemma4.c"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mailharishin"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Show HN: I benchmarked Gemma 4 E2B \u2013 the 2B model beat the 12B on multi-turn"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"https://aiexplr.com/post/gemma-4-e2b-benchmark"}},"_tags":["story","author_mailharishin","story_47756892","show_hn"],"author":"mailharishin","children":[47766250],"created_at":"2026-04-13T19:39:54Z","created_at_i":1776109194,"num_comments":1,"objectID":"47756892","points":8,"story_id":47756892,"title":"Show HN: I benchmarked Gemma 4 E2B \u2013 the 2B model beat the 12B on multi-turn","updated_at":"2026-04-17T11:18:18Z","url":"https://aiexplr.com/post/gemma-4-e2b-benchmark"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"yukunqiu"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Over the past few months, as we scaled our internal AI Agents, we hit a dead end: Running LLM-generated arbitrary code in Docker is basically running naked on security due to container escape risks. But using full traditional VMs takes minutes to boot and eats too much memory to support high-density concurrency. We loved the developer experience of SaaS sandboxes on the market, but they are closed-source, expensive, and have too high a barrier to entry for self-hosting.

So, our team decided to build our own. After months of grinding, using RustVMM and KVM, we built a blazing-fast, ultra-lightweight secure sandbox service from the ground up: CubeSandbox. Today, we are officially open-sourcing it.

To balance security and performance, we stripped the underlying OS to the absolute extreme. Here\u2019s what it can do right now:

1. <60ms blazing-fast cold start: End-to-end latency is under 60ms, making it 2.5x to 50x faster than traditional secure sandbox solutions.

2. <5MB extreme memory footprint: Memory per instance is kept under 5MB. A single 96-vCPU physical machine can easily run 2,000+ sandboxes concurrently, reducing storage consumption by 90%.

3. Massive concurrency scheduling: Capable of spinning up hundreds of thousands of instances in minutes.

4. True kernel-level isolation: Every Agent gets its own dedicated Guest OS kernel.

5. Native E2B SDK compatibility: Just swap a single URL environment variable. Zero code changes required for smooth migration and hosting.

Also, a millisecond-level \u201csnapshot rollback\u201d feature is coming soon\u2026

Before opening the repo today, CubeSandbox has been running silently behind the scenes in Tencent Cloud, serving massive real-world AI Agent workloads in production. As we open-source it today, it is no longer a prototype, but battle-tested, production-ready infrastructure.

Today, we hand it over to the community. Because we believe that high-performance agent infrastructure shouldn\u2019t be exclusive to a few\u2014it belongs to every developer worldwide who demands ultimate security and freedom.

The project is still in its very early open-source stages, and we are really looking forward to your hardest critiques and architectural roasts. I\u2019ll be hanging out here all day to answer your questions. The source code and deployment guides are all in the README. Come play with it!\nhttps://github.com/TencentCloud/CubeSandbox"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Show HN: We built a <60ms, open-source alternative to E2B using RustVMM and KVM"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/TencentCloud/CubeSandbox"}},"_tags":["story","author_yukunqiu","story_47863430","show_hn"],"author":"yukunqiu","children":[47863548,47863925,47864075,47864086,47872519,47872556,47872602,47874948],"created_at":"2026-04-22T13:36:32Z","created_at_i":1776864992,"num_comments":2,"objectID":"47863430","points":7,"story_id":47863430,"story_text":"Over the past few months, as we scaled our internal AI Agents, we hit a dead end: Running LLM-generated arbitrary code in Docker is basically running naked on security due to container escape risks. But using full traditional VMs takes minutes to boot and eats too much memory to support high-density concurrency. We loved the developer experience of SaaS sandboxes on the market, but they are closed-source, expensive, and have too high a barrier to entry for self-hosting.

So, our team decided to build our own. After months of grinding, using RustVMM and KVM, we built a blazing-fast, ultra-lightweight secure sandbox service from the ground up: CubeSandbox. Today, we are officially open-sourcing it.

To balance security and performance, we stripped the underlying OS to the absolute extreme. Here\u2019s what it can do right now:

1. <60ms blazing-fast cold start: End-to-end latency is under 60ms, making it 2.5x to 50x faster than traditional secure sandbox solutions.

2. <5MB extreme memory footprint: Memory per instance is kept under 5MB. A single 96-vCPU physical machine can easily run 2,000+ sandboxes concurrently, reducing storage consumption by 90%.

3. Massive concurrency scheduling: Capable of spinning up hundreds of thousands of instances in minutes.

4. True kernel-level isolation: Every Agent gets its own dedicated Guest OS kernel.

5. Native E2B SDK compatibility: Just swap a single URL environment variable. Zero code changes required for smooth migration and hosting.

Also, a millisecond-level \u201csnapshot rollback\u201d feature is coming soon\u2026

Before opening the repo today, CubeSandbox has been running silently behind the scenes in Tencent Cloud, serving massive real-world AI Agent workloads in production. As we open-source it today, it is no longer a prototype, but battle-tested, production-ready infrastructure.

Today, we hand it over to the community. Because we believe that high-performance agent infrastructure shouldn\u2019t be exclusive to a few\u2014it belongs to every developer worldwide who demands ultimate security and freedom.

The project is still in its very early open-source stages, and we are really looking forward to your hardest critiques and architectural roasts. I\u2019ll be hanging out here all day to answer your questions. The source code and deployment guides are all in the README. Come play with it!\nhttps://github.com/TencentCloud/CubeSandbox","title":"Show HN: We built a <60ms, open-source alternative to E2B using RustVMM and KVM","updated_at":"2026-04-29T21:39:06Z","url":"https://github.com/TencentCloud/CubeSandbox"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mikerubini"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Show HN: AI Sandbox with Kata, Firecracker, Cloud HV (E2B, Daytona Alternative)"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www.cognitora.dev"}},"_tags":["story","author_mikerubini","story_44784059","show_hn"],"author":"mikerubini","children":[44784067,44784429],"created_at":"2025-08-04T10:42:22Z","created_at_i":1754304142,"num_comments":4,"objectID":"44784059","points":6,"story_id":44784059,"title":"Show HN: AI Sandbox with Kata, Firecracker, Cloud HV (E2B, Daytona Alternative)","updated_at":"2025-09-18T16:42:26Z","url":"https://www.cognitora.dev"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"rbilgil"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"How we run OpenCode in the cloud with E2B and Convex"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"https://codecloud.dev/blog/opencode-e2b-sandbox"}},"_tags":["story","author_rbilgil","story_47223367"],"author":"rbilgil","created_at":"2026-03-02T20:06:59Z","created_at_i":1772482019,"num_comments":0,"objectID":"47223367","points":4,"story_id":47223367,"title":"How we run OpenCode in the cloud with E2B and Convex","updated_at":"2026-03-05T23:40:04Z","url":"https://codecloud.dev/blog/opencode-e2b-sandbox"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mailharishin"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"A weekend with LoRA on Gemma 4 E2B: instrumenting what fine-tuning changes"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://aiexplr.com/post/fine-tuning-5b-code-assistant-three-lessons"}},"_tags":["story","author_mailharishin","story_47911663"],"author":"mailharishin","children":[47911664,47959064],"created_at":"2026-04-26T16:44:52Z","created_at_i":1777221892,"num_comments":0,"objectID":"47911663","points":3,"story_id":47911663,"title":"A weekend with LoRA on Gemma 4 E2B: instrumenting what fine-tuning changes","updated_at":"2026-04-30T14:36:36Z","url":"https://aiexplr.com/post/fine-tuning-5b-code-assistant-three-lessons"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ushakov"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"E2B \u2013 Run AI-Generated code in secure sandbox"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"https://e2b.dev"}},"_tags":["story","author_ushakov","story_41860413"],"author":"ushakov","children":[41862700],"created_at":"2024-10-16T15:43:58Z","created_at_i":1729093438,"num_comments":1,"objectID":"41860413","points":2,"story_id":41860413,"title":"E2B \u2013 Run AI-Generated code in secure sandbox","updated_at":"2024-10-16T19:13:53Z","url":"https://e2b.dev"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"geoctl"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Hello HN , Cordium is a FOSS, self-hosted, general-purpose sandbox platform that I've been working on for a long time now that is built on Kubernetes and Octelium https://github.com/octelium/octelium, my main work. Cordium can be used for various persistent/ephemeral long/short-lived workloads, including coding for developers with VSCode, Zed, etc. (i.e. self-hosted GitHub Codespaces alternative), AI agent tasks (i.e. FOSS alternative to AI sandbox products such as E2B, Daytona, etc.), CI/CD workloads (e.g. building and publishing Docker images etc.), and more importantly for secretless remote access to infrastructure for devs and automated workloads from within the sandboxes.

The key differentiator here for Cordium, in comparison with other dev environments and sandbox platforms, is that Cordium automatically provides identity-based, secretless secure access to resources or infrastructure (e.g. APIs, SSH, databases, k8s, etc.) without having to inject credentials (e.g. API keys/access tokens, SSH private keys, database passwords, etc.) into the sandbox where the upstream credential is held by the identity-aware proxy of the Octelium-protected resource outside the reach of the sandbox. The sandbox permissions and access to resources is determined via identity-based, L7-aware, pre-request access control through CEL/OPA policy-as-code rather than injected credentials inside the sandbox. In other words, Cordium isn't just meant as a runtime for isolated execution where filesystem, CPU, memory, storage, etc... are isolated and controlled, but more importantly meant for identity-based secure access to infrastructure and resources.

In short, Cordium is basically a genereal-purpose sandbox platform + a ZTNA/remote-access-VPN baked-in with unified identity management, L7-aware access control and visibility.

Cordium is a purely FOSS project under Apache 2.0 that's meant for self-hosting and there are no plans for a pro/SaaS/cloud/commercial version. It was developed initially as a remote development environment for Octelium users to access their resources via web-based terminals through reproducible remote sandboxes instead of having to install and run the Octelium CLI connectors on their own machines but over time it grew into a general-purpose sandbox platform that can be used for all kinds of persistent/ephemeral and short/long-lived tasks by developers or automated workloads. I also want to clarify that Cordium, while opensourced a few days ago, is not a new project, the development of the project dates back to 2022 (see the older in https://github.com/octelium/spaces) and it is already being used by a few organizations that use Octelium since last year. In other words, this is not a toy project and it's meant to be used in production even though it's not quite ready to be labeled v1.0 yet. Happy to answer any questions."},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Show HN: Cordium \u2013 FOSS self-hosted sandbox platform alt. Codespaces/E2B/Daytona"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/octelium/cordium"}},"_tags":["story","author_geoctl","story_48433738","show_hn"],"author":"geoctl","children":[48435363,48522751],"created_at":"2026-06-07T11:13:29Z","created_at_i":1780830809,"num_comments":0,"objectID":"48433738","points":2,"story_id":48433738,"story_text":"Hello HN , Cordium is a FOSS, self-hosted, general-purpose sandbox platform that I've been working on for a long time now that is built on Kubernetes and Octelium https://github.com/octelium/octelium, my main work. Cordium can be used for various persistent/ephemeral long/short-lived workloads, including coding for developers with VSCode, Zed, etc. (i.e. self-hosted GitHub Codespaces alternative), AI agent tasks (i.e. FOSS alternative to AI sandbox products such as E2B, Daytona, etc.), CI/CD workloads (e.g. building and publishing Docker images etc.), and more importantly for secretless remote access to infrastructure for devs and automated workloads from within the sandboxes.

The key differentiator here for Cordium, in comparison with other dev environments and sandbox platforms, is that Cordium automatically provides identity-based, secretless secure access to resources or infrastructure (e.g. APIs, SSH, databases, k8s, etc.) without having to inject credentials (e.g. API keys/access tokens, SSH private keys, database passwords, etc.) into the sandbox where the upstream credential is held by the identity-aware proxy of the Octelium-protected resource outside the reach of the sandbox. The sandbox permissions and access to resources is determined via identity-based, L7-aware, pre-request access control through CEL/OPA policy-as-code rather than injected credentials inside the sandbox. In other words, Cordium isn't just meant as a runtime for isolated execution where filesystem, CPU, memory, storage, etc... are isolated and controlled, but more importantly meant for identity-based secure access to infrastructure and resources.

In short, Cordium is basically a genereal-purpose sandbox platform + a ZTNA/remote-access-VPN baked-in with unified identity management, L7-aware access control and visibility.

Cordium is a purely FOSS project under Apache 2.0 that's meant for self-hosting and there are no plans for a pro/SaaS/cloud/commercial version. It was developed initially as a remote development environment for Octelium users to access their resources via web-based terminals through reproducible remote sandboxes instead of having to install and run the Octelium CLI connectors on their own machines but over time it grew into a general-purpose sandbox platform that can be used for all kinds of persistent/ephemeral and short/long-lived tasks by developers or automated workloads. I also want to clarify that Cordium, while opensourced a few days ago, is not a new project, the development of the project dates back to 2022 (see the older in https://github.com/octelium/spaces) and it is already being used by a few organizations that use Octelium since last year. In other words, this is not a toy project and it's meant to be used in production even though it's not quite ready to be labeled v1.0 yet. Happy to answer any questions.","title":"Show HN: Cordium \u2013 FOSS self-hosted sandbox platform alt. Codespaces/E2B/Daytona","updated_at":"2026-07-15T20:37:08Z","url":"https://github.com/octelium/cordium"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"sixhobbits"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Cloud Computers for Agents: Exe.dev vs. Sprites vs. Shellbox vs. E2B vs. Blaxel"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://techstackups.com/comparisons/cloud-computers-for-ai-agents/"}},"_tags":["story","author_sixhobbits","story_47890500"],"author":"sixhobbits","created_at":"2026-04-24T14:05:44Z","created_at_i":1777039544,"num_comments":0,"objectID":"47890500","points":2,"story_id":47890500,"title":"Cloud Computers for Agents: Exe.dev vs. Sprites vs. Shellbox vs. E2B vs. Blaxel","updated_at":"2026-04-28T19:58:02Z","url":"https://techstackups.com/comparisons/cloud-computers-for-ai-agents/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ushakov"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Docker and E2B: Building the Future of Trusted AI"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"https://www.docker.com/blog/docker-e2b-building-the-future-of-trusted-ai/"}},"_tags":["story","author_ushakov","story_45682658"],"author":"ushakov","created_at":"2025-10-23T15:02:03Z","created_at_i":1761231723,"num_comments":0,"objectID":"45682658","points":2,"story_id":45682658,"title":"Docker and E2B: Building the Future of Trusted AI","updated_at":"2026-03-05T22:53:33Z","url":"https://www.docker.com/blog/docker-e2b-building-the-future-of-trusted-ai/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"nkov47as"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"The Sandboxed Open-Source Agent that is 70% cheaper than E2B"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://coasty.ai:443/"}},"_tags":["story","author_nkov47as","story_47268582"],"author":"nkov47as","children":[47268583],"created_at":"2026-03-05T23:16:26Z","created_at_i":1772752586,"num_comments":1,"objectID":"47268582","points":1,"story_id":47268582,"title":"The Sandboxed Open-Source Agent that is 70% cheaper than E2B","updated_at":"2026-03-05T23:42:33Z","url":"https://coasty.ai:443/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ohstep23"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"I don\u2019t believe agents should \u201cthink\u201d without being allowed to check their work. So I built a CLI agent that can actually run code while it reasons.

Subconscious handles the reasoning and tool-calling while actively pruning context as it goes, so the agent runs faster, costs less, and has more tokens available for complex tasks.

Code runs in one of the fastest sandboxed environments available using E2B, keeping execution quick while your local machine stays clean, organized, and protected from prompt injection.

The result feels less like a language model and more like working with a tool that can verify its own work, stay disciplined, and generate real reports and charts.

Run it with:

npx @subcon/e2b-cli

If you care about reliable agents, CLIs, or running code safely without touching your local machine, this is worth a look. Especially if you are tired of deleting random .md files."},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Show HN: A CLI agent that runs code while it reasons (Subconscious and E2B)"}},"_tags":["story","author_ohstep23","story_46768999","show_hn"],"author":"ohstep23","created_at":"2026-01-26T17:51:02Z","created_at_i":1769449862,"num_comments":0,"objectID":"46768999","points":1,"story_id":46768999,"story_text":"I don\u2019t believe agents should \u201cthink\u201d without being allowed to check their work. So I built a CLI agent that can actually run code while it reasons.

Subconscious handles the reasoning and tool-calling while actively pruning context as it goes, so the agent runs faster, costs less, and has more tokens available for complex tasks.

Code runs in one of the fastest sandboxed environments available using E2B, keeping execution quick while your local machine stays clean, organized, and protected from prompt injection.

The result feels less like a language model and more like working with a tool that can verify its own work, stay disciplined, and generate real reports and charts.

Run it with:

npx @subcon/e2b-cli

If you care about reliable agents, CLIs, or running code safely without touching your local machine, this is worth a look. Especially if you are tired of deleting random .md files.","title":"Show HN: A CLI agent that runs code while it reasons (Subconscious and E2B)","updated_at":"2026-03-05T23:26:06Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"bhavaniravi"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"E2B has been my favorite Sandbox environment for the AI agents I build. On actively using it for the last few months, CLI-ing every time is a lot of work. So here is a VSCode Extension"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Show HN: VSCode Extension for E2B Sandbox"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"https://marketplace.visualstudio.com/items?itemName=bhavaniravi.e2b-sandbox-explorer"}},"_tags":["story","author_bhavaniravi","story_46735089","show_hn"],"author":"bhavaniravi","created_at":"2026-01-23T17:23:17Z","created_at_i":1769188997,"num_comments":0,"objectID":"46735089","points":1,"story_id":46735089,"story_text":"E2B has been my favorite Sandbox environment for the AI agents I build. On actively using it for the last few months, CLI-ing every time is a lot of work. So here is a VSCode Extension","title":"Show HN: VSCode Extension for E2B Sandbox","updated_at":"2026-03-05T23:24:14Z","url":"https://marketplace.visualstudio.com/items?itemName=bhavaniravi.e2b-sandbox-explorer"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ushakov"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Fragments by E2B, open-source Claude Artifacts for full-stack apps"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"https://github.com/e2b-dev/fragments"}},"_tags":["story","author_ushakov","story_41731748"],"author":"ushakov","created_at":"2024-10-03T15:38:17Z","created_at_i":1727969897,"num_comments":0,"objectID":"41731748","points":1,"story_id":41731748,"title":"Fragments by E2B, open-source Claude Artifacts for full-stack apps","updated_at":"2024-10-03T15:42:04Z","url":"https://github.com/e2b-dev/fragments"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"HenryNdubuaku"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Hey HN, Henry & Roman here from Cactus.

A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-trained to know when it's wrong. Every response comes with a confidence score between 0 and 1. Developers can accept the on-device when it's high, hand off to a bigger cloud model when it's low. By routing only 15-35% of queries to Gemini 3.1 Flash-Lite, Gemma-4-E2B matches Gemini 3.1 Flash-Lite on most benchmarks.

- ChartQA: 15-20%

- LibriSpeech: 25-30%

- MMBench, GigaSpeech, MMAU: 30-35%

- MMLU-Pro: 45-55%

We were always frustrated by the routing signals hybrid apps rely on: asking the model to rate itself in text (unreliable, and you're parsing prose), or token entropy heuristics (barely better than a coin flip in our tests). So we did mechanistic studies on small models, Gemma 4 particularly, and found the hidden state for different layers carry meaningful self-awareness signal for various situations.

SO we extended the model with a 68k params probe layer (LayerNorm, low-rank projection, attention pooling, small MLP head) reads one intermediate layer during decoding and predicts p(wrong); confidence = 1 - p(wrong), returned as structured data, never parsed out of the answer text.

Across 12 hold-out benchmarks spanning text, vision and audio, the probe averages 0.814 AUROC vs 0.549 for token entropy. The result that convinced us this is real: the probe was trained on zero audio data, yet scores 0.79-0.88 AUROC on four audio benchmarks where entropy is near-random or worse (0.32-0.52). It's reading a modality-independent correctness signal from the hidden state, not memorizing patterns from its training data.

We published all weights on HuggingFace and provide copy-pase codes to run it on Transformers, MLX, Llama.cpp or Cactus. With Ollama, vLLM, SGLang etc in the works. For llama.cpp we ship a patch series you compile in once (upstreaming is planned). The code is MIT licensed; Gemma model use remains subject to the Gemma terms.

GitHub: https://github.com/cactus-compute/cactus-hybrid

Weights: https://huggingface.co/collections/Cactus-Compute/cactus-hyb...

Some caveats:

- The probe scores single-sequence decoding only, up to the first 1024 generated tokens.

- Handoff works best when routing per task in a multi-step process, not per step.

- Hierarchical routing is still in the works: try on-device, then DeepSeek v4 Flash, before Fable/GPT5.5/Gemini/Muse/Grok.

- The technique is boutique for each model, we will share each weights as they roll out.

These issues are currently being tackled at Cactus and updated weights will be shipped directly into the HuggingFace collection and GitHub repository straight up. Please let us know your thoughts, it helps us find ways to improve the design progressively.

Thanks a million!"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/cactus-compute/cactus-hybrid"}},"_tags":["story","author_HenryNdubuaku","story_49010782","show_hn"],"author":"HenryNdubuaku","children":[49011476,49014892,49015274,49015412,49015640,49015646,49015694,49015777,49016689,49018377,49018808,49019286,49019400,49019611,49019984,49020816,49023018,49027579,49046044,49048870],"created_at":"2026-07-22T17:56:29Z","created_at_i":1784742989,"num_comments":44,"objectID":"49010782","points":191,"story_id":49010782,"story_text":"Hey HN, Henry & Roman here from Cactus.

A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-trained to know when it's wrong. Every response comes with a confidence score between 0 and 1. Developers can accept the on-device when it's high, hand off to a bigger cloud model when it's low. By routing only 15-35% of queries to Gemini 3.1 Flash-Lite, Gemma-4-E2B matches Gemini 3.1 Flash-Lite on most benchmarks.

- ChartQA: 15-20%

- LibriSpeech: 25-30%

- MMBench, GigaSpeech, MMAU: 30-35%

- MMLU-Pro: 45-55%

We were always frustrated by the routing signals hybrid apps rely on: asking the model to rate itself in text (unreliable, and you're parsing prose), or token entropy heuristics (barely better than a coin flip in our tests). So we did mechanistic studies on small models, Gemma 4 particularly, and found the hidden state for different layers carry meaningful self-awareness signal for various situations.

SO we extended the model with a 68k params probe layer (LayerNorm, low-rank projection, attention pooling, small MLP head) reads one intermediate layer during decoding and predicts p(wrong); confidence = 1 - p(wrong), returned as structured data, never parsed out of the answer text.

Across 12 hold-out benchmarks spanning text, vision and audio, the probe averages 0.814 AUROC vs 0.549 for token entropy. The result that convinced us this is real: the probe was trained on zero audio data, yet scores 0.79-0.88 AUROC on four audio benchmarks where entropy is near-random or worse (0.32-0.52). It's reading a modality-independent correctness signal from the hidden state, not memorizing patterns from its training data.

We published all weights on HuggingFace and provide copy-pase codes to run it on Transformers, MLX, Llama.cpp or Cactus. With Ollama, vLLM, SGLang etc in the works. For llama.cpp we ship a patch series you compile in once (upstreaming is planned). The code is MIT licensed; Gemma model use remains subject to the Gemma terms.

GitHub: https://github.com/cactus-compute/cactus-hybrid

Weights: https://huggingface.co/collections/Cactus-Compute/cactus-hyb...

Some caveats:

- The probe scores single-sequence decoding only, up to the first 1024 generated tokens.

- Handoff works best when routing per task in a multi-step process, not per step.

- Hierarchical routing is still in the works: try on-device, then DeepSeek v4 Flash, before Fable/GPT5.5/Gemini/Muse/Grok.

- The technique is boutique for each model, we will share each weights as they roll out.

These issues are currently being tackled at Cactus and updated weights will be shipped directly into the HuggingFace collection and GitHub repository straight up. Please let us know your thoughts, it helps us find ways to improve the design progressively.

Thanks a million!","title":"Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong","updated_at":"2026-08-01T08:13:49Z","url":"https://github.com/cactus-compute/cactus-hybrid"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"gustrigos"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Hey HN, We're Gus and Carlos from Runtime (https://runtm.com). We're building infra that lets your whole team (including non-engineers) ship with Claude Code, Codex, and other agents without engineering having to handhold every session.

After Mentum (YC S21) was acquired, I personally shipped 4 full-stack products in 3 months using coding agents. When I tried to roll the same workflow out to the rest of the team, it fell apart: Most PRs were unmergeable slop - Every repo required an engineer doing one-off local setup. - Skills and context lived in one person's head. - There was no safe way for a PM to touch a real codebase without risking a bad deploy or a secrets leak.

Carlos comes from building agentic reconciliation systems at Modern Treasury and had a similar experience when letting his support team use devin.

We ended up building internal background agent infra but it quickly became a nightmare to mantain and develop. We built Runtime so you don't have to do this kind of thing.

Runtime work like as follows. Engineering defines the context once: system instructions, skills, and scoped integrations installable via CLI, mise, npm, or any package manager. Then Runtime snapshots your full running environment including multi-service Docker Compose setups, Kafka, Redis, seeded DBs, so it comes up in milliseconds with every server already running.

We orchestrate across sandbox providers like E2B, Daytona, EC2 or self-hosted K8s depending on your setup. Secrets are injected through our managed proxy so they never touch the agent directly, and guardrails run at the infrastructure level: command allow/deny lists, network egress controls, and RBAC scoped per human and per agent. Every session also gets a shareable preview URL, so internal builds go from sandbox to the rest of the team without needing production access.

Runtime works with whichever agent your team already uses: Claude Code, Codex, Cursor, Copilot, Gemini, Devin. You can trigger sandboxes from our web app, CLI, Slack, Linear, GitHub, or API.

One of our customers built an on-call inspector that wires PagerDuty, Sentry, and their repo so when an alert fires, the agent finds the cause and opens a PR with a unit test before anyone gets paged. Another runs a finance agent in a private Slack channel pulling from Stripe, NetSuite, and Snowflake to run reconciliations in minutes with source rows attached.

A fintech unicorn and several YC scaleups are live on Runtime, including a few teams who had built similar infrastructure internally and handed it to us to take over.

The core is open source at https://github.com/runtm-ai/runtm. Hosted version is live at https://app.runtm.com, free tier included. We're charging a flat platform fee plus compute, no token markup.

Check our demo: https://www.youtube.com/watch?v=wLwj__aEEh4

We'd love to hear how you're thinking about the infra for letting more people across your org use coding agents without creating chaos!"},"title":{"matchLevel":"none","matchedWords":[],"value":"Launch HN: Runtime (YC P26) \u2013 Sandboxed coding agents for everyone on a team"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www.runtm.com/"}},"_tags":["story","author_gustrigos","story_48225040","launch_hn"],"author":"gustrigos","children":[48225950,48226222,48226448,48226658,48227200,48227224,48227408,48227469,48227607,48227611,48228001,48228619,48228715,48229525,48230821,48231141,48231595,48231655,48231759,48233249,48233300,48233461,48233706,48234697,48234824,48243345,48329413,48515532],"created_at":"2026-05-21T16:07:13Z","created_at_i":1779379633,"num_comments":30,"objectID":"48225040","points":103,"story_id":48225040,"story_text":"Hey HN, We're Gus and Carlos from Runtime (https://runtm.com). We're building infra that lets your whole team (including non-engineers) ship with Claude Code, Codex, and other agents without engineering having to handhold every session.

After Mentum (YC S21) was acquired, I personally shipped 4 full-stack products in 3 months using coding agents. When I tried to roll the same workflow out to the rest of the team, it fell apart: Most PRs were unmergeable slop - Every repo required an engineer doing one-off local setup. - Skills and context lived in one person's head. - There was no safe way for a PM to touch a real codebase without risking a bad deploy or a secrets leak.

Carlos comes from building agentic reconciliation systems at Modern Treasury and had a similar experience when letting his support team use devin.

We ended up building internal background agent infra but it quickly became a nightmare to mantain and develop. We built Runtime so you don't have to do this kind of thing.

Runtime work like as follows. Engineering defines the context once: system instructions, skills, and scoped integrations installable via CLI, mise, npm, or any package manager. Then Runtime snapshots your full running environment including multi-service Docker Compose setups, Kafka, Redis, seeded DBs, so it comes up in milliseconds with every server already running.

We orchestrate across sandbox providers like E2B, Daytona, EC2 or self-hosted K8s depending on your setup. Secrets are injected through our managed proxy so they never touch the agent directly, and guardrails run at the infrastructure level: command allow/deny lists, network egress controls, and RBAC scoped per human and per agent. Every session also gets a shareable preview URL, so internal builds go from sandbox to the rest of the team without needing production access.

Runtime works with whichever agent your team already uses: Claude Code, Codex, Cursor, Copilot, Gemini, Devin. You can trigger sandboxes from our web app, CLI, Slack, Linear, GitHub, or API.

One of our customers built an on-call inspector that wires PagerDuty, Sentry, and their repo so when an alert fires, the agent finds the cause and opens a PR with a unit test before anyone gets paged. Another runs a finance agent in a private Slack channel pulling from Stripe, NetSuite, and Snowflake to run reconciliations in minutes with source rows attached.

A fintech unicorn and several YC scaleups are live on Runtime, including a few teams who had built similar infrastructure internally and handed it to us to take over.

The core is open source at https://github.com/runtm-ai/runtm. Hosted version is live at https://app.runtm.com, free tier included. We're charging a flat platform fee plus compute, no token markup.

Check our demo: https://www.youtube.com/watch?v=wLwj__aEEh4

We'd love to hear how you're thinking about the infra for letting more people across your org use coding agents without creating chaos!","title":"Launch HN: Runtime (YC P26) \u2013 Sandboxed coding agents for everyone on a team","updated_at":"2026-08-25T05:02:19Z","url":"https://www.runtm.com/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"NathanFlurry"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"We\u2019ve been working with automating coding agents in sandboxes as of late. It\u2019s bewildering how poorly standardized and difficult to use each agent varies between each other.

We open-sourced the Sandbox Agent SDK based on tools we built internally to solve 3 problems:

1. Universal agent API: interact with any coding agent using the same API

2. Running agents inside the sandbox: Agent Sandbox provides a Rust binary that serves the universal agent API over HTTP, instead of having to futz with undocumented interfaces

3. Universal session schema: persisting sessions is always problematic, since we don\u2019t want the source of truth for the conversation to live inside the container in a schema we don\u2019t control

Agent Sandbox SDK has:

- Any coding agent: Universal API to interact with all agents with full feature coverage

- Server or SDK mode: Run as an HTTP server or with the TypeScript SDK

- Universal session schema: Universal schema to store agent transcripts

- Supports your sandbox provider: Daytona, E2B, Vercel Sandboxes, and more

- Lightweight, portable Rust binary: Install anywhere with 1 curl command

- OpenAPI spec: Well documented and easy to integrate

We will be adding much more in the coming weeks \u2013 would love to hear any feedback or questions."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Sandbox Agent SDK \u2013 unified API for automating coding agents"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/rivet-dev/sandbox-agent"}},"_tags":["story","author_NathanFlurry","story_46795584","show_hn"],"author":"NathanFlurry","children":[46852816,46853121,46853918,46856280],"created_at":"2026-01-28T14:06:48Z","created_at_i":1769609208,"num_comments":7,"objectID":"46795584","points":41,"story_id":46795584,"story_text":"We\u2019ve been working with automating coding agents in sandboxes as of late. It\u2019s bewildering how poorly standardized and difficult to use each agent varies between each other.

We open-sourced the Sandbox Agent SDK based on tools we built internally to solve 3 problems:

1. Universal agent API: interact with any coding agent using the same API

2. Running agents inside the sandbox: Agent Sandbox provides a Rust binary that serves the universal agent API over HTTP, instead of having to futz with undocumented interfaces

3. Universal session schema: persisting sessions is always problematic, since we don\u2019t want the source of truth for the conversation to live inside the container in a schema we don\u2019t control

Agent Sandbox SDK has:

- Any coding agent: Universal API to interact with all agents with full feature coverage

- Server or SDK mode: Run as an HTTP server or with the TypeScript SDK

- Universal session schema: Universal schema to store agent transcripts

- Supports your sandbox provider: Daytona, E2B, Vercel Sandboxes, and more

- Lightweight, portable Rust binary: Install anywhere with 1 curl command

- OpenAPI spec: Well documented and easy to integrate

We will be adding much more in the coming weeks \u2013 would love to hear any feedback or questions.","title":"Show HN: Sandbox Agent SDK \u2013 unified API for automating coding agents","updated_at":"2026-04-03T09:39:20Z","url":"https://github.com/rivet-dev/sandbox-agent"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ankit219"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"GitHub: https://github.com/ClioAI/kw-sdk

Most AI agent frameworks target code. Write code, run tests, fix errors, repeat. That works because code has a natural verification signal. It works or it doesn't.

This SDK treats knowledge work like an engineering problem:

Task \u2192 Brief \u2192 Rubric (hidden from executor) \u2192 Work \u2192 Verify \u2192 Fail? \u2192 Retry \u2192 Pass \u2192 Submit

The orchestrator coordinates subagents, web search, code execution, and file I/O. then checks its own work against criteria it can't game (the rubric is generated in a separate call and the executor never sees it directly).

We originally built this as a harness for RL training on knowledge tasks. The rubric is the reward function. If you're training models on knowledge work, the brief\u2192rubric\u2192execute\u2192verify loop gives you a structured reward signal for tasks that normally don't have one.

What makes Knowledge work different from code? (apart from feedback loop)\nI believe there is some functionality missing from today's agents when it comes to knowledge work. I tried to include that in this release. Example:

Explore mode: Mapping the solution space, identifying the set level gaps, and giving options.

Most agents optimize for a single answer, and end up with a median one. For strategy, design, creative problems, you want to see the options, what are the tradeoffs, and what can you do? Explore mode generates N distinct approaches, each with explicit assumptions and counterfactuals ("this works if X, breaks if Y"). The output ends with set-level gaps ie what angles the entire set missed. The gaps are often more valuable than the takes. I think this is what many of us do on a daily basis, but no agent directly captures it today. See https://github.com/ClioAI/kw-sdk/blob/main/examples/explore_... and the output for a sense of how this is different.

Checkpointing: With many ai agents and especially multi agent systems, i can see where it went wrong, but cant run inference from same stage. (or you may want multiple explorations once an agent has done some tasks like search and is now looking at ideas). I used this for rollouts a lot, and think its a great feature to run again, or fork from a specific checkpoint.

A note on Verification loop:\nThe verify step is where the real leverage is. A model that can accurately assess its own work against a rubric is more valuable than one that generates slightly better first drafts. The rubric makes quality legible \u2014 to the agent, to the human, and potentially to a training signal.

Some things i like about this: \n- You can pass a remote execution environment (including your browser as a sandbox) and it would work. It can be docker, e2b, your local env, anything, the model will execute commands in your context, and will iterate based on feedback loop. Code execution is a protocol here.

- Tool calling: I realize you don't need complex functions. Models are good at writing terminal code, and can iterate based on feedback, so you can just pass either functions in context and model will execute or you can pass docs and model will write the code. (same as anthropic's programmatic tool calling). Details: https://github.com/ClioAI/kw-sdk/blob/main/TOOL_CALLING_GUID...

Lastly, some guides: \n- SDK guide: https://github.com/ClioAI/kw-sdk/blob/main/SDK_GUIDE.md\n- Extensible. See bizarro example where i add a new mode: https://github.com/ClioAI/kw-sdk/blob/main/examples/custom_m...\n- working with files: https://github.com/ClioAI/kw-sdk/blob/main/examples/with_fil... \n- this is simple but i love the csv example: https://github.com/ClioAI/kw-sdk/blob/main/examples/csv_rese...\n- remote execution: https://github.com/ClioAI/kw-sdk/blob/main/examples/with_cus...

And a lot more. This was completely refactored by opus and given the rework, probably would have taken a lot of time to release it.

MIT licensed. Would love your feedback."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Open-Source SDK for AI Knowledge Work"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/ClioAI/kw-sdk"}},"_tags":["story","author_ankit219","story_46963026","show_hn"],"author":"ankit219","children":[46963327],"created_at":"2026-02-10T17:06:00Z","created_at_i":1770743160,"num_comments":1,"objectID":"46963026","points":21,"story_id":46963026,"story_text":"GitHub: https://github.com/ClioAI/kw-sdk

Most AI agent frameworks target code. Write code, run tests, fix errors, repeat. That works because code has a natural verification signal. It works or it doesn't.

This SDK treats knowledge work like an engineering problem:

Task \u2192 Brief \u2192 Rubric (hidden from executor) \u2192 Work \u2192 Verify \u2192 Fail? \u2192 Retry \u2192 Pass \u2192 Submit

The orchestrator coordinates subagents, web search, code execution, and file I/O. then checks its own work against criteria it can't game (the rubric is generated in a separate call and the executor never sees it directly).

We originally built this as a harness for RL training on knowledge tasks. The rubric is the reward function. If you're training models on knowledge work, the brief\u2192rubric\u2192execute\u2192verify loop gives you a structured reward signal for tasks that normally don't have one.

What makes Knowledge work different from code? (apart from feedback loop)\nI believe there is some functionality missing from today's agents when it comes to knowledge work. I tried to include that in this release. Example:

Explore mode: Mapping the solution space, identifying the set level gaps, and giving options.

Most agents optimize for a single answer, and end up with a median one. For strategy, design, creative problems, you want to see the options, what are the tradeoffs, and what can you do? Explore mode generates N distinct approaches, each with explicit assumptions and counterfactuals ("this works if X, breaks if Y"). The output ends with set-level gaps ie what angles the entire set missed. The gaps are often more valuable than the takes. I think this is what many of us do on a daily basis, but no agent directly captures it today. See https://github.com/ClioAI/kw-sdk/blob/main/examples/explore_... and the output for a sense of how this is different.

Checkpointing: With many ai agents and especially multi agent systems, i can see where it went wrong, but cant run inference from same stage. (or you may want multiple explorations once an agent has done some tasks like search and is now looking at ideas). I used this for rollouts a lot, and think its a great feature to run again, or fork from a specific checkpoint.

A note on Verification loop:\nThe verify step is where the real leverage is. A model that can accurately assess its own work against a rubric is more valuable than one that generates slightly better first drafts. The rubric makes quality legible \u2014 to the agent, to the human, and potentially to a training signal.

Some things i like about this: \n- You can pass a remote execution environment (including your browser as a sandbox) and it would work. It can be docker, e2b, your local env, anything, the model will execute commands in your context, and will iterate based on feedback loop. Code execution is a protocol here.

- Tool calling: I realize you don't need complex functions. Models are good at writing terminal code, and can iterate based on feedback, so you can just pass either functions in context and model will execute or you can pass docs and model will write the code. (same as anthropic's programmatic tool calling). Details: https://github.com/ClioAI/kw-sdk/blob/main/TOOL_CALLING_GUID...

Lastly, some guides: \n- SDK guide: https://github.com/ClioAI/kw-sdk/blob/main/SDK_GUIDE.md\n- Extensible. See bizarro example where i add a new mode: https://github.com/ClioAI/kw-sdk/blob/main/examples/custom_m...\n- working with files: https://github.com/ClioAI/kw-sdk/blob/main/examples/with_fil... \n- this is simple but i love the csv example: https://github.com/ClioAI/kw-sdk/blob/main/examples/csv_rese...\n- remote execution: https://github.com/ClioAI/kw-sdk/blob/main/examples/with_cus...

And a lot more. This was completely refactored by opus and given the rework, probably would have taken a lot of time to release it.

MIT licensed. Would love your feedback.","title":"Show HN: Open-Source SDK for AI Knowledge Work","updated_at":"2026-03-05T23:34:07Z","url":"https://github.com/ClioAI/kw-sdk"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ATechGuy"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"In the last couple of months, several new solutions for sandboxing AI agents have launched (microVMs, WASM runtimes, browser isolation, hardened tool containers, etc.). Curious to hear from people using them in production. Are they working as advertised, or are there still major tradeoffs around security, cost, and performance?

Here's my list of sandboxing solutions launched in the last year alone: E2B, AIO Sandbox, Sandboxer, AgentSphere, Yolobox, Exe.dev, yolo-cage, SkillFS, ERA Jazzberry Computer, Vibekit, Daytona, Modal, Cognitora, YepCode, Run Compute, CLI Fence, Landrun, Sprites, pctx-sandbox, pctx Sandbox, Agent SDK, Lima-devbox, OpenServ, Browser Agent Playground, Flintlock Agent, Quickstart, Bouvet Sandbox, Arrakis, Cellmate (ceLLMate), AgentFence, Tasker, DenoSandbox, Capsule (WASM-based), Volant, Nono, NetFence"},"title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: The new wave of AI agent sandboxes?"}},"_tags":["story","author_ATechGuy","story_47444917","ask_hn"],"author":"ATechGuy","children":[47447412,47448449,47452369,47454677,47461589,47462125,47474500,47475562,47478137,47480110,47518400,47553356,47568556,47574159],"created_at":"2026-03-19T19:51:38Z","created_at_i":1773949898,"num_comments":5,"objectID":"47444917","points":12,"story_id":47444917,"story_text":"In the last couple of months, several new solutions for sandboxing AI agents have launched (microVMs, WASM runtimes, browser isolation, hardened tool containers, etc.). Curious to hear from people using them in production. Are they working as advertised, or are there still major tradeoffs around security, cost, and performance?

Here's my list of sandboxing solutions launched in the last year alone: E2B, AIO Sandbox, Sandboxer, AgentSphere, Yolobox, Exe.dev, yolo-cage, SkillFS, ERA Jazzberry Computer, Vibekit, Daytona, Modal, Cognitora, YepCode, Run Compute, CLI Fence, Landrun, Sprites, pctx-sandbox, pctx Sandbox, Agent SDK, Lima-devbox, OpenServ, Browser Agent Playground, Flintlock Agent, Quickstart, Bouvet Sandbox, Arrakis, Cellmate (ceLLMate), AgentFence, Tasker, DenoSandbox, Capsule (WASM-based), Volant, Nono, NetFence","title":"Ask HN: The new wave of AI agent sandboxes?","updated_at":"2026-07-13T14:42:45Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mlejva"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Hi everyone, I'm Vasek (https://x.com/mlejva), the CEO of the company behind this - https://e2b.dev. The company is called E2B. We're an open-source (https://github.com/e2b-dev) devtool that makes it easy to run untrusted AI-generated code in our secure sandboxes. You can think of us as coding runtime for LLMs.\nYou can self host us on GCP (https://github.com/e2b-dev/infra/blob/main/self-host.md) and we're working on AWS, then Azure, and any Linux machine.

This repo is one of our open-source projects that we're releasing to show developers what they can build with E2B. We used our sandboxes that are powered by AWS's Firecracker and gave them GUI with Linux. At the same time we made it easy to control this cloud computer with our Desktop SDK (https://github.com/e2b-dev/desktop). Essentially building a virtual desktop computer for AI and we gave it to LLMs to control it.\nHere's a demo we showed at an event - https://x.com/tereza_tizkova/status/1878834392891838556

We think computer use is still highly experimental but it feels sort of similar like AI codegen in early 2023. You see the sparks but it's not there just yet. However, we wanted to research if open-source LLMs could at least get some results. Here we're using Llama 3.2, Llama 3.3, and an OS-Atlas as a base model (fine-tuned Qwen).

If you have any questions, happy to answer them!

We're also hiring! If you're a fullstack engineer, distributed systems engineer, AI engineer, product designer, or GTM person and based in SF send me a hello to vasek @ e2b.dev!"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Open Computer Use"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"https://github.com/e2b-dev/open-computer-use"}},"_tags":["story","author_mlejva","story_42804933","show_hn"],"author":"mlejva","children":[42804962,42805050],"created_at":"2025-01-23T15:27:13Z","created_at_i":1737646033,"num_comments":1,"objectID":"42804933","points":9,"story_id":42804933,"story_text":"Hi everyone, I'm Vasek (https://x.com/mlejva), the CEO of the company behind this - https://e2b.dev. The company is called E2B. We're an open-source (https://github.com/e2b-dev) devtool that makes it easy to run untrusted AI-generated code in our secure sandboxes. You can think of us as coding runtime for LLMs.\nYou can self host us on GCP (https://github.com/e2b-dev/infra/blob/main/self-host.md) and we're working on AWS, then Azure, and any Linux machine.

This repo is one of our open-source projects that we're releasing to show developers what they can build with E2B. We used our sandboxes that are powered by AWS's Firecracker and gave them GUI with Linux. At the same time we made it easy to control this cloud computer with our Desktop SDK (https://github.com/e2b-dev/desktop). Essentially building a virtual desktop computer for AI and we gave it to LLMs to control it.\nHere's a demo we showed at an event - https://x.com/tereza_tizkova/status/1878834392891838556

We think computer use is still highly experimental but it feels sort of similar like AI codegen in early 2023. You see the sparks but it's not there just yet. However, we wanted to research if open-source LLMs could at least get some results. Here we're using Llama 3.2, Llama 3.3, and an OS-Atlas as a base model (fine-tuned Qwen).

If you have any questions, happy to answer them!

We're also hiring! If you're a fullstack engineer, distributed systems engineer, AI engineer, product designer, or GTM person and based in SF send me a hello to vasek @ e2b.dev!","title":"Show HN: Open Computer Use","updated_at":"2025-01-23T17:57:42Z","url":"https://github.com/e2b-dev/open-computer-use"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"willydouhard"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Hi HN,

AgentBox is an SDK for running coding agents (Claude Code, Codex, OpenCode) inside sandboxes (Docker, E2B, Modal, Daytona, Vercel). One API. Swap the agent or the sandbox and your code doesn't change.

Think of it as what the AI SDK did for LLMs, but for agent + runtime.

Most wrappers call agents in non interactive mode (claude --print, codex exec). AgentBox instead boots each agent's native server inside the sandbox (Codex app-server over JSON-RPC, OpenCode serve over HTTP/SSE, Claude Code's SDK WebSocket transport) and drives it from the host, behaving like an interactive session.

Live demo: https://agentbox-demo-175164121374.us-west1.run.app

Feedback welcome, especially on the abstractions and which providers to add next!"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: AgentBox \u2013 SDK to Run Claude Code, Codex, or OpenCode in Any Sandbox"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/TwillAI/agentbox-sdk"}},"_tags":["story","author_willydouhard","story_47876788","show_hn"],"author":"willydouhard","created_at":"2026-04-23T15:19:20Z","created_at_i":1776957560,"num_comments":0,"objectID":"47876788","points":8,"story_id":47876788,"story_text":"Hi HN,

AgentBox is an SDK for running coding agents (Claude Code, Codex, OpenCode) inside sandboxes (Docker, E2B, Modal, Daytona, Vercel). One API. Swap the agent or the sandbox and your code doesn't change.

Think of it as what the AI SDK did for LLMs, but for agent + runtime.

Most wrappers call agents in non interactive mode (claude --print, codex exec). AgentBox instead boots each agent's native server inside the sandbox (Codex app-server over JSON-RPC, OpenCode serve over HTTP/SSE, Claude Code's SDK WebSocket transport) and drives it from the host, behaving like an interactive session.

Live demo: https://agentbox-demo-175164121374.us-west1.run.app

Feedback welcome, especially on the abstractions and which providers to add next!","title":"Show HN: AgentBox \u2013 SDK to Run Claude Code, Codex, or OpenCode in Any Sandbox","updated_at":"2026-06-03T02:04:39Z","url":"https://github.com/TwillAI/agentbox-sdk"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ajaysheoran2323"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"I've been building Sandflare for the past few months \u2014 it launches Firecracker microVMs for AI agents in ~300ms cold start. The idea came from running LLM-generated code in production. Docker felt too risky (shared kernel), full VMs too slow (5\u201310s). Firecracker hits the middle: real VM isolation, fast boot.

I also added managed Postgres because almost every agent I built needed persistent state. One call wires a database into a sandbox.

There are great tools in this space already (E2B, Modal, Daytona) \u2014 I wanted something with batteries-included Postgres, and simpler pricing

What I'm trying to figure out: how do I get cold start below 100ms? Currently the bottleneck is the Firecracker API + network setup. Would love to hear from anyone who's pushed Firecracker further.

https://sandflare.io"},"title":{"matchLevel":"none","matchedWords":[],"value":"Sandflare \u2013 I built a sandbox that launches AI agent VMs in ~300ms"}},"_tags":["story","author_ajaysheoran2323","story_47583255","ask_hn"],"author":"ajaysheoran2323","children":[47583347,47583442,47584466,47587177,47595552],"created_at":"2026-03-31T05:54:41Z","created_at_i":1774936481,"num_comments":4,"objectID":"47583255","points":5,"story_id":47583255,"story_text":"I've been building Sandflare for the past few months \u2014 it launches Firecracker microVMs for AI agents in ~300ms cold start. The idea came from running LLM-generated code in production. Docker felt too risky (shared kernel), full VMs too slow (5\u201310s). Firecracker hits the middle: real VM isolation, fast boot.

I also added managed Postgres because almost every agent I built needed persistent state. One call wires a database into a sandbox.

There are great tools in this space already (E2B, Modal, Daytona) \u2014 I wanted something with batteries-included Postgres, and simpler pricing

What I'm trying to figure out: how do I get cold start below 100ms? Currently the bottleneck is the Firecracker API + network setup. Would love to hear from anyone who's pushed Firecracker further.

https://sandflare.io","title":"Sandflare \u2013 I built a sandbox that launches AI agent VMs in ~300ms","updated_at":"2026-04-08T01:17:43Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"soham123"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Hey HN,

I built SWE-Kit, LLM toolkit (Function callable tools) which makes building agents specialised in coding like Devin very easy.

I noticed a typical pattern while building local agents: creating & perfecting LLM tools to interact with system or codebase was the repeated and time-consuming. We created a layer that simplifies building agents that can interact with code, file system, git, shell and allows you to quickly solve for a wide variety of coding agent use cases.

Aren\u2019t there open coding agents already? Well, yes, but most folks would want to solve their specific use case like a large refactor and current coding agents aren\u2019t customisable to your specific use case or aren\u2019t meant to be molded to different workflows.

The idea is to provide a library of tools so you can build software engineering agents with a few lines of code in agentic framework of your choice.

We have solved following hard parts for everyone - \n- Optimized Coding Tools: Includes Code Analysis, File Operations, and Shell tools for seamless interaction with codebases and operating systems.\n- Browser Interaction Tool: Enables navigation and interaction with UI-based applications and codebases.\n- Framework Agnostic: Compatible with frameworks like LangChain, LlamaIndex, CrewAI, and Autogen, this allows you to work with your preferred setup.\n- Third-Party Integrations: Connects with applications like GitHub, Slack, Jira, and Gmail to build fully autonomous, end-to-end AI coding agents.\n- Flexible Deployment: Run on Local, Docker, Fly.io, E2b, AWS Lambda (soon!)

Is this the 10x Coding Agent I was looking for?

No this is not a coding agent but allows you to build your custom coding agent in framework of your choice.

We have created some templates to get started quickly though:\n- GitHub PR Agent: Autonomously reviews GitHub pull requests with full codebase context.\n- SWE Agent: Writes new features, debugs code, refactors, and creates tests.\n- Codebase Q&A Agent: Enables natural language interactions with the codebase.

To better showcase the SWE kit's capability, we tested it on [swebench.com](https://www.swebench.com/), the benchmark for testing coding agents. It scored 48.60%, whereas Devin scored only 13.86%.

If you end up using this, please do provide feedback and if you need help building coding agent feel free to reach out to us

I (Soham) & my cofounder Karan are both active on this thread to answer any questions!"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: LLM Function Calling Library to Interact with File, Shell, Git and Code"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://swekit.dev/"}},"_tags":["story","author_soham123","story_42063098","show_hn"],"author":"soham123","created_at":"2024-11-06T14:47:44Z","created_at_i":1730904464,"num_comments":0,"objectID":"42063098","points":5,"story_id":42063098,"story_text":"Hey HN,

I built SWE-Kit, LLM toolkit (Function callable tools) which makes building agents specialised in coding like Devin very easy.

I noticed a typical pattern while building local agents: creating & perfecting LLM tools to interact with system or codebase was the repeated and time-consuming. We created a layer that simplifies building agents that can interact with code, file system, git, shell and allows you to quickly solve for a wide variety of coding agent use cases.

Aren\u2019t there open coding agents already? Well, yes, but most folks would want to solve their specific use case like a large refactor and current coding agents aren\u2019t customisable to your specific use case or aren\u2019t meant to be molded to different workflows.

The idea is to provide a library of tools so you can build software engineering agents with a few lines of code in agentic framework of your choice.

We have solved following hard parts for everyone - \n- Optimized Coding Tools: Includes Code Analysis, File Operations, and Shell tools for seamless interaction with codebases and operating systems.\n- Browser Interaction Tool: Enables navigation and interaction with UI-based applications and codebases.\n- Framework Agnostic: Compatible with frameworks like LangChain, LlamaIndex, CrewAI, and Autogen, this allows you to work with your preferred setup.\n- Third-Party Integrations: Connects with applications like GitHub, Slack, Jira, and Gmail to build fully autonomous, end-to-end AI coding agents.\n- Flexible Deployment: Run on Local, Docker, Fly.io, E2b, AWS Lambda (soon!)

Is this the 10x Coding Agent I was looking for?

No this is not a coding agent but allows you to build your custom coding agent in framework of your choice.

We have created some templates to get started quickly though:\n- GitHub PR Agent: Autonomously reviews GitHub pull requests with full codebase context.\n- SWE Agent: Writes new features, debugs code, refactors, and creates tests.\n- Codebase Q&A Agent: Enables natural language interactions with the codebase.

To better showcase the SWE kit's capability, we tested it on [swebench.com](https://www.swebench.com/), the benchmark for testing coding agents. It scored 48.60%, whereas Devin scored only 13.86%.

If you end up using this, please do provide feedback and if you need help building coding agent feel free to reach out to us

I (Soham) & my cofounder Karan are both active on this thread to answer any questions!","title":"Show HN: LLM Function Calling Library to Interact with File, Shell, Git and Code","updated_at":"2024-11-18T04:36:30Z","url":"https://swekit.dev/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"steadyelk"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"I built an MCP to handle all my wearables data, and it was super helpful, but the types of questions that agent could answer without access to write its own programs was limited. A single wearables stream can have 20k data points for 1 hr of activity (1Hz across GPS, barometer, temperature, and HR), and the typical MCP design either has the MCP author manually defining the aggregation methods (e.g., get_average_heartrate) or the LLM has to hold in it's context the large data representation.

I recently gave the MCP tools to create and access an iPython kernel, where it has the ability to specify what package to download before the session is created, and it can manipulate a copy of the data by writing its own code.

Making sure data was kept private and tools / code were secure was what took most of the time, and I'm wondering if there are any tools folks are using to make this easier. I know there are tools like e2b.dev which provide code sandboxes, but I feel like other data providing MCPs will run into this issue, so there must be some solution / architecture design I'm missing."},"title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: Is anyone here giving their MCP server a code execution environment?"}},"_tags":["story","author_steadyelk","story_47513339","ask_hn"],"author":"steadyelk","children":[47513583],"created_at":"2026-03-25T04:42:58Z","created_at_i":1774413778,"num_comments":1,"objectID":"47513339","points":4,"story_id":47513339,"story_text":"I built an MCP to handle all my wearables data, and it was super helpful, but the types of questions that agent could answer without access to write its own programs was limited. A single wearables stream can have 20k data points for 1 hr of activity (1Hz across GPS, barometer, temperature, and HR), and the typical MCP design either has the MCP author manually defining the aggregation methods (e.g., get_average_heartrate) or the LLM has to hold in it's context the large data representation.

I recently gave the MCP tools to create and access an iPython kernel, where it has the ability to specify what package to download before the session is created, and it can manipulate a copy of the data by writing its own code.

Making sure data was kept private and tools / code were secure was what took most of the time, and I'm wondering if there are any tools folks are using to make this easier. I know there are tools like e2b.dev which provide code sandboxes, but I feel like other data providing MCPs will run into this issue, so there must be some solution / architecture design I'm missing.","title":"Ask HN: Is anyone here giving their MCP server a code execution environment?","updated_at":"2026-03-25T07:02:43Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"iacguy"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Hi HN, today we're launching Open Sandbox, an open-source, self-hostable Linux sandbox written in Rust. It runs commands in isolated environments using process level sandboxing rather than micro VMs.

AI agents and LLMs generate code, but you can't just exec() untrusted code on your machine. You need a sandbox, an isolated environment where that code runs without access to your hostsystem/data.

The idea came from a conversation between my co-founders and me about slow startup times in Firecracker/micro-VM-based sandboxes. He mentioned that during his PhD in the UK, he'd used process-level sandboxes in competitive programming, and they were fast. That sent us down a rabbit hole.

We looked at existing implementations of sandboxes with process level isolation like Isolate, Minijail, nsjail and found that process-level sandboxes have very low resource overhead and surprisingly fast startup times. So we built our own in Rust.

How this compares to E2B, Modal etc? Those are great products, but they're hosted services built on micro-VMs or containers. You send your workloads to their infrastructure and pay per usage.

Open Sandbox is different in three ways:

1. Self-hosted and open source. Your code never leaves your machines. (although, yes, e2b is open source but it is far from easy to self-host)

2. Process-level isolation instead of VMs. This means ~100ms startup and very low resource overhead per sandbox, vs the micro-VM approaches.

3. The trade-off is weaker isolation. A kernel exploit could escape the\nsandbox.

We also ran some benchmarks vs E2B: Open Sandbox was 2x faster at sandbox creation, faster at Git Clone, and also had 6x concurrency. E2B was faster on command execution. More details in the README of the repo with exact numbers.

This is incredibly early - curious on feedback!"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Open Sandbox \u2013 an open-source self-hostable Linux sandbox for AI agents"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/diggerhq/opensandbox"}},"_tags":["story","author_iacguy","story_46832620","show_hn"],"author":"iacguy","created_at":"2026-01-31T02:07:57Z","created_at_i":1769825277,"num_comments":0,"objectID":"46832620","points":4,"story_id":46832620,"story_text":"Hi HN, today we're launching Open Sandbox, an open-source, self-hostable Linux sandbox written in Rust. It runs commands in isolated environments using process level sandboxing rather than micro VMs.

AI agents and LLMs generate code, but you can't just exec() untrusted code on your machine. You need a sandbox, an isolated environment where that code runs without access to your hostsystem/data.

The idea came from a conversation between my co-founders and me about slow startup times in Firecracker/micro-VM-based sandboxes. He mentioned that during his PhD in the UK, he'd used process-level sandboxes in competitive programming, and they were fast. That sent us down a rabbit hole.

We looked at existing implementations of sandboxes with process level isolation like Isolate, Minijail, nsjail and found that process-level sandboxes have very low resource overhead and surprisingly fast startup times. So we built our own in Rust.

How this compares to E2B, Modal etc? Those are great products, but they're hosted services built on micro-VMs or containers. You send your workloads to their infrastructure and pay per usage.

Open Sandbox is different in three ways:

1. Self-hosted and open source. Your code never leaves your machines. (although, yes, e2b is open source but it is far from easy to self-host)

2. Process-level isolation instead of VMs. This means ~100ms startup and very low resource overhead per sandbox, vs the micro-VM approaches.

3. The trade-off is weaker isolation. A kernel exploit could escape the\nsandbox.

We also ran some benchmarks vs E2B: Open Sandbox was 2x faster at sandbox creation, faster at Git Clone, and also had 6x concurrency. E2B was faster on command execution. More details in the README of the repo with exact numbers.

This is incredibly early - curious on feedback!","title":"Show HN: Open Sandbox \u2013 an open-source self-hostable Linux sandbox for AI agents","updated_at":"2026-03-05T23:27:38Z","url":"https://github.com/diggerhq/opensandbox"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"avyvar"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"We're Dari (YC F25). We're building browser automation agents with a deterministic caching layer; for that to work, we needed a way to save and rerun code across sessions. There's also no audit trail with traditional sandboxes; when something goes wrong, you have no idea what the agent actually did.

So we built SkillFS. The core idea is simple: every agent sandbox is a git repo. The agent does its work, commits its progress, and the session ends. When you start a new session, SkillFS restores the repo from a git bundle and the agent continues where it left off. And because it's git, you have a complete history of every action the agent took.

The workflow looks like this:

  1. Agent works in a sandboxed environment, making changes to files\n  2. Agent commits progress at meaningful checkpoints\n  3. Session ends, git bundle gets saved to storage (local or GCS)\n  4. Next session starts, bundle is restored, agent resumes\n  5. Need to debug? git log shows you exactly what happened\n
\nKey features:

  - Persistent state via git bundles, with pluggable storage backends (local filesystem or GCS), so skills and scripts from previous sessions are reusable.\n\n  - MCP integration that lets you plug in any server. SkillFS also generates Python wrappers from MCP tool definitions and uploads them to the sandbox, so your agent can call any MCP tool as regular code to cache sequences of MCP interactions into deterministic scripts.\n\n  - Built-in LLM runner with standard tools (glob, grep, read/write/edit files, run commands), or you can bring your own agent loop.\n\n  - Runs on E2B sandboxes so agents execute code in isolated environments. Agent skills can also be easily imported from local files or from GitHub repos.\n
\nWe're open sourcing this because agent persistence is a problem every team building with agents ends up solving differently, and we think bash+git is a good answer that more people should be using.

  > pip install skillfs
"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: SkillFS \u2013 Git-backed persistent sandboxes for AI agents"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/mupt-ai/skillfs"}},"_tags":["story","author_avyvar","story_46543093","show_hn"],"author":"avyvar","created_at":"2026-01-08T16:38:43Z","created_at_i":1767890323,"num_comments":0,"objectID":"46543093","points":4,"story_id":46543093,"story_text":"We're Dari (YC F25). We're building browser automation agents with a deterministic caching layer; for that to work, we needed a way to save and rerun code across sessions. There's also no audit trail with traditional sandboxes; when something goes wrong, you have no idea what the agent actually did.

So we built SkillFS. The core idea is simple: every agent sandbox is a git repo. The agent does its work, commits its progress, and the session ends. When you start a new session, SkillFS restores the repo from a git bundle and the agent continues where it left off. And because it's git, you have a complete history of every action the agent took.

The workflow looks like this:

  1. Agent works in a sandboxed environment, making changes to files\n  2. Agent commits progress at meaningful checkpoints\n  3. Session ends, git bundle gets saved to storage (local or GCS)\n  4. Next session starts, bundle is restored, agent resumes\n  5. Need to debug? git log shows you exactly what happened\n
\nKey features:

  - Persistent state via git bundles, with pluggable storage backends (local filesystem or GCS), so skills and scripts from previous sessions are reusable.\n\n  - MCP integration that lets you plug in any server. SkillFS also generates Python wrappers from MCP tool definitions and uploads them to the sandbox, so your agent can call any MCP tool as regular code to cache sequences of MCP interactions into deterministic scripts.\n\n  - Built-in LLM runner with standard tools (glob, grep, read/write/edit files, run commands), or you can bring your own agent loop.\n\n  - Runs on E2B sandboxes so agents execute code in isolated environments. Agent skills can also be easily imported from local files or from GitHub repos.\n
\nWe're open sourcing this because agent persistence is a problem every team building with agents ends up solving differently, and we think bash+git is a good answer that more people should be using.

  > pip install skillfs
","title":"Show HN: SkillFS \u2013 Git-backed persistent sandboxes for AI agents","updated_at":"2026-03-05T23:19:21Z","url":"https://github.com/mupt-ai/skillfs"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"markoh49"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Hello Hacker News! I'm Mark. I'm building Moru, an open-source runtime for AI agents that runs each session in an isolated Firecracker microVM. It started as a fork of E2B, and most of the low-level Firecracker runtime is still from upstream.

It lets you run agent harnesses like Claude Code or Codex in the cloud, giving each session its own isolated microVM with filesystem and shell access.

The repo is: https://github.com/moru-ai/moru

Each VM is a snapshot of a Docker build. You define a Dockerfile, CPU, memory limits, and Moru runs the build inside a Firecracker VM, then pauses and saves the exact state: CPU, dirty memory pages, and changed filesystem blocks.

When you spawn a new VM, it resumes from that template snapshot. Memory snapshot is lazy-loaded via userfaultfd, which helps sandboxes start within a second.

Each VM runs on Firecracker with KVM isolation and a dedicated kernel. Network uses namespaces for isolation and iptables for access control.

From outside, you talk to the VM through the Moru CLI or TypeScript/Python SDK. Inside, it's just Linux. Run commands, read/write files, anything you'd do on a normal machine.

I've been building AI apps since the ChatGPT launch. These days, when an agent needs to solve complex problems, I just give it filesystem + shell access. This works well because it (1) handles large data without pushing everything into the model context window, and (2) reuses tools that already work (Python, Bash, etc.). This has become much more practical as frontier models have gotten good at tool use and multi-step workflows.

Now models run for hours on real tasks. As models get smarter, the harness should give models more autonomy, but with safe guardrails. I want Moru to help developers focus on building agents, not the underlying runtime and infra.

You can try the cloud version without setting up your own infra. It's fully self-hostable including the infra and the dashboard. I'm planning to keep this open like the upstream repo (Apache 2.0).

Give it a spin: https://github.com/moru-ai/moru \nLet me know what you think!

Next features I'm working toward:

- Richer streaming: today it's mostly stdin/stdout. That pushes me to overload print/console.log for control-plane communication, which gets messy fast. I want a separate streaming channel for structured events and coordination with the control plane (often an app server), while keeping stdout/stderr for debugging.

- Seamless deployment: a deploy experience closer to Vercel/Fly.io.

- A storage primitive: save and resume sessions without always having to manually sync workspace and session state.

Open to your feature requests or suggestions.

I'm focusing on making it easy to deploy and run local-first agent harnesses (e.g., Claude Agent SDK) inside isolated VMs. If you've built or are building those, I'd appreciate any notes on what's missing, or what you'd prioritize first."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: I'm building an open-source AI agent runtime using Firecracker microVMs"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/moru-ai/moru"}},"_tags":["story","author_markoh49","story_46635859","show_hn"],"author":"markoh49","children":[46643917],"created_at":"2026-01-15T17:18:30Z","created_at_i":1768497510,"num_comments":1,"objectID":"46635859","points":3,"story_id":46635859,"story_text":"Hello Hacker News! I'm Mark. I'm building Moru, an open-source runtime for AI agents that runs each session in an isolated Firecracker microVM. It started as a fork of E2B, and most of the low-level Firecracker runtime is still from upstream.

It lets you run agent harnesses like Claude Code or Codex in the cloud, giving each session its own isolated microVM with filesystem and shell access.

The repo is: https://github.com/moru-ai/moru

Each VM is a snapshot of a Docker build. You define a Dockerfile, CPU, memory limits, and Moru runs the build inside a Firecracker VM, then pauses and saves the exact state: CPU, dirty memory pages, and changed filesystem blocks.

When you spawn a new VM, it resumes from that template snapshot. Memory snapshot is lazy-loaded via userfaultfd, which helps sandboxes start within a second.

Each VM runs on Firecracker with KVM isolation and a dedicated kernel. Network uses namespaces for isolation and iptables for access control.

From outside, you talk to the VM through the Moru CLI or TypeScript/Python SDK. Inside, it's just Linux. Run commands, read/write files, anything you'd do on a normal machine.

I've been building AI apps since the ChatGPT launch. These days, when an agent needs to solve complex problems, I just give it filesystem + shell access. This works well because it (1) handles large data without pushing everything into the model context window, and (2) reuses tools that already work (Python, Bash, etc.). This has become much more practical as frontier models have gotten good at tool use and multi-step workflows.

Now models run for hours on real tasks. As models get smarter, the harness should give models more autonomy, but with safe guardrails. I want Moru to help developers focus on building agents, not the underlying runtime and infra.

You can try the cloud version without setting up your own infra. It's fully self-hostable including the infra and the dashboard. I'm planning to keep this open like the upstream repo (Apache 2.0).

Give it a spin: https://github.com/moru-ai/moru \nLet me know what you think!

Next features I'm working toward:

- Richer streaming: today it's mostly stdin/stdout. That pushes me to overload print/console.log for control-plane communication, which gets messy fast. I want a separate streaming channel for structured events and coordination with the control plane (often an app server), while keeping stdout/stderr for debugging.

- Seamless deployment: a deploy experience closer to Vercel/Fly.io.

- A storage primitive: save and resume sessions without always having to manually sync workspace and session state.

Open to your feature requests or suggestions.

I'm focusing on making it easy to deploy and run local-first agent harnesses (e.g., Claude Agent SDK) inside isolated VMs. If you've built or are building those, I'd appreciate any notes on what's missing, or what you'd prioritize first.","title":"Show HN: I'm building an open-source AI agent runtime using Firecracker microVMs","updated_at":"2026-07-28T15:43:36Z","url":"https://github.com/moru-ai/moru"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"danterolle"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"I got tired of sending every text I translate to Google/DeepL. Even with all the opt-out options and privacy policies, it never felt right especially for some work documents, personal writing, or anything sensitive. So I decided to build this tool, which lets me use LLMs for context translations and also a standard translation engine like Argos. It works with Ollama, llama.cpp or argos-translate, and you can configure the model you want to use.

loqi translate --model phi4-mini --from it --to en "Ciao mondo"

Obviously, the quality of the translation depends entirely on the model used, but I've noticed that you can get good, if not excellent, results even with a small model (such as Gemma 4 E2B or Phi4-mini).

So there you have it: Loqi is open source, cross-platform (MacOS, GNU/Linux, Windows), written in Go with Bubble Tea for the TUI. It allows the model to translate individual sentences or process entire files (whether plain text, Markdown or JSON).

I used LLMs to help write parts of the code, including self-hostable ones (those that run on AMD GPU with 16 GB of VRAM), but I've tried to set up the project as much as possible, and there's probably a lot more work to be done and some ideas to implement.

I'd be more than happy to accept contributions."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Loqi, a \"local-first\" translation tool using Ollama/llama.cpp"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/danterolle/loqi"}},"_tags":["story","author_danterolle","story_48632814","show_hn"],"author":"danterolle","created_at":"2026-06-22T17:01:54Z","created_at_i":1782147714,"num_comments":0,"objectID":"48632814","points":3,"story_id":48632814,"story_text":"I got tired of sending every text I translate to Google/DeepL. Even with all the opt-out options and privacy policies, it never felt right especially for some work documents, personal writing, or anything sensitive. So I decided to build this tool, which lets me use LLMs for context translations and also a standard translation engine like Argos. It works with Ollama, llama.cpp or argos-translate, and you can configure the model you want to use.

loqi translate --model phi4-mini --from it --to en "Ciao mondo"

Obviously, the quality of the translation depends entirely on the model used, but I've noticed that you can get good, if not excellent, results even with a small model (such as Gemma 4 E2B or Phi4-mini).

So there you have it: Loqi is open source, cross-platform (MacOS, GNU/Linux, Windows), written in Go with Bubble Tea for the TUI. It allows the model to translate individual sentences or process entire files (whether plain text, Markdown or JSON).

I used LLMs to help write parts of the code, including self-hostable ones (those that run on AMD GPU with 16 GB of VRAM), but I've tried to set up the project as much as possible, and there's probably a lot more work to be done and some ideas to implement.

I'd be more than happy to accept contributions.","title":"Show HN: Loqi, a \"local-first\" translation tool using Ollama/llama.cpp","updated_at":"2026-06-29T18:45:11Z","url":"https://github.com/danterolle/loqi"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"joegibbs"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Experiment that I've made. The models get access to an E2B sandbox and are instructed to create an ad according to the specifications (they can choose whatever tools they want to use for it, e.g. Pillow, Chromium) as a proxy for their ability to use tools, create other kinds of images, do complex layouts etc. Currently Opus 4.8 is on top (not surprising, but it did take 66 conversation turns to create the image) and GLM-5.2 is on fifth (which I do find surprising because it doesn't have image capabilty)."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: AdvertBench, ranking the ability of LLMs to create image ads"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://advertbench.com"}},"_tags":["story","author_joegibbs","story_48618085","show_hn"],"author":"joegibbs","created_at":"2026-06-21T11:50:31Z","created_at_i":1782042631,"num_comments":0,"objectID":"48618085","points":3,"story_id":48618085,"story_text":"Experiment that I've made. The models get access to an E2B sandbox and are instructed to create an ad according to the specifications (they can choose whatever tools they want to use for it, e.g. Pillow, Chromium) as a proxy for their ability to use tools, create other kinds of images, do complex layouts etc. Currently Opus 4.8 is on top (not surprising, but it did take 66 conversation turns to create the image) and GLM-5.2 is on fifth (which I do find surprising because it doesn't have image capabilty).","title":"Show HN: AdvertBench, ranking the ability of LLMs to create image ads","updated_at":"2026-06-21T20:18:08Z","url":"https://advertbench.com"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"geoctl"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Hello HN, Cordium is a general-purpose sandbox platform built on Kubernetes and Octelium, may main work https://github.com/octelium/octelium, that can be used for various use cases, including coding for developers with VSCode, Zed, etc. (i.e. self-hosted GitHub Codespaces alternative), AI agent tasks (i.e. FOSS alternative to AI sandbox products such as E2B, Daytona, etc.), CI/CD workloads (e.g. building and publishing Docker images etc.), and more importantly for secretless remote access to infrastructure for devs and automated workloads.

The main _differentiator_ here, compared to other dev environments and sandbox platforms, is that Cordium automatically provides identity-based, secretless secure access to resources/infrastructure (e.g. APIs, SSH, databases, k8s, etc.) without having to inject credentials (e.g. API keys, SSH private keys, database passwords, etc.) into the sandbox where the upstream credential is held by the identity-aware proxy of the Octelium-protected resource outside the reach of the sandbox. You can simply think of it as a sandbox + ZTNA/remote-access-VPN baked-in where access to infrastructure is based on identity and policy-as-code rather than credentials.

Cordium is a purely FOSS project under Apache 2.0 that's meant for self-hosting and there are no plans for a pro/SaaS/cloud version. The development of the project started back in 2022 and it is already being used by a few organizations that use Octelium since last year. Happy to answer any questions."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Cordium: FOSS sandbox platform that eliminates credential injection"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/octelium/cordium"}},"_tags":["story","author_geoctl","story_48344623","show_hn"],"author":"geoctl","children":[48365858],"created_at":"2026-05-31T10:41:00Z","created_at_i":1780224060,"num_comments":0,"objectID":"48344623","points":3,"story_id":48344623,"story_text":"Hello HN, Cordium is a general-purpose sandbox platform built on Kubernetes and Octelium, may main work https://github.com/octelium/octelium, that can be used for various use cases, including coding for developers with VSCode, Zed, etc. (i.e. self-hosted GitHub Codespaces alternative), AI agent tasks (i.e. FOSS alternative to AI sandbox products such as E2B, Daytona, etc.), CI/CD workloads (e.g. building and publishing Docker images etc.), and more importantly for secretless remote access to infrastructure for devs and automated workloads.

The main _differentiator_ here, compared to other dev environments and sandbox platforms, is that Cordium automatically provides identity-based, secretless secure access to resources/infrastructure (e.g. APIs, SSH, databases, k8s, etc.) without having to inject credentials (e.g. API keys, SSH private keys, database passwords, etc.) into the sandbox where the upstream credential is held by the identity-aware proxy of the Octelium-protected resource outside the reach of the sandbox. You can simply think of it as a sandbox + ZTNA/remote-access-VPN baked-in where access to infrastructure is based on identity and policy-as-code rather than credentials.

Cordium is a purely FOSS project under Apache 2.0 that's meant for self-hosting and there are no plans for a pro/SaaS/cloud version. The development of the project started back in 2022 and it is already being used by a few organizations that use Octelium since last year. Happy to answer any questions.","title":"Show HN: Cordium: FOSS sandbox platform that eliminates credential injection","updated_at":"2026-06-02T04:00:35Z","url":"https://github.com/octelium/cordium"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"krunkworx"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"I'm genuinely surprised at what people are willing to share with AI companions. Read Replika's privacy policy. Then Character.AI's. These apps store your most personal conversations on their servers, linked to your email address. A breach or subpoena and your identity is attached to everything you ever told your "AI friend." Eek.

The only thing I think actually solves this is local inference. I remember browing r/LocalLLaMA and years ago and thinking this is the future. Local models are finally good enough. I was playing with the bonsai 8B 1-bit quant model a few weeks back and I think we're almost there. I built friendAI to see if there's market demand for local inference. Everything runs on your phone.

What's actually on-device:

- Bonsai-8B (1-bit quantized Qwen3-8B, ~1.3GB) via MLX for speed\n- Gemma 4 E2B (~4.5GB, GGUF) via llama.cpp for vision\n- A unified client that routes between them

A few things I'm reasonably proud of solving in about a week:

- Turns out the hardest part was actually managing the background model downloads that survive crashes, network drops and reboots. You can start chatting before the download finishes.\n- Runtime thread auto-tuning that benchmarks your actual device at startup rather than guessing with a static heuristic\n- Local memory without a vector DB. TF-IDF style ranking with recency decay. No embedding model needed.

Happy to go deep on any of it.\nwww.friendai.pro"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: I read Replika's privacy policy and then built a competitor"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://apps.apple.com/us/app/friend-ai-private-chat/id6761649790"}},"_tags":["story","author_krunkworx","story_47915214","show_hn"],"author":"krunkworx","created_at":"2026-04-26T22:09:10Z","created_at_i":1777241350,"num_comments":0,"objectID":"47915214","points":3,"story_id":47915214,"story_text":"I'm genuinely surprised at what people are willing to share with AI companions. Read Replika's privacy policy. Then Character.AI's. These apps store your most personal conversations on their servers, linked to your email address. A breach or subpoena and your identity is attached to everything you ever told your "AI friend." Eek.

The only thing I think actually solves this is local inference. I remember browing r/LocalLLaMA and years ago and thinking this is the future. Local models are finally good enough. I was playing with the bonsai 8B 1-bit quant model a few weeks back and I think we're almost there. I built friendAI to see if there's market demand for local inference. Everything runs on your phone.

What's actually on-device:

- Bonsai-8B (1-bit quantized Qwen3-8B, ~1.3GB) via MLX for speed\n- Gemma 4 E2B (~4.5GB, GGUF) via llama.cpp for vision\n- A unified client that routes between them

A few things I'm reasonably proud of solving in about a week:

- Turns out the hardest part was actually managing the background model downloads that survive crashes, network drops and reboots. You can start chatting before the download finishes.\n- Runtime thread auto-tuning that benchmarks your actual device at startup rather than guessing with a static heuristic\n- Local memory without a vector DB. TF-IDF style ranking with recency decay. No embedding model needed.

Happy to go deep on any of it.\nwww.friendai.pro","title":"Show HN: I read Replika's privacy policy and then built a competitor","updated_at":"2026-05-07T17:41:00Z","url":"https://apps.apple.com/us/app/friend-ai-private-chat/id6761649790"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"probiruk"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"I built this TypeScript agent framework to make it easier to build AI agents that are type-safe end-to-end (client to server). It also lets you package common runtime behavior as plugins and reuse them across agents, like auth, rate limits, and sandboxing (e.g. Daytona, E2B, or other sandbox providers). It\u2019s fully event-driven.

Most setups I tried didn\u2019t handle end-to-end typing well, didn\u2019t fit cleanly into an existing app, or made it hard to reuse behavior across agents. This tries to keep things consistent across client and server, make plugins easy to share, and support stateful, resumable runs without a lot of glue code.

Website: https://better-agent.com\nGitHub: https://github.com/better-agent/better-agent"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Better Agent \u2013 A composable AI agent framework in TypeScript"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www.better-agent.com/"}},"_tags":["story","author_probiruk","story_47694452","show_hn"],"author":"probiruk","created_at":"2026-04-08T18:42:36Z","created_at_i":1775673756,"num_comments":0,"objectID":"47694452","points":3,"story_id":47694452,"story_text":"I built this TypeScript agent framework to make it easier to build AI agents that are type-safe end-to-end (client to server). It also lets you package common runtime behavior as plugins and reuse them across agents, like auth, rate limits, and sandboxing (e.g. Daytona, E2B, or other sandbox providers). It\u2019s fully event-driven.

Most setups I tried didn\u2019t handle end-to-end typing well, didn\u2019t fit cleanly into an existing app, or made it hard to reuse behavior across agents. This tries to keep things consistent across client and server, make plugins easy to share, and support stateful, resumable runs without a lot of glue code.

Website: https://better-agent.com\nGitHub: https://github.com/better-agent/better-agent","title":"Show HN: Better Agent \u2013 A composable AI agent framework in TypeScript","updated_at":"2026-04-08T18:52:45Z","url":"https://www.better-agent.com/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ATechGuy"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"In the last couple of months, several new solutions for sandboxing AI agents have launched (microVMs, WASM runtimes, browser isolation, hardened tool containers, etc.). Curious to hear from people using them in production. Are they working as advertised, or are there still major tradeoffs around security, cost, and performance?

Here's my list of sandboxing solutions launched in the last year alone: E2B, AIO Sandbox, Sandboxer, AgentSphere, Yolobox, Exe.dev, yolo-cage, SkillFS, ERA Jazzberry Computer, Vibekit, Daytona, Modal, Cognitora, YepCode, Run Compute, CLI Fence, Landrun, Sprites, pctx-sandbox, pctx Sandbox, Agent SDK, Lima-devbox, OpenServ, Browser Agent Playground, Flintlock Agent, Quickstart, Bouvet Sandbox, Arrakis, Cellmate (ceLLMate), AgentFence, Tasker, DenoSandbox, Capsule (WASM-based), Volant, Nono, NetFence"},"title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: The new wave of AI agent sandboxes?"}},"_tags":["story","author_ATechGuy","story_47254841","ask_hn"],"author":"ATechGuy","created_at":"2026-03-04T22:26:27Z","created_at_i":1772663187,"num_comments":0,"objectID":"47254841","points":3,"story_id":47254841,"story_text":"In the last couple of months, several new solutions for sandboxing AI agents have launched (microVMs, WASM runtimes, browser isolation, hardened tool containers, etc.). Curious to hear from people using them in production. Are they working as advertised, or are there still major tradeoffs around security, cost, and performance?

Here's my list of sandboxing solutions launched in the last year alone: E2B, AIO Sandbox, Sandboxer, AgentSphere, Yolobox, Exe.dev, yolo-cage, SkillFS, ERA Jazzberry Computer, Vibekit, Daytona, Modal, Cognitora, YepCode, Run Compute, CLI Fence, Landrun, Sprites, pctx-sandbox, pctx Sandbox, Agent SDK, Lima-devbox, OpenServ, Browser Agent Playground, Flintlock Agent, Quickstart, Bouvet Sandbox, Arrakis, Cellmate (ceLLMate), AgentFence, Tasker, DenoSandbox, Capsule (WASM-based), Volant, Nono, NetFence","title":"Ask HN: The new wave of AI agent sandboxes?","updated_at":"2026-03-05T23:41:47Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"adlkiarash"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"You can watch an agent destroy a file and instantly revert it in an interactive browser terminal here (no signup): https://mcp.undisk.app/e2b"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: An edge MCP file system with a 50ms undo button for AI agents"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://mcp.undisk.app"}},"_tags":["story","author_adlkiarash","story_47773437","show_hn"],"author":"adlkiarash","children":[47773452,47773461],"created_at":"2026-04-15T01:05:30Z","created_at_i":1776215130,"num_comments":2,"objectID":"47773437","points":2,"story_id":47773437,"story_text":"You can watch an agent destroy a file and instantly revert it in an interactive browser terminal here (no signup): https://mcp.undisk.app/e2b","title":"Show HN: An edge MCP file system with a 50ms undo button for AI agents","updated_at":"2026-04-15T01:39:38Z","url":"https://mcp.undisk.app"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"notanaiagent"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Hey HN \u2014 I'm Julius, cofounder of Righthand. We've been building AI employees (we call them "Righthands") that actually live inside your workflow instead of waiting for prompts. Each one has a name, email address and phone number + unified memory across all communication channels. They run in their own E2B sandboxes, with a filesystem and custom CLI.

We built this because of anxiety. Everyone is anxious about something, especially at work, and we realized that a lot of the tasks people are anxious about are trivial things that we just put off. Not for any good reason, but because we just don't do them. The idea behind Righthand is that you should just be able to forward away the anxiety-driving email or call in for a summary of your to dos on your way to work.

Real customer uses from the last 3 months: \n - automated inbox triage\n - managing maintenance staff for a property in Malta\n - booking a meeting (hundreds, actually)\n - booking a doctor's appointment\n - customized daily briefs

Each one of these things was anxiety driving but it wasn't getting done because it was always lower priority than everything else.

The last two months we rebuilt the skill system from scratch (Skills V4), updated our pricing model, onboarded more customers, improved visibility into the Righthand's "brain" and onboarded a lot of customers by hand. Other things that shipped: nightly self-review (the Righthand reviews its own day and writes notes), goal requests from inside the sandbox, communication-style presets, Bedrock + Codex fallback routing for over-time personas, Parallel web search, Slack app auto-provisioning via Browser Use (yes, we literally drive api.slack.com headlessly to provision each Righthand its own app \u2014 happy to go into why).

Trial is card-free now: https://www.righthand.ai. Pricing is $99 Starter / $199 Pro - all with 1 wk free trial. Ask: would love feedback on UI / interaction paradigm. Go through the onboarding and tell us what you liked and what you hated. You won't be charged for 7 days and can easily cancel."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Righthand \u2013 Autonomous AI assistants with skills, goals, and a CLI"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www.rigthhand.ai"}},"_tags":["story","author_notanaiagent","story_48578130","show_hn"],"author":"notanaiagent","children":[48578853],"created_at":"2026-06-17T22:52:52Z","created_at_i":1781736772,"num_comments":1,"objectID":"48578130","points":2,"story_id":48578130,"story_text":"Hey HN \u2014 I'm Julius, cofounder of Righthand. We've been building AI employees (we call them "Righthands") that actually live inside your workflow instead of waiting for prompts. Each one has a name, email address and phone number + unified memory across all communication channels. They run in their own E2B sandboxes, with a filesystem and custom CLI.

We built this because of anxiety. Everyone is anxious about something, especially at work, and we realized that a lot of the tasks people are anxious about are trivial things that we just put off. Not for any good reason, but because we just don't do them. The idea behind Righthand is that you should just be able to forward away the anxiety-driving email or call in for a summary of your to dos on your way to work.

Real customer uses from the last 3 months: \n - automated inbox triage\n - managing maintenance staff for a property in Malta\n - booking a meeting (hundreds, actually)\n - booking a doctor's appointment\n - customized daily briefs

Each one of these things was anxiety driving but it wasn't getting done because it was always lower priority than everything else.

The last two months we rebuilt the skill system from scratch (Skills V4), updated our pricing model, onboarded more customers, improved visibility into the Righthand's "brain" and onboarded a lot of customers by hand. Other things that shipped: nightly self-review (the Righthand reviews its own day and writes notes), goal requests from inside the sandbox, communication-style presets, Bedrock + Codex fallback routing for over-time personas, Parallel web search, Slack app auto-provisioning via Browser Use (yes, we literally drive api.slack.com headlessly to provision each Righthand its own app \u2014 happy to go into why).

Trial is card-free now: https://www.righthand.ai. Pricing is $99 Starter / $199 Pro - all with 1 wk free trial. Ask: would love feedback on UI / interaction paradigm. Go through the onboarding and tell us what you liked and what you hated. You won't be charged for 7 days and can easily cancel.","title":"Show HN: Righthand \u2013 Autonomous AI assistants with skills, goals, and a CLI","updated_at":"2026-06-18T00:18:08Z","url":"https://www.rigthhand.ai"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"geoctl"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Hello HN. Cordium is a FOSS, self-hosted, identity-based, general-purpose sandbox platform that I've been working on for a long time now that is built on Kubernetes and Octelium, my main work. The key difference here for Cordium, when compared to other dev environments (e.g. GitHub Codespaces) and sandbox platforms (e.g. E2B, Daytona, etc.), is that Cordium automatically provides identity-based, secretless secure access to resources/infrastructure (e.g. APIs, SSH, databases, k8s, etc.) without having to inject credentials (e.g. API keys and access tokens, SSH private keys/passwords, database passwords, mTLS private keys, etc.) into the sandbox where the upstream credential is held by the identity-aware proxy of the Octelium-protected resource outside the reach of the sandbox and inject on-the-fly by the identity-aware proxy if the user is authorized via the Octelium policies.

In short, Cordium is not just an isolated execution environment that can replace remote development environments and sandbox platforms, but also equally a secure access platform to infrastructure/resources. It's basically a sandbox platform + a ZTNA/remote-access-VPN baked-in with unified identity management, L7-aware access control and visibility."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: FOSS sandbox platform that hides infra secrets from devs and AI agents"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/octelium/cordium"}},"_tags":["story","author_geoctl","story_48609062","show_hn"],"author":"geoctl","created_at":"2026-06-20T13:19:20Z","created_at_i":1781961560,"num_comments":0,"objectID":"48609062","points":2,"story_id":48609062,"story_text":"Hello HN. Cordium is a FOSS, self-hosted, identity-based, general-purpose sandbox platform that I've been working on for a long time now that is built on Kubernetes and Octelium, my main work. The key difference here for Cordium, when compared to other dev environments (e.g. GitHub Codespaces) and sandbox platforms (e.g. E2B, Daytona, etc.), is that Cordium automatically provides identity-based, secretless secure access to resources/infrastructure (e.g. APIs, SSH, databases, k8s, etc.) without having to inject credentials (e.g. API keys and access tokens, SSH private keys/passwords, database passwords, mTLS private keys, etc.) into the sandbox where the upstream credential is held by the identity-aware proxy of the Octelium-protected resource outside the reach of the sandbox and inject on-the-fly by the identity-aware proxy if the user is authorized via the Octelium policies.

In short, Cordium is not just an isolated execution environment that can replace remote development environments and sandbox platforms, but also equally a secure access platform to infrastructure/resources. It's basically a sandbox platform + a ZTNA/remote-access-VPN baked-in with unified identity management, L7-aware access control and visibility.","title":"Show HN: FOSS sandbox platform that hides infra secrets from devs and AI agents","updated_at":"2026-06-20T14:28:34Z","url":"https://github.com/octelium/cordium"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"geoctl"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Cordium is a FOSS, self-hosted, identity-based, general-purpose sandbox platform that I've been working on for a long time now that is built on Kubernetes and Octelium, my main project.\nThe key differentiator here for Cordium, when compared to other dev environments (e.g. GitHub Codespaces) and sandbox platforms (e.g. E2B, Daytona, etc.), is that Cordium automatically provides identity-based, secretless secure access to resources/infrastructure (e.g. APIs, SSH, databases, k8s, etc.) without having to inject credentials (e.g. API keys, SSH private keys, database passwords, etc.) into the sandbox where the upstream credential is held by the identity-aware proxy of the Octelium-protected resource outside the reach of the sandbox.

In short, Cordium is not just an isolated execution environment that can replace remote development environments and sandbox platforms, but also equally a secure access platform to infrastructure/resources. It's basically a sandbox platform + a ZTNA/remote-access-VPN baked-in with unified identity management, L7-aware access control and visibility."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Cordium \u2013 FOSS identity-based sandbox platform with zero-trust access"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/octelium/cordium"}},"_tags":["story","author_geoctl","story_48533784","show_hn"],"author":"geoctl","children":[48541916],"created_at":"2026-06-14T22:47:53Z","created_at_i":1781477273,"num_comments":0,"objectID":"48533784","points":2,"story_id":48533784,"story_text":"Cordium is a FOSS, self-hosted, identity-based, general-purpose sandbox platform that I've been working on for a long time now that is built on Kubernetes and Octelium, my main project.\nThe key differentiator here for Cordium, when compared to other dev environments (e.g. GitHub Codespaces) and sandbox platforms (e.g. E2B, Daytona, etc.), is that Cordium automatically provides identity-based, secretless secure access to resources/infrastructure (e.g. APIs, SSH, databases, k8s, etc.) without having to inject credentials (e.g. API keys, SSH private keys, database passwords, etc.) into the sandbox where the upstream credential is held by the identity-aware proxy of the Octelium-protected resource outside the reach of the sandbox.

In short, Cordium is not just an isolated execution environment that can replace remote development environments and sandbox platforms, but also equally a secure access platform to infrastructure/resources. It's basically a sandbox platform + a ZTNA/remote-access-VPN baked-in with unified identity management, L7-aware access control and visibility.","title":"Show HN: Cordium \u2013 FOSS identity-based sandbox platform with zero-trust access","updated_at":"2026-06-16T06:48:53Z","url":"https://github.com/octelium/cordium"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ndeodhar"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Hi HN, I'm Neha. I spent years at Google building infrastructure that handled billions of events at 99.999% reliability. When I started building AI agents, I was surprised at how much production plumbing you're expected to own yourself.

The agent itself is the easy part. The hard part is everything around it: where does it execute safely? What happens when it fails midway through a workflow? How do you trigger it from your existing tools? How do you even know what it did?

I kept stitching together Docker, a workflow engine, a notification layer, and custom retry logic. Every team I talked to was doing the same thing. So I built Polos - an open-source runtime that handles the production layer so you just write the agent.

What it does:

- Sandboxed execution: agents run sensitive operations inside managed Docker containers with built-in tools for file I/O, bash, and web search. You don't manage the sandbox or its lifecycle, Polos does. Will support more sandboxes like E2B in the future.

- Slack integration: @mention an agent in Slack, get responses in thread. Trigger workflows from Slack, receive notifications, collect input. Agents become part of your team's existing workflow.

- Durable workflows: if an agent fails mid-run, it resumes from the exact step that failed. Built-in prompt caching with 60-80% cost savings on retries.

- Observability: OpenTelemetry tracing for every step, tool call, and decision.

- LLM agnostic: works with OpenAI, Anthropic, Google, or any provider via Vercel AI SDK and LiteLLM.

The stack is Rust orchestrator (Axum + Tokio + PostgreSQL), Python and TypeScript SDKs, and Vite UI. You can install and run a durable, sandboxed agent in under 5 minutes:

```

curl -fsSL https://install.polos.dev/install.sh | bash

npx create-polos

cd my-project && polos dev

```

Here's a 3-min demo of a coding agent that picks up a GitHub issue, fixes the code in a sandbox, and submits a PR: https://www.youtube.com/watch?v=KYVBpdZ_5eM

Happy to discuss technical decisions and more: why Rust for the orchestrator, how durable execution works without a DAG, and the sandbox lifecycle model.

GitHub: https://github.com/polos-dev/polos"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Polos: Open-source runtime for AI agents with sandbox and durable exec"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/polos-dev/polos"}},"_tags":["story","author_ndeodhar","story_47153680","show_hn"],"author":"ndeodhar","created_at":"2026-02-25T16:24:03Z","created_at_i":1772036643,"num_comments":0,"objectID":"47153680","points":2,"story_id":47153680,"story_text":"Hi HN, I'm Neha. I spent years at Google building infrastructure that handled billions of events at 99.999% reliability. When I started building AI agents, I was surprised at how much production plumbing you're expected to own yourself.

The agent itself is the easy part. The hard part is everything around it: where does it execute safely? What happens when it fails midway through a workflow? How do you trigger it from your existing tools? How do you even know what it did?

I kept stitching together Docker, a workflow engine, a notification layer, and custom retry logic. Every team I talked to was doing the same thing. So I built Polos - an open-source runtime that handles the production layer so you just write the agent.

What it does:

- Sandboxed execution: agents run sensitive operations inside managed Docker containers with built-in tools for file I/O, bash, and web search. You don't manage the sandbox or its lifecycle, Polos does. Will support more sandboxes like E2B in the future.

- Slack integration: @mention an agent in Slack, get responses in thread. Trigger workflows from Slack, receive notifications, collect input. Agents become part of your team's existing workflow.

- Durable workflows: if an agent fails mid-run, it resumes from the exact step that failed. Built-in prompt caching with 60-80% cost savings on retries.

- Observability: OpenTelemetry tracing for every step, tool call, and decision.

- LLM agnostic: works with OpenAI, Anthropic, Google, or any provider via Vercel AI SDK and LiteLLM.

The stack is Rust orchestrator (Axum + Tokio + PostgreSQL), Python and TypeScript SDKs, and Vite UI. You can install and run a durable, sandboxed agent in under 5 minutes:

```

curl -fsSL https://install.polos.dev/install.sh | bash

npx create-polos

cd my-project && polos dev

```

Here's a 3-min demo of a coding agent that picks up a GitHub issue, fixes the code in a sandbox, and submits a PR: https://www.youtube.com/watch?v=KYVBpdZ_5eM

Happy to discuss technical decisions and more: why Rust for the orchestrator, how durable execution works without a DAG, and the sandbox lifecycle model.

GitHub: https://github.com/polos-dev/polos","title":"Show HN: Polos: Open-source runtime for AI agents with sandbox and durable exec","updated_at":"2026-03-05T23:37:11Z","url":"https://github.com/polos-dev/polos"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"heygarrison"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["e2b"],"value":"Hey HN, I'm Garrison, founder of ComputeSDK (https://www.computesdk.com). ComputeSDK is a universal sandbox interface that lets you build, deploy, and interact with sandboxes across any provider (E2B, Daytona, Modal; AWS, Fly.io coming soon) using a single, consistent API with built-in secure tunneling.

Try it out for free here: https://www.computesdk.com/docs/getting-started/quick-start/

Today we are proud to announce the Compute CLI.

I built ComputeSDK after shipping a K8s-based sandbox product for a large enterprise client. Shortly after, when I started talking to AI app builders, I noticed a common frustration: they all ended up building integrations with multiple sandbox providers. Each integration meant rewriting their application, custom proxies, and routing.

As I talked with more customers, I realized sandbox providers and app builders had massive architectural overlap. Both were having to write their own custom terminals, file-watchers, proxies, and more to interact with their sandboxes.

At this point, I realized there was a missing piece in the development infrastructure stack: a way to decouple your application from your sandbox provider while still getting full, live interactivity.

Our original K8s architecture had three services: an API, a gateway, and a sidecar.\nWe extracted our sidecar component and rebuilt it as the "Compute" cli, a lightweight service that installs in any sandbox and automatically creates a secure tunnel to your app. Then we wrapped it with a universal SDK that works across providers. Now you write your sandbox implementation once, and it works everywhere: from E2B, Daytona, & Modal to general cloud platforms like Railway, Fly.io, & AWS.

ComputeSDK provides a consistent interface for sandbox creation and management. It gives you instant access to terminals, file systems, real-time file watching, and a WebSocket event system from any client (server, browser, application). You can switch or add providers easily by importing the provider module and updating environment variables.

We'd love to get any feedback from the community, ideas, or experiences building similar tools.

Thanks!"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Compute CLI \u2013 A universal sandbox SDK with direct browser access"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www.computesdk.com/blog/november-2025-update/"}},"_tags":["story","author_heygarrison","story_45916895","show_hn"],"author":"heygarrison","created_at":"2025-11-13T16:31:31Z","created_at_i":1763051491,"num_comments":0,"objectID":"45916895","points":2,"story_id":45916895,"story_text":"Hey HN, I'm Garrison, founder of ComputeSDK (https://www.computesdk.com). ComputeSDK is a universal sandbox interface that lets you build, deploy, and interact with sandboxes across any provider (E2B, Daytona, Modal; AWS, Fly.io coming soon) using a single, consistent API with built-in secure tunneling.

Try it out for free here: https://www.computesdk.com/docs/getting-started/quick-start/

Today we are proud to announce the Compute CLI.

I built ComputeSDK after shipping a K8s-based sandbox product for a large enterprise client. Shortly after, when I started talking to AI app builders, I noticed a common frustration: they all ended up building integrations with multiple sandbox providers. Each integration meant rewriting their application, custom proxies, and routing.

As I talked with more customers, I realized sandbox providers and app builders had massive architectural overlap. Both were having to write their own custom terminals, file-watchers, proxies, and more to interact with their sandboxes.

At this point, I realized there was a missing piece in the development infrastructure stack: a way to decouple your application from your sandbox provider while still getting full, live interactivity.

Our original K8s architecture had three services: an API, a gateway, and a sidecar.\nWe extracted our sidecar component and rebuilt it as the "Compute" cli, a lightweight service that installs in any sandbox and automatically creates a secure tunnel to your app. Then we wrapped it with a universal SDK that works across providers. Now you write your sandbox implementation once, and it works everywhere: from E2B, Daytona, & Modal to general cloud platforms like Railway, Fly.io, & AWS.

ComputeSDK provides a consistent interface for sandbox creation and management. It gives you instant access to terminals, file systems, real-time file watching, and a WebSocket event system from any client (server, browser, application). You can switch or add providers easily by importing the provider module and updating environment variables.

We'd love to get any feedback from the community, ideas, or experiences building similar tools.

Thanks!","title":"Show HN: Compute CLI \u2013 A universal sandbox SDK with direct browser access","updated_at":"2026-03-05T22:58:32Z","url":"https://www.computesdk.com/blog/november-2025-update/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"abhegd"},"story_text":{"matchLevel":"none","matchedWords":[],"value":"Hey HN! I'm Abhishek. I\u2019m a designer and product builder. I\u2019ve been building AI products for the past year after leaving my design job, where I worked on dev tools being used by 1M+ engineers.

Previously I built Allyce AI (https://allyce.ai/), which helps answer questions, qualify leads, and handle complaints on websites (AI for sales/SDR + CS). When I started as an independent builder, the first thing I built was a sophisticated GraphRAG pipeline that powers Allyce. The customers using Allyce needed something to analyze sales data and ad performance data, which wasn't possible with RAG - that's what led me to build Tabwise as a separate product.

The key insight was that general AI tools fail at data analysis because they skip three critical steps:\n1. Pre-processing: Cleaning data, inferring schema and structuring context specifically for analytical queries\n2. Context engineering: Crafting prompts that understand data relationships, business context, and expected output formats\n3. Post-processing: Converting raw LLM outputs into properly formatted charts, executive summaries, and actionable insights

This pipeline is why Tabwise can consistently outperform ChatGPT (GPT-5 Thinking) and Claude (Sonnet-4) on data analysis tasks - it's not just the model, it's the entire system.

Tech stack: Next.js, Vercel AI SDK, E2B, I automatically route tasks to the best model based on complexity (Powered by Claude sonnet-4 and OSS models via Fireworks AI)

Right now I'm focused on nailing spreadsheets (.xlsx, .csv), but have more data source integrations planned based on what users actually need.

Would love your thoughts."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Tabwise \u2013 AI data analyst that outperforms ChatGPT and Claude"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://tabwise.ai/"}},"_tags":["story","author_abhegd","story_44985515","show_hn"],"author":"abhegd","created_at":"2025-08-22T15:04:54Z","created_at_i":1755875094,"num_comments":0,"objectID":"44985515","points":2,"story_id":44985515,"story_text":"Hey HN! I'm Abhishek. I\u2019m a designer and product builder. I\u2019ve been building AI products for the past year after leaving my design job, where I worked on dev tools being used by 1M+ engineers.

Previously I built Allyce AI (https://allyce.ai/), which helps answer questions, qualify leads, and handle complaints on websites (AI for sales/SDR + CS). When I started as an independent builder, the first thing I built was a sophisticated GraphRAG pipeline that powers Allyce. The customers using Allyce needed something to analyze sales data and ad performance data, which wasn't possible with RAG - that's what led me to build Tabwise as a separate product.

The key insight was that general AI tools fail at data analysis because they skip three critical steps:\n1. Pre-processing: Cleaning data, inferring schema and structuring context specifically for analytical queries\n2. Context engineering: Crafting prompts that understand data relationships, business context, and expected output formats\n3. Post-processing: Converting raw LLM outputs into properly formatted charts, executive summaries, and actionable insights

This pipeline is why Tabwise can consistently outperform ChatGPT (GPT-5 Thinking) and Claude (Sonnet-4) on data analysis tasks - it's not just the model, it's the entire system.

Tech stack: Next.js, Vercel AI SDK, E2B, I automatically route tasks to the best model based on complexity (Powered by Claude sonnet-4 and OSS models via Fireworks AI)

Right now I'm focused on nailing spreadsheets (.xlsx, .csv), but have more data source integrations planned based on what users actually need.

Would love your thoughts.","title":"Show HN: Tabwise \u2013 AI data analyst that outperforms ChatGPT and Claude","updated_at":"2026-03-05T22:33:11Z","url":"https://tabwise.ai/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"rbitar"},"story_text":{"matchLevel":"none","matchedWords":[],"value":"There are a number of companies working on solving micro-VM sandboxes, using Firecracker or libkrun. This includes CodeSandbox, E2B and Microsandbox. One of the major use cases is running AI-generated code in a safe environment, with the promise of fast (~2-300 ms) bootup times, pre-built memory snapshots, and the ability hibernate and wake up instances extremely fast. The downside is these solutions still have network latency, require server side resources, and added complexity when managing cloud environments.

StackBlitz has shown that they can achieve a WASM-based Node environment with WebContainers[1] running fully in the browser, but it's a proprietary solution and requires a commercial license to use. There was a similar effort by the CodeSandbox team called Nodebox [2] which appears to be discontinued, and used browser based API polyfills and clever workarounds to emulate Node.

Is anyone working on an open-source effort to build a WASM based Node-compatible engine? These solutions are particularly useful for running AI-generated code in a browser sandbox with instant feedback. If so, I'd love to hear about your project.

[1] https://webcontainers.io/\n[2] https://github.com/Sandpack/nodebox-runtime"},"title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: Is anyone working on a WASM-based Node engine for the browser?"}},"_tags":["story","author_rbitar","story_44152138","ask_hn"],"author":"rbitar","created_at":"2025-06-01T16:41:56Z","created_at_i":1748796116,"num_comments":0,"objectID":"44152138","points":2,"story_id":44152138,"story_text":"There are a number of companies working on solving micro-VM sandboxes, using Firecracker or libkrun. This includes CodeSandbox, E2B and Microsandbox. One of the major use cases is running AI-generated code in a safe environment, with the promise of fast (~2-300 ms) bootup times, pre-built memory snapshots, and the ability hibernate and wake up instances extremely fast. The downside is these solutions still have network latency, require server side resources, and added complexity when managing cloud environments.

StackBlitz has shown that they can achieve a WASM-based Node environment with WebContainers[1] running fully in the browser, but it's a proprietary solution and requires a commercial license to use. There was a similar effort by the CodeSandbox team called Nodebox [2] which appears to be discontinued, and used browser based API polyfills and clever workarounds to emulate Node.

Is anyone working on an open-source effort to build a WASM based Node-compatible engine? These solutions are particularly useful for running AI-generated code in a browser sandbox with instant feedback. If so, I'd love to hear about your project.

[1] https://webcontainers.io/\n[2] https://github.com/Sandpack/nodebox-runtime","title":"Ask HN: Is anyone working on a WASM-based Node engine for the browser?","updated_at":"2025-09-14T19:44:25Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"mvac"},"story_text":{"matchLevel":"none","matchedWords":[],"value":"We built a tiny web app that lets you connect to any MCP server and chat with it right from the browser.

It spins up MCP servers in E2B sandboxes and converts the stdio to SSE. The chat app executes on the client\u2011side, so your API keys are sent directly and only to their respective services (E2B, model providers) - never to our server.

Since everything is client-side, you need an OpenAI/Anthropic API key and an E2B API key to try it [https://e2b.dev].

Try it here: https://netglade.github.io/mcp-chat/

We made it as a hackathon project and won the 1st price with it, though we're not really sure what would be a meaningful way to move it forward, so we welcome all types of feedback."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Browser\u2011based MCP Chat \u2013 run MCPs without local setup"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://netglade.github.io/mcp-chat/"}},"_tags":["story","author_mvac","story_43716122","show_hn"],"author":"mvac","created_at":"2025-04-17T12:55:34Z","created_at_i":1744894534,"num_comments":0,"objectID":"43716122","points":2,"story_id":43716122,"story_text":"We built a tiny web app that lets you connect to any MCP server and chat with it right from the browser.

It spins up MCP servers in E2B sandboxes and converts the stdio to SSE. The chat app executes on the client\u2011side, so your API keys are sent directly and only to their respective services (E2B, model providers) - never to our server.

Since everything is client-side, you need an OpenAI/Anthropic API key and an E2B API key to try it [https://e2b.dev].

Try it here: https://netglade.github.io/mcp-chat/

We made it as a hackathon project and won the 1st price with it, though we're not really sure what would be a meaningful way to move it forward, so we welcome all types of feedback.","title":"Show HN: Browser\u2011based MCP Chat \u2013 run MCPs without local setup","updated_at":"2025-04-17T13:14:23Z","url":"https://netglade.github.io/mcp-chat/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"olafmol"},"story_text":{"matchLevel":"none","matchedWords":[],"value":"Hi HN! My name is Olaf, I work at CircleCI as a technology advisor in the CTO office, came in through the acquisition of my company Vamp.io (progressive delivery for microservices on k8s) in 2021. Wanted to hear the HN community feedback and thoughts on a project we think could be very interesting when adding AI coding agents to the SDLC and your CI pipelines.

Our team at CircleCI built Chunk sidecars after repeatedly running into the same issue internally: by the time our CI catches a failure, the agent has already moved on and most of the useful context is gone.

The basic idea of Chunk sidecars is to move fast lightweight validation into the inner development loop.

Chunk sidecars runs scoped \u201cmicrobuilds\u201d inside a lightweight microVM that mirrors your CI environment. It tries to auto-detect your stack and test commands, syncs changes from the agent session, and runs validations before commit/push.

A few implementation details that might be interesting:

validation hooks trigger automatically during agent stop/evaluation events

warm snapshots keep startup times low

validations run against environments matching the CI stack instead of local machine state

microbuilds only run the relevant slice instead of the entire pipeline

In our own experiments we measured:

~27 second average microbuild compute

~5 minutes total billable compute for equivalent full CI runs

3x\u20135x lower token usage in retry loops

The compute comparison is billable compute vs billable compute, not wall clock time. Full CI pipelines were parallelized.

The 27s is with warm snapshots \u2014 first-time setup takes about 15 minutes. We tested this on our own pipeline, not a large corpus. Larger repos with heavier deps will vary.

Under the hood it's currently Firecracker microVMs, running on E2B infrastructure. Current spec: 4 CPU, 8GB RAM (comparable to a Docker large). Things can change in the future depending on feedback and learnings.

Short demo video (YT) here: https://circle.ci/4dq9fph

Blog post: https://circleci.com/blog/chunk-sidecars/

Chunk CLI GitHub repo: https://github.com/CircleCI-Public/chunk-cli

This works with any CircleCI account (including the free one), and integrates with Claude Code, Codex, Cursor, or your own agents. The project is open source and also has features that work without CircleCI connected. Simply install the Chunk CLI and run "chunk init" and the sidecar auto-detects your stack and test commands.

Would love all feedback, especially from people already experimenting with agentic workflows. We're especially curious whether others are seeing the same CI failure rate pattern and "widening gap" between inner and outer dev/SDLC loop with agent-generated code?"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Chunk sidecars for validating agent-generated code before pushing to CI"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://circleci.com/blog/chunk-sidecars/"}},"_tags":["story","author_olafmol","story_48281284","show_hn"],"author":"olafmol","children":[48281399],"created_at":"2026-05-26T15:41:32Z","created_at_i":1779810092,"num_comments":2,"objectID":"48281284","points":1,"story_id":48281284,"story_text":"Hi HN! My name is Olaf, I work at CircleCI as a technology advisor in the CTO office, came in through the acquisition of my company Vamp.io (progressive delivery for microservices on k8s) in 2021. Wanted to hear the HN community feedback and thoughts on a project we think could be very interesting when adding AI coding agents to the SDLC and your CI pipelines.

Our team at CircleCI built Chunk sidecars after repeatedly running into the same issue internally: by the time our CI catches a failure, the agent has already moved on and most of the useful context is gone.

The basic idea of Chunk sidecars is to move fast lightweight validation into the inner development loop.

Chunk sidecars runs scoped \u201cmicrobuilds\u201d inside a lightweight microVM that mirrors your CI environment. It tries to auto-detect your stack and test commands, syncs changes from the agent session, and runs validations before commit/push.

A few implementation details that might be interesting:

validation hooks trigger automatically during agent stop/evaluation events

warm snapshots keep startup times low

validations run against environments matching the CI stack instead of local machine state

microbuilds only run the relevant slice instead of the entire pipeline

In our own experiments we measured:

~27 second average microbuild compute

~5 minutes total billable compute for equivalent full CI runs

3x\u20135x lower token usage in retry loops

The compute comparison is billable compute vs billable compute, not wall clock time. Full CI pipelines were parallelized.

The 27s is with warm snapshots \u2014 first-time setup takes about 15 minutes. We tested this on our own pipeline, not a large corpus. Larger repos with heavier deps will vary.

Under the hood it's currently Firecracker microVMs, running on E2B infrastructure. Current spec: 4 CPU, 8GB RAM (comparable to a Docker large). Things can change in the future depending on feedback and learnings.

Short demo video (YT) here: https://circle.ci/4dq9fph

Blog post: https://circleci.com/blog/chunk-sidecars/

Chunk CLI GitHub repo: https://github.com/CircleCI-Public/chunk-cli

This works with any CircleCI account (including the free one), and integrates with Claude Code, Codex, Cursor, or your own agents. The project is open source and also has features that work without CircleCI connected. Simply install the Chunk CLI and run "chunk init" and the sidecar auto-detects your stack and test commands.

Would love all feedback, especially from people already experimenting with agentic workflows. We're especially curious whether others are seeing the same CI failure rate pattern and "widening gap" between inner and outer dev/SDLC loop with agent-generated code?","title":"Show HN: Chunk sidecars for validating agent-generated code before pushing to CI","updated_at":"2026-05-26T16:01:04Z","url":"https://circleci.com/blog/chunk-sidecars/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"scosman"},"story_text":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/scosman/Biscotti

I got tired of Electron meeting-transcription apps that upload my private audio to a server to do transcription/summary my MacBook could do locally in seconds. So I built Biscotti: a free, local meeting recorder. Calendar integration, voice isolation, speaker ID, AI summaries, native UI \u2014 and nothing ever leaves your device.

It's completely private: all AI runs locally, all data stored locally, nothing leaves your device, no cloud. It's native macOS (SwiftUI, not Electron), so it's smaller, faster, and more intuitive. And no bot ever joins your call - it uses system audio APIs, so it works with every app.

It integrates deeply with your Mac. It watches for meetings via audio APIs and offers to record when one starts, then automatically stops when the meeting ends (cancelable). It reads your calendar to see upcoming meetings, offers to start recording when they begin, and uses that metadata to help identify speakers. Speaker ID combines diarization, calendar metadata, and the transcript to figure out who's who automatically. Voice isolation separates your mic from your speaker output for clean audio.

The AI is customizable: want it to generate action items? A Slack-style summary? Talk like a pirate? Just edit the summary prompt. There's also markdown notes, and it keeps compressed audio files so you can re-transcribe later.

It's completely free. Under the PolyForm Perimeter License \u2014 fairly permissive (but not OSI open source).

Models and SDKs:

  Transcription \u2014 Whisper Large V3 Turbo via WhisperKit\n  Summarization \u2014 Gemma 4 via llama.cpp (12B or E2B depending on your memory/processor)\n  Speaker ID (diarization) \u2014 Pyannote via SpeakerKit\n
\nTech stack:

  Swift and SwiftUI\n  XPC services for all AI processing, so the app stays tiny and stable and AI memory releases immediately\n  System APIs: CoreAudio for audio, EventKit for calendars\n
\nhttps://github.com/scosman/Biscotti"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Biscotti \u2013 on-device meeting transcription for macOS"}},"_tags":["story","author_scosman","story_48894351","show_hn"],"author":"scosman","children":[48968800],"created_at":"2026-07-13T15:38:26Z","created_at_i":1783957106,"num_comments":1,"objectID":"48894351","points":1,"story_id":48894351,"story_text":"https://github.com/scosman/Biscotti

I got tired of Electron meeting-transcription apps that upload my private audio to a server to do transcription/summary my MacBook could do locally in seconds. So I built Biscotti: a free, local meeting recorder. Calendar integration, voice isolation, speaker ID, AI summaries, native UI \u2014 and nothing ever leaves your device.

It's completely private: all AI runs locally, all data stored locally, nothing leaves your device, no cloud. It's native macOS (SwiftUI, not Electron), so it's smaller, faster, and more intuitive. And no bot ever joins your call - it uses system audio APIs, so it works with every app.

It integrates deeply with your Mac. It watches for meetings via audio APIs and offers to record when one starts, then automatically stops when the meeting ends (cancelable). It reads your calendar to see upcoming meetings, offers to start recording when they begin, and uses that metadata to help identify speakers. Speaker ID combines diarization, calendar metadata, and the transcript to figure out who's who automatically. Voice isolation separates your mic from your speaker output for clean audio.

The AI is customizable: want it to generate action items? A Slack-style summary? Talk like a pirate? Just edit the summary prompt. There's also markdown notes, and it keeps compressed audio files so you can re-transcribe later.

It's completely free. Under the PolyForm Perimeter License \u2014 fairly permissive (but not OSI open source).

Models and SDKs:

  Transcription \u2014 Whisper Large V3 Turbo via WhisperKit\n  Summarization \u2014 Gemma 4 via llama.cpp (12B or E2B depending on your memory/processor)\n  Speaker ID (diarization) \u2014 Pyannote via SpeakerKit\n
\nTech stack:

  Swift and SwiftUI\n  XPC services for all AI processing, so the app stays tiny and stable and AI memory releases immediately\n  System APIs: CoreAudio for audio, EventKit for calendars\n
\nhttps://github.com/scosman/Biscotti","title":"Show HN: Biscotti \u2013 on-device meeting transcription for macOS","updated_at":"2026-07-19T15:05:34Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"bianheshan"},"story_text":{"matchLevel":"none","matchedWords":[],"value":"Hi HN,

I'm Heshan, founder of X-Pilot. We're building an AI Video Generator for online courses and educational content. Unlike most text-to-video generator that render videos directly from models (which often produce random stock footage unrelated to the actual content), we take a code-first approach: generate editable code layers, let users verify/refine them, then render to video.

The Problem We're Solving

Most AI video generators treat "education" and "marketing" the same\u2014they optimize for "looks good" rather than "logically accurate." When you feed a technical tutorial or course script into a generic video AI, you get:\n- Random B-roll that doesn't match the concept being explained\n- Incorrect visualizations (e.g., showing a "for loop" diagram when explaining recursion)\n- No way to systematically fix errors without regenerating everything

For educators, corporate trainers, and knowledge creators, accuracy matters more than aesthetics. A single incorrect diagram can break a learner's mental model.

Our Approach: Code as the Intermediate Layer

Instead of text \u2192 video blackbox, we do:\nText/PDF/Doc \u2192 Structured Code (Remotion + Visual Box Engine) \u2192 Editable Preview \u2192 Final Render

Tech Stack\n- Agent orchestration: LangGraph (with Gemini 2.5 Flash for planning, reasoning, and content structuring)\n- Video Code generation model: Gemini3.0 for Remotion Code & Veo 3 (for generative footage where needed)\n- Code-based rendering: Remotion (React-based video framework)\n- Knowledge visualization engine: Our own "Visual Box Engine"\u2014a library of parameterized educational animation components (flowcharts, comparisons, step-by-step sequences, system diagrams, etc.)\n- Voice synthesis: Fish Audio (for natural narration)\n- Rendering: Google Cloud (distributed video rendering using chrome headless)\n- Code execution sandbox: E2B (for safe, isolated code execution during generation and preview\uff0c but we will update to our own sandbox\uff0c because e2b offen time out\uff0cand low performance for bundle and render)

Why Remotion + Custom Components\uff1f\nWe chose Remotion because:\n1. Editability: Every visual element is React code. Users (or our AI agents) can modify text, swap components, adjust timing\u2014without touching raw video files.\n2. Reproducibility: Same input \u2192 same output. No model randomness in final render.\n3. Composability: We built a "Visual Box" library\u2014reusable animation patterns for education (e.g., "cause-and-effect flow," "comparison table," "hierarchical breakdown"). These aren't generic motion graphics; they're designed around pedagogical principles.

The trade-off: We sacrifice some "cinematic quality" for logical accuracy and user control. Right now, output can feel closer to "animated slides" than "documentary footage"\u2014which is actually our biggest unsolved challenge (more on that below).

What We're Struggling With (and Planning to Fix)

1. Code Error Rate\nGenerating Remotion code via LLMs is powerful but error-prone. \n2. Limited Asset Handling\nRight now, if a user wants to insert a custom image/GIF/video mid-generation, they need to upload \u2192 we process \u2192 regenerate. This breaks flow.\n3. The "PPT Feel" Problem\nThis is the hardest one. Because we prioritize structure and editability, our videos can feel like "animated PowerPoint" rather than "produced content."

We're experimenting with:\n- Hybrid rendering: Use generative video (Veo) for transitions/B-roll, but keep Visual Boxes for core explanations\n- Cinematic presets: Camera movements, depth effects, color grading\u2014applied as composable layers\n- Motion design constraints: Teaching our agent to follow motion design principles (easing curves, visual hierarchy, pacing)

Honest question for HN: Has anyone solved this trade-off between "programmatically editable" and "cinematic quality"? I'd love to hear how others have approached it (especially in contexts where correctness > vibes)."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: X-Pilot \u2013 Code-Driven AI Video Generator for Online Courses"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://www.x-pilot.ai/"}},"_tags":["story","author_bianheshan","story_46705382","show_hn"],"author":"bianheshan","children":[46705553],"created_at":"2026-01-21T13:24:33Z","created_at_i":1769001873,"num_comments":1,"objectID":"46705382","points":1,"story_id":46705382,"story_text":"Hi HN,

I'm Heshan, founder of X-Pilot. We're building an AI Video Generator for online courses and educational content. Unlike most text-to-video generator that render videos directly from models (which often produce random stock footage unrelated to the actual content), we take a code-first approach: generate editable code layers, let users verify/refine them, then render to video.

The Problem We're Solving

Most AI video generators treat "education" and "marketing" the same\u2014they optimize for "looks good" rather than "logically accurate." When you feed a technical tutorial or course script into a generic video AI, you get:\n- Random B-roll that doesn't match the concept being explained\n- Incorrect visualizations (e.g., showing a "for loop" diagram when explaining recursion)\n- No way to systematically fix errors without regenerating everything

For educators, corporate trainers, and knowledge creators, accuracy matters more than aesthetics. A single incorrect diagram can break a learner's mental model.

Our Approach: Code as the Intermediate Layer

Instead of text \u2192 video blackbox, we do:\nText/PDF/Doc \u2192 Structured Code (Remotion + Visual Box Engine) \u2192 Editable Preview \u2192 Final Render

Tech Stack\n- Agent orchestration: LangGraph (with Gemini 2.5 Flash for planning, reasoning, and content structuring)\n- Video Code generation model: Gemini3.0 for Remotion Code & Veo 3 (for generative footage where needed)\n- Code-based rendering: Remotion (React-based video framework)\n- Knowledge visualization engine: Our own "Visual Box Engine"\u2014a library of parameterized educational animation components (flowcharts, comparisons, step-by-step sequences, system diagrams, etc.)\n- Voice synthesis: Fish Audio (for natural narration)\n- Rendering: Google Cloud (distributed video rendering using chrome headless)\n- Code execution sandbox: E2B (for safe, isolated code execution during generation and preview\uff0c but we will update to our own sandbox\uff0c because e2b offen time out\uff0cand low performance for bundle and render)

Why Remotion + Custom Components\uff1f\nWe chose Remotion because:\n1. Editability: Every visual element is React code. Users (or our AI agents) can modify text, swap components, adjust timing\u2014without touching raw video files.\n2. Reproducibility: Same input \u2192 same output. No model randomness in final render.\n3. Composability: We built a "Visual Box" library\u2014reusable animation patterns for education (e.g., "cause-and-effect flow," "comparison table," "hierarchical breakdown"). These aren't generic motion graphics; they're designed around pedagogical principles.

The trade-off: We sacrifice some "cinematic quality" for logical accuracy and user control. Right now, output can feel closer to "animated slides" than "documentary footage"\u2014which is actually our biggest unsolved challenge (more on that below).

What We're Struggling With (and Planning to Fix)

1. Code Error Rate\nGenerating Remotion code via LLMs is powerful but error-prone. \n2. Limited Asset Handling\nRight now, if a user wants to insert a custom image/GIF/video mid-generation, they need to upload \u2192 we process \u2192 regenerate. This breaks flow.\n3. The "PPT Feel" Problem\nThis is the hardest one. Because we prioritize structure and editability, our videos can feel like "animated PowerPoint" rather than "produced content."

We're experimenting with:\n- Hybrid rendering: Use generative video (Veo) for transitions/B-roll, but keep Visual Boxes for core explanations\n- Cinematic presets: Camera movements, depth effects, color grading\u2014applied as composable layers\n- Motion design constraints: Teaching our agent to follow motion design principles (easing curves, visual hierarchy, pacing)

Honest question for HN: Has anyone solved this trade-off between "programmatically editable" and "cinematic quality"? I'd love to hear how others have approached it (especially in contexts where correctness > vibes).","title":"Show HN: X-Pilot \u2013 Code-Driven AI Video Generator for Online Courses","updated_at":"2026-03-05T23:22:20Z","url":"https://www.x-pilot.ai/"}],"hitsPerPage":50,"nbHits":50226,"nbPages":20,"page":0,"params":"query=e2b&tags=story&hitsPerPage=50&advancedSyntax=true&analyticsTags=backend","processingTimeMS":18,"processingTimingsMS":{"_request":{"queue":5,"roundTrip":17},"afterFetch":{"format":{"highlighting":4,"total":4},"merge":{"total":1},"total":1},"fetch":{"query":3,"scanning":12,"total":16},"total":18},"query":"e2b","serverTimeMS":28}