[{"additionalPlain":"Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.\nEqual Opportunity\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.\nAccessibility & Accommodations\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.\n","additional":"
Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"US - Headquarters","team":"Orchestration","allLocations":["US - Headquarters"]},"createdAt":1784695295232,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
The role
\nYou will set the multi-year technical vision for the platform — and prove it works by building it. This is not a strategy-only role. You perform competency analysis across the architecture, build working prototypes to validate direction, write production code, and lead cross-functional delivery. You are accountable to stakeholders for the quality and viability of what ships. You work hand-in-hand with Product to ensure the technical vision maps to customer value.
\nWhat you'll work on
\n\n
Nice to have
\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"The role
\nYou will set the multi-year technical vision for the platform — and prove it works by building it. This is not a strategy-only role. You perform competency analysis across the architecture, build working prototypes to validate direction, write production code, and lead cross-functional delivery. You are accountable to stakeholders for the quality and viability of what ships. You work hand-in-hand with Product to ensure the technical vision maps to customer value.
\nWhere you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"US - Headquarters","team":"Program Management","allLocations":["US - Headquarters"]},"createdAt":1782857495793,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n\n\nRole Overview \nThe Engineering Program Manager (EPM) will own and drive cross-functional programs across hardware and software development for networking products including switches and ASICs. This individual will lead New Product Introduction (NPI) processes through software phase-gate milestones, own software release schedules, and drive alignment across engineering, product management, and executive stakeholders. This is an opportunity to shape and scale program management practices at a rapidly scaling company, with broad ownership from day one, direct visibility to executive leadership, and the expectation to mentor and grow a program management discipline as the team expands.\nThe ideal candidate thrives in fast-paced environments, brings exceptional organizational and leadership skills, and has deep expertise in networking technologies and operating systems. You are effective at leading teams with varying levels of systems and networking experience, and you can translate complex technical context into clear, actionable information for diverse stakeholders. You bring a track record of driving programs to successful outcomes and influencing without authority across engineering organizations.\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Role Overview
\nThe Engineering Program Manager (EPM) will own and drive cross-functional programs across hardware and software development for networking products including switches and ASICs. This individual will lead New Product Introduction (NPI) processes through software phase-gate milestones, own software release schedules, and drive alignment across engineering, product management, and executive stakeholders. This is an opportunity to shape and scale program management practices at a rapidly scaling company, with broad ownership from day one, direct visibility to executive leadership, and the expectation to mentor and grow a program management discipline as the team expands.
\nThe ideal candidate thrives in fast-paced environments, brings exceptional organizational and leadership skills, and has deep expertise in networking technologies and operating systems. You are effective at leading teams with varying levels of systems and networking experience, and you can translate complex technical context into clear, actionable information for diverse stakeholders. You bring a track record of driving programs to successful outcomes and influencing without authority across engineering organizations.
\nProgram & Project Management
\nSupport end-to-end program execution for hardware NPI and/or software release programs, from concept through production release and customer delivery.
\nHelp develop and maintain integrated program schedules, track milestones, manage risks, and communicate status to stakeholders and executives.
\nCoordinate software NPI phase-gate reviews (Phase 0: Concept, Phase 1: Commit, Phase 2: Execution, Phase 3: Testing, Phase 4: First Revenue Shipment) and ensure deliverables meet quality, cost, and schedule targets at each gate.
\nFacilitate cross-functional meetings, capture action items, and drive accountability across teams.
\nContribute to defining and tracking program and release KPIs (e.g., release cadence, defect escape rate, schedule adherence) to support leadership decision-making.
\nCross-Functional Collaboration
\nPartner closely with Product Managers (PM), software engineering, QA, operations, and ASIC/silicon teams to align priorities and resolve dependencies.
\nServe as a coordination point between hardware and software workstreams, helping ensure integration readiness at each milestone.
\nBridge knowledge gaps across teams by translating hardware constraints, system-level dependencies, and NOS release requirements into actionable context for engineering leaders and stakeholders who may not have deep systems or networking backgrounds.
\nCommunicate program status, risks, and trade-offs to executive leadership, providing clear recommendations for decision-making.
\nTechnical Coordination & Release Support
\nSupport software releases for network operating systems (e.g., SONiC) running on switches, including release readiness reviews, code freeze coordination, and integration testing.
\nTrack feature readiness, bug triage, and release criteria across software and firmware teams.
\nSupport ASIC bring-up and integration milestones by helping align hardware and software schedules.
\nMaintain awareness of platform-level dependencies across network operating systems such as SONiC and related NOS environments.
\nAssist in maintaining release governance processes, including branching strategies, freeze windows, and rollback procedures in collaboration with engineering leads.
\nSupport post-release retrospectives and help capture lessons learned to improve future release cycles.
\nProcess Improvement
\nIdentify and recommend process improvements to increase program execution efficiency and predictability.
\nHelp develop dashboards, templates, and reporting mechanisms for program health and risk visibility.
\nContribute to building scalable program management practices as the organization grows.
\nHelp ensure systems are in place to capture and act on product feedback from customers and internal stakeholders.
\n• Own end-to-end program execution for software NPI and release programs, from concept through production release and customer delivery.
\n• Develop and maintain integrated program schedules, track milestones, manage risks, and communicate status to stakeholders and executives.
\n• Lead software NPI phase-gate reviews (Phase 0: Concept, Phase 1: Commit, Phase 2: Execution, Phase 3: Testing, Phase 4: First Revenue Shipment) and ensure deliverables meet quality, cost, and schedule targets at each gate.
\n• Lead cross-functional meetings, capture action items, and drive accountability across teams.
\n• Define and own program and release KPIs (e.g., release cadence, defect escape rate, schedule adherence) to drive leadership decision-making and continuous improvement.
\n• Lead cross-functional alignment with Product Managers (PM), software engineering, QA, operations, and ASIC/silicon teams to drive priorities and resolve dependencies.
\n• Serve as the primary integration point between hardware and software workstreams, ensuring integration readiness at each milestone.
\n• Bridge knowledge gaps across teams by translating hardware constraints, system-level dependencies, and NOS release requirements into actionable context for engineering leaders and stakeholders who may not have deep systems or networking backgrounds.
\n• Own executive communication of program status, risks, and trade-offs, providing clear recommendations and escalation paths for decision-making.
\n• Drive software releases for network operating systems (e.g., SONiC) running on switches, including release readiness reviews, code freeze coordination, and integration testing.
\n• Own feature readiness tracking, bug triage prioritization, and release criteria enforcement across software and firmware teams.
\n• Drive ASIC bring-up and integration milestones by aligning hardware and software schedules and proactively resolving cross-team blockers.
\n• Maintain deep understanding of platform-level dependencies across network operating systems such as SONiC and related NOS environments.
\n• Own release governance processes, including branching strategies, freeze windows, and rollback procedures in collaboration with engineering leads.
\n• Lead post-release retrospectives and ensure lessons learned are captured, tracked, and applied to improve future release cycles.
\n• Drive process improvements to increase program execution efficiency, predictability, and scalability across the organization.
\n• Build and own dashboards, templates, and reporting mechanisms for program health and risk visibility.
\n• Establish and scale program management practices, frameworks, and tooling as the organization grows.
\n• Ensure systems are in place to capture and act on product feedback from customers and internal stakeholders.
\n• Bachelor's degree in Engineering, Computer Science, or a related technical field.
\n• 7+ years of experience in engineering program management, technical project management, or a related role within a hardware, software, or networking company.
\n• Deep experience with software NPI phase-gate processes (e.g., Concept, Commit, Execution, Testing, FRS) or equivalent milestone-driven development in Agile environments.
\n• Strong command of the software release lifecycle, including feature planning, code freeze, regression testing, and GA releases.
\n• Proven track record of leading cross-functional programs involving software engineering, PM, QA, and executive leadership—particularly in environments where not all stakeholders share the same technical background.
\n• Executive presence with exceptional written and verbal communication skills; demonstrated ability to distill complex technical programs into clear status updates and executive summaries.
\n• Proficiency with project management tools such as Jira, Confluence, Smartsheet, or equivalent.
\n• Demonstrated ability to influence without authority and drive accountability across teams and organizational boundaries.
\n• Significant experience in the networking or data center infrastructure industry.
\n• Hands-on experience with network operating systems (e.g., SONiC, Cumulus, NX-OS, EOS) or Linux-based networking platforms.
\n• Experience with AI/ML infrastructure or high-performance computing networking environments.
\n• Strong background in Agile methodologies and SDLC practices adapted to hardware-coupled development cycles.
\n• Proven ownership of release governance processes, including branching strategies, freeze windows, and rollback procedures.
\n• Track record of building program/release KPI frameworks and driving continuous improvement initiatives.
\n• Deep understanding of open-source networking platforms (SONiC, Open Network Linux).
\n• Experience mentoring or growing junior program managers and establishing PM practices at scaling organizations.
\n• PMP or Agile/Scrum certification.
\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Role Overview
\nThe Engineering Program Manager (EPM) will own and drive cross-functional programs across hardware and software development for networking products including switches and ASICs. This individual will lead New Product Introduction (NPI) processes through software phase-gate milestones, own software release schedules, and drive alignment across engineering, product management, and executive stakeholders. This is an opportunity to shape and scale program management practices at a rapidly scaling company, with broad ownership from day one, direct visibility to executive leadership, and the expectation to mentor and grow a program management discipline as the team expands.
\nThe ideal candidate thrives in fast-paced environments, brings exceptional organizational and leadership skills, and has deep expertise in networking technologies and operating systems. You are effective at leading teams with varying levels of systems and networking experience, and you can translate complex technical context into clear, actionable information for diverse stakeholders. You bring a track record of driving programs to successful outcomes and influencing without authority across engineering organizations.
\nWhere you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"India - Bangalore","team":"AI Network Software","allLocations":["India - Bangalore"]},"createdAt":1781409292294,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n\n\nUpscale is building the engineering team that drives product improvements though hands-on field engineering. The Forward-Deployed Engineer (FDE) owns complex fabric and system issues end-to-end: from first report through root-cause analysis, workaround delivery, and verified fix. You resolve them with the rest of the team, or you drive them to resolution across whatever boundary stands in the way. \nThis role is new to UpscaleAI and sits at the intersection of support, engineering, and product management. You will work directly with customers running production AI fabrics, reproduce issues in the lab, develop workarounds under pressure, and contribute fixes and diagnostic content back into the product and our AI-driven monitoring agent. \nThe environment is often unstructured. Problem definitions are incomplete. Documentation may not exist yet. The right candidate sees that as an opportunity, not an obstacle. \nLarge AI infrastructure operators increasingly expect their vendor's engineering team to function as an extension of their own infrastructure organization. This role is built to meet that expectation. \nThis is a small team. There will be an on-call component, but we are building a global team to reduce out-of-hours calls. We work with customers directly, so some travel is involved. Remote meeting tools handle much of the collaboration but building strong trust relationships will require time on site. \n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Upscale is building the engineering team that drives product improvements though hands-on field engineering. The Forward-Deployed Engineer (FDE) owns complex fabric and system issues end-to-end: from first report through root-cause analysis, workaround delivery, and verified fix. You resolve them with the rest of the team, or you drive them to resolution across whatever boundary stands in the way.
\nThis role is new to UpscaleAI and sits at the intersection of support, engineering, and product management. You will work directly with customers running production AI fabrics, reproduce issues in the lab, develop workarounds under pressure, and contribute fixes and diagnostic content back into the product and our AI-driven monitoring agent.
\nThe environment is often unstructured. Problem definitions are incomplete. Documentation may not exist yet. The right candidate sees that as an opportunity, not an obstacle.
\nLarge AI infrastructure operators increasingly expect their vendor's engineering team to function as an extension of their own infrastructure organization. This role is built to meet that expectation.
\nThis is a small team. There will be an on-call component, but we are building a global team to reduce out-of-hours calls. We work with customers directly, so some travel is involved. Remote meeting tools handle much of the collaboration but building strong trust relationships will require time on site.
\nTriage and resolve complex AI fabric issues: including silent packet drops, queue anomalies, NCCL stalls, gray failures, and performance degradation in production environments. You own the problem from intake to resolution.
\nReproduce issues in the lab: building test scenarios that isolate L2 and L3 problems (packet loss, latency, retransmits, ECMP behavior, PFC/ECN interactions) and deliver reproducible cases to the development team when code fixes are needed.
\nAct as the escalation buffer for engineering: resolving issues without engaging development when possible, and packaging clean, reproducible problem statements when development engagement is required.
\nTrain and improve the on-box and off-box AI agents: contributing field-validated detection signatures, classification logic, and resolution recommendations based on real cases. Your field experience directly shapes what the agent can identify and handle autonomously.
\nValidate AI agent accuracy: reviewing the agent's data collection, anomaly detection, and triage classifications against real-world outcomes. You determine when the agent is ready to advance from data collection to active triage to mitigation
\nWork with customers directly: understanding their fabric topology, workload patterns, and operational constraints. Communicate findings clearly to both technical and non-technical stakeholders.
\nContribute to the product: identifying supportability gaps, proposing diagnostic improvements, and feeding field insights into the product management process. You are an active voice in what the product needs to become.
\nMaintain and evolve the AI agent's detection capabilities: update detection signatures, retrain classification models, and tune thresholds as customer fabrics scale, new hardware is deployed, and new failure modes are discovered in the field. The AI agent is an evolving system, not a shipped product.
\nRead and work with source code: engaging with developers at the code level when needed to understand behavior, identify root causes, or validate fixes. You do not need to be a full-time developer, but you must be comfortable in the codebase.
\n
Requirements
5+ years of hands-on experience with high speed data center switching platforms at scale. Experience with AI fabric deployments is a strong plus but not required if switching and troubleshooting background depth is there.
\nPeople who take ownership. The FDE role requires someone who treats every problem as theirs until it's resolved, regardless of where the root cause sits. If the issue crosses into ASIC behavior, NOS code, optics firmware, or customer configuration, you follow it there. Pointing to another team is not a resolution.
\nDeep Ethernet troubleshooting experience, including L2/L3 forwarding, ECMP, and packet-level analysis
\nBGP operations experience, including route reflectors, convergence behavior, and fabric-scale deployments
\nQoS configuration and troubleshooting: PFC, ECN, DSCP, queue management
\nHardware support background: ASICs, FPGAs, chassis-based platforms, pluggable optic modules
\nFiber and optical link troubleshooting, including DOM telemetry interpretation
\nSoftware support history: working with NOS, firmware, and driver-level issues
\nLinux proficiency (command line, system administration, log analysis)
\nTelemetry and monitoring: experience with streaming telemetry, counters, and event-driven diagnostics
\nExperience contributing to or training ML/AI systems (classification models, labeled data, feedback loops)
\nLab skills: ability to design, build, and execute complex test scenarios that isolate specific failure modes
\nClear written and verbal communication, including customer-facing interaction
\nNice to have
\nExperience with SONiC or other open network operating systems
\nFamiliarity with AI data center designs: low-latency fabrics, GPU cluster networking, NCCL, RDMA/RoCEv2
\nExperience with NVIDIA Spectrum switch family, BlueField SuperNICs, or ConnectX adapters
\nUnderstanding of Ultra Ethernet Consortium specifications and goals
\nPCIe architecture knowledge (relevant to NIC and accelerator integration)
\nExperience with Ixia/Keysight or similar network test equipment and test suites
\nExposure to high-frequency telemetry (HFT) or WJH (What Just Happened) event data
\nWhat This Role Is Not
\nThis is not a traditional TAC or support queue role. You will not be working from runbooks or following scripts. Many of the problems you encounter will not have documented solutions. You will be expected to figure it out, document it, and make sure the next engineer (or the agent) can handle it, next time, without you.
\nThis is also not a development role. You will read code and work closely with developers, but your primary output is resolved customer issues, validated agent training data, and product improvement recommendations, not shipped features.
\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Upscale is building the engineering team that drives product improvements though hands-on field engineering. The Forward-Deployed Engineer (FDE) owns complex fabric and system issues end-to-end: from first report through root-cause analysis, workaround delivery, and verified fix. You resolve them with the rest of the team, or you drive them to resolution across whatever boundary stands in the way.
\nThis role is new to UpscaleAI and sits at the intersection of support, engineering, and product management. You will work directly with customers running production AI fabrics, reproduce issues in the lab, develop workarounds under pressure, and contribute fixes and diagnostic content back into the product and our AI-driven monitoring agent.
\nThe environment is often unstructured. Problem definitions are incomplete. Documentation may not exist yet. The right candidate sees that as an opportunity, not an obstacle.
\nLarge AI infrastructure operators increasingly expect their vendor's engineering team to function as an extension of their own infrastructure organization. This role is built to meet that expectation.
\nThis is a small team. There will be an on-call component, but we are building a global team to reduce out-of-hours calls. We work with customers directly, so some travel is involved. Remote meeting tools handle much of the collaboration but building strong trust relationships will require time on site.
\nWhere you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"US - Headquarters","team":"AI Network Software","allLocations":["US - Headquarters"]},"createdAt":1781304223490,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n\n\nUpscale is building the engineering team that drives product improvements though hands-on field engineering. The Forward-Deployed Engineer (FDE) owns complex fabric and system issues end-to-end: from first report through root-cause analysis, workaround delivery, and verified fix. You resolve them with the rest of the team, or you drive them to resolution across whatever boundary stands in the way. \nThis role is new to UpscaleAI and sits at the intersection of support, engineering, and product management. You will work directly with customers running production AI fabrics, reproduce issues in the lab, develop workarounds under pressure, and contribute fixes and diagnostic content back into the product and our AI-driven monitoring agent. \nThe environment is often unstructured. Problem definitions are incomplete. Documentation may not exist yet. The right candidate sees that as an opportunity, not an obstacle. \nLarge AI infrastructure operators increasingly expect their vendor's engineering team to function as an extension of their own infrastructure organization. This role is built to meet that expectation. \nThis is a small team. There will be an on-call component, but we are building a global team to reduce out-of-hours calls. We work with customers directly, so some travel is involved. Remote meeting tools handle much of the collaboration but building strong trust relationships will require time on site. \n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Upscale is building the engineering team that drives product improvements though hands-on field engineering. The Forward-Deployed Engineer (FDE) owns complex fabric and system issues end-to-end: from first report through root-cause analysis, workaround delivery, and verified fix. You resolve them with the rest of the team, or you drive them to resolution across whatever boundary stands in the way.
\nThis role is new to UpscaleAI and sits at the intersection of support, engineering, and product management. You will work directly with customers running production AI fabrics, reproduce issues in the lab, develop workarounds under pressure, and contribute fixes and diagnostic content back into the product and our AI-driven monitoring agent.
\nThe environment is often unstructured. Problem definitions are incomplete. Documentation may not exist yet. The right candidate sees that as an opportunity, not an obstacle.
\nLarge AI infrastructure operators increasingly expect their vendor's engineering team to function as an extension of their own infrastructure organization. This role is built to meet that expectation.
\nThis is a small team. There will be an on-call component, but we are building a global team to reduce out-of-hours calls. We work with customers directly, so some travel is involved. Remote meeting tools handle much of the collaboration but building strong trust relationships will require time on site.
\nKey Responsibilities
\nTriage and resolve complex AI fabric issues: including silent packet drops, queue anomalies, NCCL stalls, gray failures, and performance degradation in production environments. You own the problem from intake to resolution.
\nReproduce issues in the lab: building test scenarios that isolate L2 and L3 problems (packet loss, latency, retransmits, ECMP behavior, PFC/ECN interactions) and deliver reproducible cases to the development team when code fixes are needed.
\nAct as the escalation buffer for engineering: resolving issues without engaging development when possible, and packaging clean, reproducible problem statements when development engagement is required.
\nTrain and improve the on-box and off-box AI agents: contributing field-validated detection signatures, classification logic, and resolution recommendations based on real cases. Your field experience directly shapes what the agent can identify and handle autonomously.
\nValidate AI agent accuracy: reviewing the agent's data collection, anomaly detection, and triage classifications against real-world outcomes. You determine when the agent is ready to advance from data collection to active triage to mitigation
\nWork with customers directly: understanding their fabric topology, workload patterns, and operational constraints. Communicate findings clearly to both technical and non-technical stakeholders.
\nContribute to the product: identifying supportability gaps, proposing diagnostic improvements, and feeding field insights into the product management process. You are an active voice in what the product needs to become.
\nMaintain and evolve the AI agent's detection capabilities: update detection signatures, retrain classification models, and tune thresholds as customer fabrics scale, new hardware is deployed, and new failure modes are discovered in the field. The AI agent is an evolving system, not a shipped product.
\nRead and work with source code: engaging with developers at the code level when needed to understand behavior, identify root causes, or validate fixes. You do not need to be a full-time developer, but you must be comfortable in the codebase.
\nRequirements
\nDeep Ethernet troubleshooting experience, including L2/L3 forwarding, ECMP, and packet-level analysis
\nBGP operations experience, including route reflectors, convergence behavior, and fabric-scale deployments
\nQoS configuration and troubleshooting: PFC, ECN, DSCP, queue management
\nHardware support background: ASICs, FPGAs, chassis-based platforms, pluggable optic modules
\nFiber and optical link troubleshooting, including DOM telemetry interpretation
\nSoftware support history: working with NOS, firmware, and driver-level issues
\nLinux proficiency (command line, system administration, log analysis)
\nTelemetry and monitoring: experience with streaming telemetry, counters, and event-driven diagnostics
\nExperience contributing to or training ML/AI systems (classification models, labeled data, feedback loops)
\nLab skills: ability to design, build, and execute complex test scenarios that isolate specific failure modes
\nClear written and verbal communication, including customer-facing interaction
\nNice to have
\nExperience with SONiC or other open network operating systems
\nFamiliarity with AI data center designs: low-latency fabrics, GPU cluster networking, NCCL, RDMA/RoCEv2
\nExperience with NVIDIA Spectrum switch family, BlueField SuperNICs, or ConnectX adapters
\nUnderstanding of Ultra Ethernet Consortium specifications and goals
\nPCIe architecture knowledge (relevant to NIC and accelerator integration)
\nExperience with Ixia/Keysight or similar network test equipment and test suites
\nExposure to high-frequency telemetry (HFT) or WJH (What Just Happened) event data
\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Upscale is building the engineering team that drives product improvements though hands-on field engineering. The Forward-Deployed Engineer (FDE) owns complex fabric and system issues end-to-end: from first report through root-cause analysis, workaround delivery, and verified fix. You resolve them with the rest of the team, or you drive them to resolution across whatever boundary stands in the way.
\nThis role is new to UpscaleAI and sits at the intersection of support, engineering, and product management. You will work directly with customers running production AI fabrics, reproduce issues in the lab, develop workarounds under pressure, and contribute fixes and diagnostic content back into the product and our AI-driven monitoring agent.
\nThe environment is often unstructured. Problem definitions are incomplete. Documentation may not exist yet. The right candidate sees that as an opportunity, not an obstacle.
\nLarge AI infrastructure operators increasingly expect their vendor's engineering team to function as an extension of their own infrastructure organization. This role is built to meet that expectation.
\nThis is a small team. There will be an on-call component, but we are building a global team to reduce out-of-hours calls. We work with customers directly, so some travel is involved. Remote meeting tools handle much of the collaboration but building strong trust relationships will require time on site.
\nWhere you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"India - Bangalore","team":"AI Network Software","allLocations":["India - Bangalore"]},"createdAt":1783950182311,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n\n\nWe are seeking a highly skilled gNMI Engineer with strong expertise in SONiC (Software for Open Networking in the Cloud), network programmability, telemetry, and automation. The ideal candidate will be responsible for designing, developing, integrating, and maintaining gNMI-based management and telemetry solutions for SONiC-enabled network devices. The role requires deep understanding of YANG models, OpenConfig, gRPC/gNMI protocols, and modern datacentre networking technologies.\n \nThe engineer will work closely with SONiC developers, network architects, automation teams, and cloud infrastructure teams to build scalable network management and observability solutions.\n \nWhat will you do ?\n• Design and implement gNMI-based configuration and telemetry solutions for SONiC switches. Maintain gNMI clients, servers, and related services.\n• Integrate SONiC management interfaces with OpenConfig and vendor-specific YANG models.\n• Enhance manageability solutions using gNMI Subscribe, Get, Set, and Capabilities operations. Implement streaming telemetry solutions using: gNMI and OpenConfig\n• Support integration of SONiC with data center orchestration platforms.\n• Troubleshoot and optimize gRPC/gNMI communication for large-scale deployments.\n• Contribute to SONiC management framework enhancements.\n• Develop automation frameworks using: Python, Go (Golang)\n• Debug protocol-level issues involving: gRPC, gNMI, OpenConfig, YANG, perform packet captures and protocol analysis.\n• Resolve scalability and performance issues in telemetry systems. Conduct root cause analysis and provide long-term solutions.\n• Contribute to SONiC Community.\n \nWho you are?\n \nMust Have Skills:\n• Engineering background with 15+ years of software development experience\n• Strong experience with: SONiC, gNMI, gRPC, OpenConfig, Yang Data Models\n• Hands-on experience in: Python, Go (Golang)\n• Solid understanding of: Linux system commands, Container technologies (Docker)\n• Experience with: NETCONF, gNOI, OpenTelemetry\n• Deep understanding of Networking, Switching platforms, TCP/IP, BGP, EVPN-VXLAN, Data Centre Networking\n• Strong analytical and problem-solving skills, excellent debugging and troubleshooting abilities.\n• Ability to work in Agile development environments.\n• Self-driven with ownership mindset. Strong communication and cross-functional collaboration skills.\n \nGood to have\n• Contributions to SONiC open-source community.\n• Familiarity with SONiC Management Framework.\n• Experience with hyperscale datacenter environments.\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"India - Bangalore","team":"ASIC AI","allLocations":["India - Bangalore"]},"createdAt":1784014118218,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n\n\nWe are looking for a highly technical, hands-on Principal DV Engineer to anchor the verification of our next-generation, high-performance networking silicon. In this role, you will be the primary technical driver for complex IPs and critical sub-systems. You will balance advanced architectural environment creation with rigorous execution, leading by example on the engineering floor. This position is ideal for a veteran engineer who loves breaking complex designs, building elite UVM environments, and mentoring the next generation of verification talent.\nKey Responsibilities\nIP & Sub-system Leadership: Own the complete verification lifecycle for critical high-performance networking IPs and large-scale sub-systems.\nUVM Environment Architecture: Architect, develop, and maintain advanced, scalable constrained-random verification environments from scratch using System Verilog and UVM.\nTest Planning & Execution: Author comprehensive, bulletproof verification plans. Drive test case development, regressions, triage, and coverage closure (functional, code, and assertion-based metrics).\nAdvanced Debugging: Serve as the team's go-to expert for isolating and resolving the most complex, deep-seated hardware bugs.\nLead by Example: Actively participate in design and verification code reviews. Maintain high standards of code quality, reuse, and documentation within the execution team.\nCross-Functional Alignment: Partner closely with Design and Architecture teams to resolve ambiguities early in the design cycle and ensure high-quality delivery.\nRequired Qualifications & Skills\nExperience: 15+ years of solid experience in ASIC/SoC Design Verification with a proven track record of successful tape-outs.\nDomain Expertise: Extensive experience in the networking domain, verifying high-throughput, data-intensive designs.\nMethodology Expert: Mastery of System Verilog and UVM. Must have a proven history of writing complex testbenches and VIP integrations from scratch.\nVerification Scope: Proven experience leading verification at the IP and Sub-system level, with strong exposure to SoC integration.\nEducation: B Tech/M Tech in Electronics/Electrical Engineering, Computer Engineering, or a related field.\nPreferred / Added Advantages\nHands-on experience with Formal Property Verification (FPV) to exhaustively verify block-level control logic.\nProficiency in Python or Perl for testbench automation and triage scripting.\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"India - Bangalore","team":"ASIC Engineering","allLocations":["India - Bangalore"]},"createdAt":1783616029729,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"India - Bangalore","team":"AI Network Software","allLocations":["India - Bangalore"]},"createdAt":1784698633730,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n\n\nPrincipal Engineer: SDK, SAI & Platform Debugging\n \nWe are looking for a seasoned networking software professional with 12+ years of industry experience in switch software, ASIC SDKs, and network operating systems. This role is intended for an experienced engineer who can independently lead complex technical investigations, drive cross-functional problem resolution, and act as the go-to expert for SDK, SAI, and platform-related issues.\n \nThe ideal candidate should have extensive experience debugging across the entire networking software stack, including SAI, SDK, NOS, Linux kernel, platform drivers, and hardware interfaces. Beyond troubleshooting, the candidate is expected to influence architecture, guide engineering teams, engage directly with silicon vendors, and establish best practices for platform stability and supportability.\n \nWhat will you do ?\nProvide guidance on SDK-SAI integration and platform software design. Drive technical reviews and influence software design decisions to improve debuggability, reliability and maintainability of the Network Operating System (NOS).\nAnalyze SDK-SAI interactions and identify functional, performance, or scalability bottlenecks.\nServe as a lead for SDK- and SAI-related escalations across product development and customer deployments. Lead investigation and resolution of complex issues involving ASIC SDKs, SAI implementations, and switching platforms.\nOwn end-to-end root cause analysis of critical platform, feature, performance, and scale issues. Coordinate across development, QA, platform, silicon vendors, and support teams to drive issue resolution. Work closely with ASIC vendors while debugging SDK defects, feature gaps, and interoperability issues. \nAnalyze crashes, memory leaks, packet drops, forwarding inconsistencies, and hardware-software interaction issues. Debug issues spanning SDK, SAI, Platform software, Linux kernel, Hardware abstraction layers, Control and data plane components\nMentor and coach engineers on troubleshooting methodologies, debugging techniques, and networking fundamentals.\n \nWho you are?\nBachelor's or master's degree in computer science, Electronics, Telecommunications.\n12–18 years of experience in networking, software development, SDK Debugging & Troubleshooting. \nProven track record of leading complex SDK debugging and issue resolution efforts. Strong Linux system debugging and networking stack knowledge.\nStrong hands-on expertise with at least one networking switch SDKs viz. Broadcom, Marvell, Cisco Silicon One, or similar merchant silicon SDKs\nDeep understanding of SAI architecture, APIs, and SDK-SAI Interface and models.\nGood understanding of Switching ASICs, Linux Networking, Linux Networking, Packet Flow Analysis\nStrong expertise in:\nEthernet Switching\nL2/L3 Forwarding\nVLAN, STP, LACP, ACLs, QOS, ECMP, ACLs, QOS, ECMP, BGP, VXLAN/EVPN and Routing protocols\nProficiency in C/C++ and Python.\nDesirable:\nDirect engagement with silicon vendors for issue triage and feature enablement.\nExposure to hyperscale, cloud, or datacenter networking environments.\nExposure on SONiC, Platform Bring-up, NVIDIA Spectrum Series SDK Diagnostics\nAutomation & Test Frameworks\n \n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
What will you do ?
\nWho you are?
\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"What will you do ?
\nWho you are?
\nWhere you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"US - Headquarters","team":"ASIC Architecture","allLocations":["US - Headquarters"]},"createdAt":1781637117940,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n\n\nPosition Overview\n\n\nAbout the Role: We're building the simulation platform for UpscaleAI's next-generation AI networking ASICs. The platform is a functionally accurate software model of our ASIC architecture, enabling software development, verification, and customer validation before silicon is available. You'll be responsible for developing, testing and maintaining the platform across multiple ASIC generations, ensuring it accurately represents hardware behavior and scales to support our customers' demanding AI infrastructure workloads. \n\n\nWhat You'll Work On: \n\nASIC Architectural Model: Our current-generation ASIC simulator\nNext-gen ASIC models: Future silicon with enhanced AI networking features\nMulti-chip simulation: Modeling systems with multiple interconnected ASICs\nPerformance optimization: Making large-scale simulations practical\n \n\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Position Overview
\n
About the Role: We're building the simulation platform for UpscaleAI's next-generation AI networking ASICs. The platform is a functionally accurate software model of our ASIC architecture, enabling software development, verification, and customer validation before silicon is available. You'll be responsible for developing, testing and maintaining the platform across multiple ASIC generations, ensuring it accurately represents hardware behavior and scales to support our customers' demanding AI infrastructure workloads.
What You'll Work On:
ASIC Architectural Model: Our current-generation ASIC simulator
\nNext-gen ASIC models: Future silicon with enhanced AI networking features
\nMulti-chip simulation: Modeling systems with multiple interconnected ASICs
\nPerformance optimization: Making large-scale simulations practical
\n\n
Design and implement software model components that accurately simulate ASIC behavior (forwarding engines, schedulers, memory subsystems, control plane interfaces)
\nTranslate architectural specifications and microarchitecture documents into high-fidelity C++ models, ranging from functionally accurate to performance-accurate depending on the subsystem
\nDevelop and maintain the register model, memory map, and configuration interfaces
\nImplement packet processing pipelines for multiple forwarding modes
\nOptimize model performance for large-scale simulations
\nCollaborate with RTL engineers to ensure model-to-silicon correlation
\nWork with the architecture team to validate algorithm and architectural choices through simulation
\nWork with SDK and platform software teams to enable software development and validation on the model
\nSupport multiple ASIC versions with maintainable, configurable model architecture
\n
Requirements
7+ years of experience in C/C++ architecture model development, cycle-aware simulation, or ASIC verification
\nExpert-level C++ programming (modern C++17)
\nDeep understanding of network switch/router architecture: forwarding tables, scheduling, QoS, ACLs
\nExperience with packet processing pipelines and dataplane simulation
\nFamiliarity with hardware description languages (Verilog/SystemVerilog) and reading RTL
\nExperience with gRPC, Protocol Buffers, or similar IPC mechanisms
\nExperience with C++ unit testing frameworks (e.g., Google Test)
\nFamiliarity with Python for test automation, tooling, and scripting
\nStrong debugging skills across software and hardware domains
\n
Nice to Have
Experience with AI/ML networking (scale-up fabrics, RDMA, collective operations)
\nExperience using ns-3 and/or HTSim to model high performance network architectures.
\nFamiliarity with SAI (Switch Abstraction Interface) or SONiC
\nExperience with FPGA prototyping or emulation platforms
\nBackground in memory subsystem modeling (DDR, HBM, caches)
\nExperience with configuration-driven simulation frameworks (YAML, JSON, protobuf-based config)
\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Position Overview
\n
About the Role: We're building the simulation platform for UpscaleAI's next-generation AI networking ASICs. The platform is a functionally accurate software model of our ASIC architecture, enabling software development, verification, and customer validation before silicon is available. You'll be responsible for developing, testing and maintaining the platform across multiple ASIC generations, ensuring it accurately represents hardware behavior and scales to support our customers' demanding AI infrastructure workloads.
What You'll Work On:
ASIC Architectural Model: Our current-generation ASIC simulator
\nNext-gen ASIC models: Future silicon with enhanced AI networking features
\nMulti-chip simulation: Modeling systems with multiple interconnected ASICs
\nPerformance optimization: Making large-scale simulations practical
\n\n
Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"India - Bangalore","team":"AI Network Software","allLocations":["India - Bangalore"]},"createdAt":1783950260994,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n\n\nWe are seeking a highly skilled gNMI Engineer with strong expertise in SONiC (Software for Open Networking in the Cloud), network programmability, telemetry, and automation. The ideal candidate will be responsible for developing, integrating, and maintaining gNMI-based management and telemetry solutions for SONiC-enabled network devices. The role requires good understanding of YANG models, OpenConfig, gRPC/gNMI protocols, and modern datacentre networking technologies.\n \nThe engineer will work closely with SONiC developers, automation teams, and cloud infrastructure teams to build scalable network management and observability solutions.\n \nWhat will you do ?\n• Design and implement gNMI-based configuration and telemetry solutions for SONiC switches. Maintain gNMI clients, servers, and related services.\n• Integrate SONiC management interfaces with OpenConfig and vendor-specific YANG models.\n• Enhance manageability solutions using gNMI Subscribe, Get, Set, and Capabilities operations. Implement streaming telemetry solutions using: gNMI and OpenConfig\n• Support integration of SONiC with data center orchestration platforms.\n• Troubleshoot and optimize gRPC/gNMI communication for large-scale deployments.\n• Contribute to SONiC management framework enhancements.\n• Develop automation frameworks using: Python, Go (Golang)\n• Debug protocol-level issues involving: gRPC, gNMI, OpenConfig, YANG, perform packet captures and protocol analysis.\n• Resolve scalability and performance issues in telemetry systems. Conduct root cause analysis and provide long-term solutions.\n• Contribute to SONiC Community.\n \nWho you are?\n \nMust Have Skills:\n• Engineering background with 5+ years of software development experience\n• Strong experience with: SONiC, gNMI, gRPC, OpenConfig, Yang Data Models\n• Hands-on experience in: Python, Go (Golang)\n• Solid understanding of: Container technologies (Docker)\n• Experience with: NETCONF, gNOI, OpenTelemetry\n• Basic understanding of Networking, Switching platforms, TCP/IP, BGP, EVPN-VXLAN, Data Centre Networking\n• Strong analytical and problem-solving skills, excellent debugging and troubleshooting abilities.\n• Ability to work in Agile development environments.\n \nGood to have\n• Contributions to SONiC open-source community.\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"India - Bangalore","team":"ASIC Engineering","allLocations":["India - Bangalore"]},"createdAt":1783616132290,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Role Summary:
As a Senior Back End Manager – Physical Design, you will be responsible for leading high-performance ASIC/SoC physical implementation teams delivering complex designs from RTL to GDSII. You will drive execution strategy, build scalable methodologies, manage technical execution across multiple blocks/subsystems, and ensure successful design convergence across advanced technology nodes.
The role requires a strong technical leader with hands-on expertise in physical design, proven ability to manage engineering teams, and experience delivering large-scale silicon programs with aggressive Power, Performance, Area, and Schedule targets.
\nKey Responsibilities:
\nLead and manage a team of Physical Design engineers responsible for implementation and signoff of complex ASIC/SoC designs.
\nOwn overall back-end execution strategy including floorplanning, placement, CTS, routing, timing closure, physical verification, and tape-out readiness.
\nDrive project planning, resource allocation, execution tracking, risk identification, and mitigation across multiple design milestones.
\nProvide technical leadership for architecture decisions including hierarchy definition, partitioning strategy, clock architecture, power planning, and implementation methodology.
\nCollaborate closely with Architecture, RTL, DFT, STA, CAD, Packaging, Foundry, and Program Management teams to achieve successful silicon delivery.
\nDrive convergence across PPA, timing, power, IR/EM, signal integrity, DRC/LVS, reliability, and manufacturability requirements.
\nEstablish robust execution methodologies, quality checks, automation flows, and best practices across the Back End organization.
\nReview design quality at key milestones and ensure strong execution discipline through structured reviews and predictable delivery.
\nLead debug and resolution of complex design challenges across timing, congestion, power, physical verification, and ECO implementation.
\nPartner with CAD and EDA vendors to improve flows, evaluate new technologies, and enhance design productivity.
\nMentor engineers, technical leads, and future leaders while building a culture of ownership, innovation, accountability, and engineering excellence.
\nRequired Skills & Experience:
\n15–20 years of experience in ASIC/SoC Physical Design with strong experience managing large Back End implementation teams.
\nProven track record of delivering complex SoC designs successfully through tape-out on advanced technology nodes.
\nDeep expertise in physical implementation flows including floorplan, power planning, placement, CTS, routing, ECO, and signoff closure.
\nStrong experience with Cadence Innovus based implementation flows.
\nExpert-level understanding of STA (PrimeTime), IR/EM analysis (Voltus/RedHawk), physical verification (Calibre/ICV), and extraction methodologies.
\nHands-on experience with advanced process nodes (5 nm / 3 nm / N3E/N3P) and foundry signoff requirements.
\nStrong understanding of hierarchical design methodologies, multi-voltage designs, UPF, clock distribution strategies, and MCMM closure.
\nExperience defining metrics, dashboards, quality checks, and execution processes for predictable project delivery.
\nStrong scripting and automation mindset using Tcl, Perl, Python, or equivalent languages.
\nExcellent leadership, communication, decision-making, and cross-functional collaboration skills.
\nPreferred Skills:
\nExperience leading full-chip implementation and large subsystem ownership.
\nExposure to AI/HPC/networking class SoCs with aggressive performance and power requirements.
\nExperience building teams, defining organizational processes, and scaling execution methodologies.
\nKnowledge of ML-driven EDA optimization and next-generation design methodologies is a plus.
\nMaster’s Degree in Electrical Engineering / Electronics / VLSI Engineering preferred.
\nKey Traits for Success:
\nStrong technical depth combined with execution leadership.
\nOwnership mindset with the ability to drive complex programs to closure.
\nAbility to identify risks early and create structured recovery plans.
\nPassion for building high-performing teams and mentoring engineers.
\nData-driven decision making with focus on quality, predictability, and continuous improvement.
\nAbility to balance innovation, schedule pressure, and silicon delivery excellence.
\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Role Summary:
As a Senior Back End Manager – Physical Design, you will be responsible for leading high-performance ASIC/SoC physical implementation teams delivering complex designs from RTL to GDSII. You will drive execution strategy, build scalable methodologies, manage technical execution across multiple blocks/subsystems, and ensure successful design convergence across advanced technology nodes.
The role requires a strong technical leader with hands-on expertise in physical design, proven ability to manage engineering teams, and experience delivering large-scale silicon programs with aggressive Power, Performance, Area, and Schedule targets.
\nKey Responsibilities:
\nLead and manage a team of Physical Design engineers responsible for implementation and signoff of complex ASIC/SoC designs.
\nOwn overall back-end execution strategy including floorplanning, placement, CTS, routing, timing closure, physical verification, and tape-out readiness.
\nDrive project planning, resource allocation, execution tracking, risk identification, and mitigation across multiple design milestones.
\nProvide technical leadership for architecture decisions including hierarchy definition, partitioning strategy, clock architecture, power planning, and implementation methodology.
\nCollaborate closely with Architecture, RTL, DFT, STA, CAD, Packaging, Foundry, and Program Management teams to achieve successful silicon delivery.
\nDrive convergence across PPA, timing, power, IR/EM, signal integrity, DRC/LVS, reliability, and manufacturability requirements.
\nEstablish robust execution methodologies, quality checks, automation flows, and best practices across the Back End organization.
\nReview design quality at key milestones and ensure strong execution discipline through structured reviews and predictable delivery.
\nLead debug and resolution of complex design challenges across timing, congestion, power, physical verification, and ECO implementation.
\nPartner with CAD and EDA vendors to improve flows, evaluate new technologies, and enhance design productivity.
\nMentor engineers, technical leads, and future leaders while building a culture of ownership, innovation, accountability, and engineering excellence.
\nRequired Skills & Experience:
\n15–20 years of experience in ASIC/SoC Physical Design with strong experience managing large Back End implementation teams.
\nProven track record of delivering complex SoC designs successfully through tape-out on advanced technology nodes.
\nDeep expertise in physical implementation flows including floorplan, power planning, placement, CTS, routing, ECO, and signoff closure.
\nStrong experience with Cadence Innovus based implementation flows.
\nExpert-level understanding of STA (PrimeTime), IR/EM analysis (Voltus/RedHawk), physical verification (Calibre/ICV), and extraction methodologies.
\nHands-on experience with advanced process nodes (5 nm / 3 nm / N3E/N3P) and foundry signoff requirements.
\nStrong understanding of hierarchical design methodologies, multi-voltage designs, UPF, clock distribution strategies, and MCMM closure.
\nExperience defining metrics, dashboards, quality checks, and execution processes for predictable project delivery.
\nStrong scripting and automation mindset using Tcl, Perl, Python, or equivalent languages.
\nExcellent leadership, communication, decision-making, and cross-functional collaboration skills.
\nPreferred Skills:
\nExperience leading full-chip implementation and large subsystem ownership.
\nExposure to AI/HPC/networking class SoCs with aggressive performance and power requirements.
\nExperience building teams, defining organizational processes, and scaling execution methodologies.
\nKnowledge of ML-driven EDA optimization and next-generation design methodologies is a plus.
\nMaster’s Degree in Electrical Engineering / Electronics / VLSI Engineering preferred.
\nKey Traits for Success:
\nStrong technical depth combined with execution leadership.
\nOwnership mindset with the ability to drive complex programs to closure.
\nAbility to identify risks early and create structured recovery plans.
\nPassion for building high-performing teams and mentoring engineers.
\nData-driven decision making with focus on quality, predictability, and continuous improvement.
\nAbility to balance innovation, schedule pressure, and silicon delivery excellence.
\nWhere you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"US - Headquarters","team":"Orchestration","allLocations":["US - Headquarters"]},"createdAt":1784013735344,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n\n\nAbout the role\nOwn the reliability, deployment, and operational infrastructure behind Orchestrator and the AI Fabric environments it manages.\nYou will build and maintain Kubernetes clusters across on-prem and cloud, design CI/CD pipelines for continuous delivery, manage Terraform-driven infrastructure-as-code, and handle secret and certificate rotations.\nYou will stand up and operate the full observability stack — Prometheus, Grafana, Loki, Splunk, Datadog, and Timestream — ensuring end-to-end visibility across customer deployments. When things break, you are the person who troubleshoots and debugs infrastructure issues at the platform and customer site level.\nBeyond keeping things running, you will build internal tooling and analytics that improve operational efficiency, reduce incident response time, and scale our infrastructure as deployments grow. We are looking for someone who has been through production pain, knows what good SRE looks like, and can bring that discipline to a fast-moving team.\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
About the role
\nOwn the reliability, deployment, and operational infrastructure behind Orchestrator and the AI Fabric environments it manages.
\nYou will build and maintain Kubernetes clusters across on-prem and cloud, design CI/CD pipelines for continuous delivery, manage Terraform-driven infrastructure-as-code, and handle secret and certificate rotations.
\nYou will stand up and operate the full observability stack — Prometheus, Grafana, Loki, Splunk, Datadog, and Timestream — ensuring end-to-end visibility across customer deployments. When things break, you are the person who troubleshoots and debugs infrastructure issues at the platform and customer site level.
\nBeyond keeping things running, you will build internal tooling and analytics that improve operational efficiency, reduce incident response time, and scale our infrastructure as deployments grow. We are looking for someone who has been through production pain, knows what good SRE looks like, and can bring that discipline to a fast-moving team.
\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"About the role
\nOwn the reliability, deployment, and operational infrastructure behind Orchestrator and the AI Fabric environments it manages.
\nYou will build and maintain Kubernetes clusters across on-prem and cloud, design CI/CD pipelines for continuous delivery, manage Terraform-driven infrastructure-as-code, and handle secret and certificate rotations.
\nYou will stand up and operate the full observability stack — Prometheus, Grafana, Loki, Splunk, Datadog, and Timestream — ensuring end-to-end visibility across customer deployments. When things break, you are the person who troubleshoots and debugs infrastructure issues at the platform and customer site level.
\nBeyond keeping things running, you will build internal tooling and analytics that improve operational efficiency, reduce incident response time, and scale our infrastructure as deployments grow. We are looking for someone who has been through production pain, knows what good SRE looks like, and can bring that discipline to a fast-moving team.
\nWhere you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"India - Bangalore","team":"ASIC AI","allLocations":["India - Bangalore"]},"createdAt":1784014406752,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n\n\nExperience Level: 17+ Years (or 15+ with exceptional strategic track record)\nLocation: Bangalore, India (On-Site)\nPosition Type: Full-Time\nAbout the Role\nWe are seeking a visionary Senior Principal DV Engineer to define the global verification strategy for our massive-scale networking SoCs. This is a high-impact, elite technical leadership role. You will be responsible for defining the overarching SoC verification methodology, driving cross-functional alignment from architecture to post-silicon, and pioneering next-generation verification technologies (such as Formal and Emulation). You will influence design architectures for verifiability and serve as a strategic technical advisor to executive leadership while steering the broader engineering organization.\nKey Responsibilities\nSoC Verification Strategy: Define, architect, and govern the overarching verification roadmap and methodology for full-scale, multi-million gate networking SoCs.\nMethodology Innovation: Drive the adoption of advanced methodologies. Seamlessly integrate Formal Property Verification (FPV), hardware emulation, and advanced shift-left strategies into the standard SoC timeline.\nHigh-Level Governance: Establish, track, and enforce rigorous quality metrics and sign-off criteria across the entire project lifecycle, ensuring zero-defect silicon.\nArchitectural Influence: Work directly with Silicon Architects and Design Directors to review specs, optimize micro-architecture for testability/verifiability (DFV), and mitigate risk early.\nOrganizational Mentorship: Mentor Principal engineers and technical leads across the site. Champion best practices, drive technical brainstorming, and cultivate a culture of engineering excellence.\nCrisis Debugging & Escalation: Act as the ultimate technical authority for critical, complex issues that block project milestones or require deep cross-functional triage.\nRequired Qualifications & Skills\nExperience: 15+ to 17+ years of distinguished experience in ASIC/SoC Verification, having successfully guided multiple complex, large-scale tape-outs from concept to production.\nDomain Leadership: Deep, comprehensive knowledge of the networking domain, with an understanding of how system-level traffic flow impacts full-chip verification constraints.\nFull-Scale SoC Scope: Proven experience leading full SoC-level verification environments, managing top-level testbenches, and coordinating multi-subsystem integration.\nExpertise: Authoritative knowledge of System Verilog, UVM, and structural verification planning.\nInfluence & Communication: Outstanding executive communication skills. Ability to articulate complex technical risks and trade-offs clearly to both engineering teams and executive management.\nEducation: B Tech/M Tech in Electronics/Electrical Engineering, Computer Engineering, or a related field.\nPreferred / Added Advantages\nStrong, proven track record of deploying Formal Verification (FPV) at scale to displace traditional simulation.\nExperience integrating hardware emulation (e.g., Palladium, Zebu) into a unified simulation workflow.\n \n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"US - Headquarters","team":"Orchestration","allLocations":["US - Headquarters"]},"createdAt":1784743178028,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
The role
\nYou will define the cross-cutting architecture that the engineering team builds within. You own platform-scale initiatives from vision through delivery — authoring architecture decision documents, building POCs to validate direction, and partnering with Product to shape the roadmap. You write code, you ship, and you are accountable for the quality of what reaches stakeholders.
\nHow we expect you to work
\n\n
Nice to have
\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"The role
\nYou will define the cross-cutting architecture that the engineering team builds within. You own platform-scale initiatives from vision through delivery — authoring architecture decision documents, building POCs to validate direction, and partnering with Product to shape the roadmap. You write code, you ship, and you are accountable for the quality of what reaches stakeholders.
\nWhere you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"India - Bangalore","team":"Orchestration","allLocations":["India - Bangalore"]},"createdAt":1784026021350,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n\n\nLocation: India · Experience: 12–15 years · Type: Full-time, IC\nAbout us\nWe build a distributed control plane for datacenter network fabrics. Our platform orchestrates intent-driven configuration, change management, drift remediation, and full-stack observability across cloud-hosted services and on-prem edge appliances deployed in customer datacenters worldwide.\nThe role\nYou will own critical platform subsystems end-to-end — from design through production. You'll build and ship core distributed systems components, drive the technical quality of the codebase, and unblock the team through direct hands-on contribution. You work closely with Product, QA, and Customer Engineering to deliver quality product to stakeholders.\nWhat you'll work on\nDistributed control plane spanning cloud services and on-prem edge appliances connected via mTLS gRPC streams\nObservability and telemetry at scale: OpenTelemetry collection pipelines, stream processing (Kafka), time-series storage (ClickHouse/VictoriaMetrics), real-time fabric state views, packet-event analysis, and fleet-wide health aggregation\nAgentic AI operations: design and build autonomous infrastructure agents that evaluate prerequisites, orchestrate multi-step workflows (onboarding, upgrades, drift remediation), handle failure recovery, and interact with the control plane through tool-use patterns (LangGraph, MCP)\nFleet orchestration: aggregate health/compliance/drift APIs, cross-site campaign execution, template promotion workflows, parallel site onboarding\nData architecture across Postgres, ArangoDB, ClickHouse, Redis, Kafka, and Git-backed content stores\nMulti-tenant SaaS with site-scoped RBAC, session lifecycle, and enterprise IdP integration\nDevice lifecycle management: zero-touch provisioning, enrollment protocols, config push via edge relay, drift detection and remediation\nIntent compilation engine: workspace management, merge request lifecycle, change request execution with gate-based verification and rollback\nWhat you bring\n12–15 years building and operating distributed systems serving enterprise customers across cloud and on-prem environments\nDeep proficiency in Go (or comparable systems language) with strong distributed systems fundamentals\nSignificant experience designing observability and telemetry platforms: collection agents, stream processing, time-series databases, alerting pipelines, and real-time dashboards at scale\nProduction experience with microservices architecture: gRPC, Protocol Buffers, spec-first REST APIs (OpenAPI)\nHands-on with multiple storage paradigms: relational, graph, time-series, key-value, and streaming\nTrack record building multi-tenant platforms with tenant isolation, RBAC, and identity federation\nExperience with Kubernetes, Helm, and hybrid cloud/on-prem deployment models\nStrong API design sense: versioning, backward compatibility, contract-first development\nAbility to communicate architectural decisions clearly through writing and diagrams\nHow we expect you to work\nCross-functional by default — you work closely with Product, QA, Design, and Customer Engineering, not in isolation\nSolution-oriented — when the team is stuck, you unblock them by driving toward answers and building what's needed\nAccountable for delivery — you take personal responsibility for shipping quality product to stakeholders on time\nHands-on always — you write code daily and prove ideas by building them\nNice to have\nAI/ML agent architectures for infrastructure operations: LangGraph, AutoGen, MCP tool-use, human-in-the-loop gating, and autonomous workflow orchestration\nNetwork automation or infrastructure management platforms (Apstra, NSO, Terraform, Crossplane)\nDatacenter networking experience: OpenConfig, gNMI, or fabric management at scale\nHub-and-spoke / edge computing / control-plane-data-plane separation architectures\nOpenTelemetry contributor experience or deep familiarity with the collector ecosystem\nDevice enrollment or zero-touch provisioning systems\nOpen-source contributions or published work in distributed systems\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Location: India · Experience: 12–15 years · Type: Full-time, IC
\nAbout us
\nWe build a distributed control plane for datacenter network fabrics. Our platform orchestrates intent-driven configuration, change management, drift remediation, and full-stack observability across cloud-hosted services and on-prem edge appliances deployed in customer datacenters worldwide.
\nThe role
\nYou will own critical platform subsystems end-to-end — from design through production. You'll build and ship core distributed systems components, drive the technical quality of the codebase, and unblock the team through direct hands-on contribution. You work closely with Product, QA, and Customer Engineering to deliver quality product to stakeholders.
\nWhat you'll work on
\nWhat you bring
\nHow we expect you to work
\nNice to have
\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Location: India · Experience: 12–15 years · Type: Full-time, IC
\nAbout us
\nWe build a distributed control plane for datacenter network fabrics. Our platform orchestrates intent-driven configuration, change management, drift remediation, and full-stack observability across cloud-hosted services and on-prem edge appliances deployed in customer datacenters worldwide.
\nThe role
\nYou will own critical platform subsystems end-to-end — from design through production. You'll build and ship core distributed systems components, drive the technical quality of the codebase, and unblock the team through direct hands-on contribution. You work closely with Product, QA, and Customer Engineering to deliver quality product to stakeholders.
\nWhat you'll work on
\nWhat you bring
\nHow we expect you to work
\nNice to have
\nWhere you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"US - Headquarters","team":"Hardware Engineering","allLocations":["US - Headquarters"]},"createdAt":1782093479396,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Support SI/PI analysis for high-speed PCB designs from architecture through product validation.
Assist with channel modeling, simulation, and optimization for high-speed interfaces.
Perform signal quality measurements using laboratory equipment.
Support PCB stackup reviews, routing reviews, and design guideline development.
Analyze transmission lines, vias, connectors, cables, and package-to-board interfaces.
Participate in correlation between simulation and laboratory measurements.
Work closely with PCB designers, hardware engineers, package engineers, and silicon teams.
Document simulation results, measurements, and design recommendations.
Assist in root-cause analysis and debugging of signal integrity issues.
Preferred Qualifications
BS or MS in Electrical Engineering or related field.
0–3 years of experience in Signal Integrity, PCB Design, Hardware Design, or related areas.
Basic understanding of:
- High-speed digital design
- Transmission line theory
- PCB stackups and routing
- S-parameters and eye diagrams
- SI and PI fundamentals
Exposure to one or more simulation tools such as:
Cadence Clarity
Cadence Sigrity
Keysight ADS
Ansys HFSS
Exposure to laboratory instruments such as:
Vector Network Analyzer (VNA)
TDR/TDT
High-speed Oscilloscopes
Strong analytical and problem-solving skills.
Excellent communication and teamwork abilities.
Nice to Have
Experience with 56G/112G/224G SerDes systems.
Knowledge of PCIe, Ethernet, UCIe, DDR interfaces
Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"US - Headquarters","team":"Hardware Engineering","allLocations":["US - Headquarters"]},"createdAt":1782096130051,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
- Lead Signal Integrity and Power Integrity analysis for complex PCB and system designs.
- Define SI/PI architecture, design rules, and channel budgets for high-speed interfaces.
- Perform detailed channel simulations and optimization using industry-standard tools.
- Drive correlation between simulation and laboratory measurements.
- Develop and validate models for packages, PCBs, connectors, cables, and complete channels.
- Conduct pre-layout and post-layout SI/PI analysis.
- Perform compliance and interoperability analysis for industry standards.
- Debug complex signal integrity issues and drive root-cause resolution.
- Define PCB stackups, routing requirements, via structures, and return path strategies.
- Mentor junior engineers and provide technical leadership across projects.
- Collaborate with silicon, package, board, mechanical, and manufacturing teams.
- Interface with vendors and customers on SI/PI-related technical topics.
Required Qualifications:
- BS/MS in Electrical Engineering or related field.
- 5+ years of experience in Signal Integrity, High-Speed Design, or related areas.
Strong understanding of:
- Signal Integrity and Power Integrity principles
- Electromagnetic theory
- Transmission line analysis
- High-speed SerDes architectures
- PCB and package design
- Channel compliance methodologies
Extensive hands-on experience with:
- Ansys HFSS
- Cadence Clarity
- Cadence Sigrity
- Keysight ADS
Strong laboratory experience with:
- Vector Network Analyzer (VNA)
- TDR/TDT systems
- High-bandwidth oscilloscopes
- De-embedding and calibration techniques
Experience analyzing:
PCB channels, Connectors, Cables, Packages, Backplanes and midplanes
Ability to correlate simulations with measured results and drive design improvements.
Preferred Qualifications:
- Experience with 112G/224G PAM4 SerDes systems.
- Experience with Ethernet, PCIe, UCIe, CXL, DDR, or AI networking products.
- Knowledge of COM, ERL, eye analysis, jitter decomposition, and channel compliance methodologies.
- Experience with Python, MATLAB, or automation frameworks.
- Experience leading SI strategy for large-scale hardware programs.
Desired Attributes:
- Strong technical leadership and mentoring skills.
- Excellent cross-functional communication.
- Ability to drive projects independently from concept through production.
- Customer-facing and vendor collaboration experience.
- Passion for solving challenging high-speed design pro
Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Where you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"US - Headquarters","team":"Support","allLocations":["US - Headquarters"]},"createdAt":1784687061525,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Role Overview
\nWe are looking for a Technical Support Lead Engineer with 12+ years of experience who is passionate about building reliable networking switches and solving challenging infrastructure problems.
\nOwn end-to-end network bring-up and performance for GPU AI clusters during customer PoCs. You’ll be the leader for switching ASIC pipeline tuning, platform readiness, and fabric performance (Ethernet/RoCEv2 and/or InfiniBand), turning requirements into reproducible configurations that hit utilization and tail-latency targets at scale.
\nAs part of the AI Networking team, you'll debug software, develop automation, validate networking platforms, troubleshoot complex system issues, and work closely with engineering teams and customers to ensure the best customer experience. You'll have the opportunity to work on technologies powering next-generation GPU clusters and AI data centers while learning from experienced engineers in a fast-paced startup environment.
\nDesign, develop, and maintain features and enhancements for the SONiC NOS platform.
\n· Debug, troubleshoot, and resolve issues on SONiC platforms.
\n· Develop and execute debugging and troubleshoot infrastructure.
\n· Collaborate closely with cross-functional teams including hardware engineers and Test teams.
\n· Participate in code reviews, architecture discussions, and documentation efforts.
\n· Develop support strategies to root-cause Networking ASICs and Networking Systems issues.
\n· Develop debugging tools, for Traffic monitoring and Performance measurements.
\n· Debug issues across software, Linux systems, networking stacks, and distributed infrastructure.
\n· Be a point of contact for customer deployments, integration testing, proof-of-concepts, and field issue resolution.
\n· Build tools that improve deployment efficiency, observability, telemetry, and automated testing.
\n· Collaborate with software, infrastructure, QA, and product teams to identify root causes and deliver robust solutions.
\n· Contribute to backend services, APIs, orchestration components, and infrastructure automation.
\n· Document technical findings and communicate effectively with engineering teams and customers.
\nRequired Qualifications
\n· Bachelor’s or master’s degree in computer science, Electrical Engineering, or a related field.
\n· Minimum of 12 years of work experience is required, with at least 3 years of hands-on SONiC or equivalent Network Operating System (NOS) development experience preferred.
\n· Strong programming skills in Python, Go, or a similar language.
\n· Solid understanding of Linux, TCP/IP networking, routing, switching, VLANs, and network troubleshooting.
\n· Experience with PTF (Packet Test Framework) and SPyTest for network validation.
\n· Familiarity with Linux internals, docker containers.
\n· Excellent problem-solving skills and ability to work in a fast-paced, collaborative environment.
\n· Knowledge of network ASICs and switch hardware architecture is mandatory.
\n· Excellent written and verbal communication skills.
\n· Ability to thrive in a collaborative, fast-paced startup environment with a strong sense of ownership.
\n
Preferred Qualifications
· Data-center networking with hands-on switch ASIC tuning and platform bring-up.
\n· Proven RoCEv2 deployments at 200/400/800 G: ECN/PFC design, DCQCN tuning, DSCP/PCP mapping, queue/WRR shaping.
\n· Deep buffer/queueing knowledge (headroom, xon/xoff, dynamic thresholds, VOQ vs shared pools).
\n· SONiC (buffers.json, qos.json, pfcwd, ecn), Enterprise-OS (class-map/type qos & network-qos, policy-map, queuing).
\n· Optics & PHY: PAM4 signal integrity, FEC modes, DOM/RS-FEC counters, AN/LT quirks, DAC/AOC selection.
\n· Tooling: ethtool, devlink, mlnx_qos, perfquery, switch telemetry (INT/sFlow/ERSPAN), gNMI/REST, Prometheus.
\n· Benchmarking: perftest, nccl-tests, iPerf3, traffic generators; reading queue stats, ECN marks, and CNP behavior.
\n· EVPN/VXLAN leaf-spine for AI pods; flowlet or latency-aware hashing.
\n· BlueField DPU offloads, GPUDirect RDMA; NIC QoS (DSCP-to-TC, PFCx, GEARBOX/FW).
\n· Storage fabrics for AI (NFS-RDMA, NVMe-oF) and their QoS interactions.
\n· Python/Ansible for templated ASIC profiles; Git workflows for config promotion.
\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Role Overview
\nWe are looking for a Technical Support Lead Engineer with 12+ years of experience who is passionate about building reliable networking switches and solving challenging infrastructure problems.
\nOwn end-to-end network bring-up and performance for GPU AI clusters during customer PoCs. You’ll be the leader for switching ASIC pipeline tuning, platform readiness, and fabric performance (Ethernet/RoCEv2 and/or InfiniBand), turning requirements into reproducible configurations that hit utilization and tail-latency targets at scale.
\nAs part of the AI Networking team, you'll debug software, develop automation, validate networking platforms, troubleshoot complex system issues, and work closely with engineering teams and customers to ensure the best customer experience. You'll have the opportunity to work on technologies powering next-generation GPU clusters and AI data centers while learning from experienced engineers in a fast-paced startup environment.
\nWhere you fall within that range depends on your experience, skills, and impact—we benchmark against internal levels to keep things fair and consistent.
\nEqual Opportunity
\nUpscale AI is building a team that reflects a wide range of perspectives, backgrounds, and experiences. We’re proud to be an Equal Opportunity Employer and consider all qualified applicants regardless of race, color, religion, national origin, sex, sexual orientation, gender identity, disability, or veteran status.
\nAccessibility & Accommodations
\nWe’re committed to making our hiring process accessible to everyone. If you need accommodations at any stage, just reach out to us at hiring@upscaleai.com—we’re happy to help. Note: This inbox is only for accommodation requests.
","categories":{"commitment":"Reg - Full-Time","location":"US - Headquarters","team":"Support","allLocations":["US - Headquarters"]},"createdAt":1784689337950,"descriptionPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","description":"Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
Role Overview
\nWe are looking for a Technical Support Engineer with 10+ years of experience who is passionate about building reliable networking switches and solving challenging infrastructure problems.
\nOwn end-to-end network bring-up and performance for GPU AI clusters during customer PoCs. You’ll be the leader for switching ASIC pipeline tuning, platform readiness, and fabric performance (Ethernet/RoCEv2 and/or InfiniBand), turning requirements into reproducible configurations that hit utilization and tail-latency targets at scale.
\nAs part of the AI Networking team, you'll debug software, develop automation, validate networking platforms, troubleshoot complex system issues, and work closely with engineering teams and customers to ensure the best customer experience. You'll have the opportunity to work on technologies powering next-generation GPU clusters and AI data centers while learning from experienced engineers in a fast-paced startup environment.
\n· Design, develop, and maintain features and enhancements for the SONiC NOS platform.
\n· Debug, troubleshoot, and resolve issues on SONiC platforms.
\n· Develop and execute debugging and troubleshoot infrastructure.
\n· Collaborate closely with cross-functional teams including hardware engineers and Test teams.
\n· Participate in code reviews, architecture discussions, and documentation efforts.
\n· Develop support strategies to root-cause Networking ASICs and Networking Systems issues.
\n· Develop debugging tools, for Traffic monitoring and Performance measurements.
\n· Debug issues across software, Linux systems, networking stacks, and distributed infrastructure.
\n· Be a point of contact for customer deployments, integration testing, proof-of-concepts, and field issue resolution.
\n· Build tools that improve deployment efficiency, observability, telemetry, and automated testing.
\n· Collaborate with software, infrastructure, QA, and product teams to identify root causes and deliver robust solutions.
\n· Contribute to backend services, APIs, orchestration components, and infrastructure automation.
\n· Document technical findings and communicate effectively with engineering teams and customers.
\nRequired Qualifications
\n· Bachelor’s or master’s degree in computer science, Electrical Engineering, or a related field.
\n· Minimum of 10 years of work experience is required, with at least 5 years of hands-on SONiC or equivalent Network Operating System (NOS) development experience preferred.
\n· Strong programming skills in Python, Go, or a similar language.
\n· Solid understanding of Linux, TCP/IP networking, routing, switching, VLANs, and network troubleshooting.
\n· Experience with PTF (Packet Test Framework) and SPyTest for network validation.
\n· Familiarity with Linux internals, docker containers.
\n· Excellent problem-solving skills and ability to work in a fast-paced, collaborative environment.
\n· Knowledge of network ASICs and switch hardware architecture is mandatory.
\n· Excellent written and verbal communication skills.
\n· Ability to thrive in a collaborative, fast-paced startup environment with a strong sense of ownership.
\n· Data-center networking with hands-on switch ASIC tuning and platform bring-up.
\n· Deep buffer/queueing knowledge (headroom, xon/xoff, dynamic thresholds, VOQ vs shared pools).
\n· SONiC, Enterprise-OS
\n· Optics & PHY: PAM4 signal integrity, FEC modes, DOM/RS-FEC counters, AN/LT quirks, DAC/AOC selection.
\n· Tooling: ethtool, devlink, mlnx_qos, perfquery, switch telemetry (INT/sFlow/ERSPAN), gNMI/REST, Prometheus.
\n· Benchmarking: perftest, nccl-tests, iPerf3, traffic generators; reading queue stats, ECN marks, and CNP behavior.
\n· EVPN/VXLAN leaf-spine for AI pods; flowlet or latency-aware hashing.
\n· BlueField DPU offloads, GPUDirect RDMA; NIC QoS (DSCP-to-TC, PFCx, GEARBOX/FW).
\n· Storage fabrics for AI (NFS-RDMA, NVMe-oF) and their QoS interactions.
\n· Python/Ansible for templated ASIC profiles; Git workflows for config promotion.
\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
","openingPlain":"Why join Upscale AI\nUpscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.\nWe focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.\nIf you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.\n","descriptionBody":"Role Overview
\nWe are looking for a Technical Support Engineer with 10+ years of experience who is passionate about building reliable networking switches and solving challenging infrastructure problems.
\nOwn end-to-end network bring-up and performance for GPU AI clusters during customer PoCs. You’ll be the leader for switching ASIC pipeline tuning, platform readiness, and fabric performance (Ethernet/RoCEv2 and/or InfiniBand), turning requirements into reproducible configurations that hit utilization and tail-latency targets at scale.
\nAs part of the AI Networking team, you'll debug software, develop automation, validate networking platforms, troubleshoot complex system issues, and work closely with engineering teams and customers to ensure the best customer experience. You'll have the opportunity to work on technologies powering next-generation GPU clusters and AI data centers while learning from experienced engineers in a fast-paced startup environment.
\n