PUNE · FULLTIME
Principal Engineer ((Distributed Inference Systems)
TensorMem Inc.
Pune · onsite · Posted 13d ago
Your match
Sign in to see your match score, skill gaps & tailored resume.
Section · 01
About this role
Company Description TensorMem Inc is a fast-paced deep-tech startup at the intersection of Cloud Scale Storage and Artificial Intelligence. Our mission is to build the next generation of distributed computing and infrastructure solutions for AI workflows. Our software-defined, inference-native context-memory orchestration platform for AI, focuses on solving performance and cost challenges in large-scale inference. As AI inference grows with long-context models, RAG pipelines, agentic systems, and copilots, TensorMem addresses memory walls, KV cache pressure, and inefficient data movement that limit throughput and efficiency. The platform treats inference state, context, and memory as first-class resources across the AI memory hierarchy, enabling predictable performance, fewer GPU stalls, and better infrastructure ROI. We are looking for an extraordinary Principal Engineer to join our core engineering team to architect, design, and build TensorMem's next-generation distributed inference systems, leveraging modern AI-assisted development practices.
Role and Responsibilities As a Principal Engineer for Distributed Inference Systems, you will serve as a primary technical authority driving the architecture, design, and performance strategy for TensorMem's high-scale distributed storage and compute platforms. You will solve complex systems-level engineering challenges to deliver ultra-low latency and high-throughput AI inference across heterogeneous hardware clusters. Your primary responsibilities will include:
- Architecting, designing, and building production-grade, large-scale distributed storage and compute systems tailored for distributed AI inference workloads.
- Designing and implementing distributed algorithms, fault-tolerant consensus protocols, data partitioning, and replication strategies across multi-node clusters.
- Applying deep Linux systems expertise (NUMA, PCIe, NVMe, RDMA, custom memory management) to optimize system-level data movement and execution paths.
- Integrating distributed filesystems, object stores, or databases seamlessly with GPU/XPU acceleration layers to maximize inference throughput and minimize latency.
- Collaborating with AI infrastructure teams to optimize modern inference engines (e.g., vLLM, TensorRT-LLM, Triton) for GPU-accelerated environments.
- Setting engineering standards, conducting high-impact architecture reviews, and mentoring senior engineers across the organization.
Qualifications Education & Experience
- Bachelor’s, Master's, or Ph.D. degree in Computer Science, Computer Engineering, Electrical Engineering, or a closely related technical field from a reputed institution.
- 15–20 years of hands-on, professional experience designing, building, and scaling large-scale distributed compute, storage, or cloud infrastructure systems.
- Proven track record of architectural leadership and technical strategy in high-performance or deep-tech startup environments.
Technical Skills (Must-have)
- Exceptional mastery of Computer Science fundamentals, including advanced data structures, core system algorithms, and distributed systems algorithms/protocols (e.g., Raft, Paxos, gossip protocols, distributed locking).
- Deep, expert-level understanding of Linux systems architecture, including NUMA architectures, PCIe bus topologies, NVMe storage, RDMA network interfaces, and low-level memory management.
- Excellent programming skills in
C++ ,
Go , and
Python .
- Significant hands-on experience designing and building distributed filesystems, distributed databases, or object storage systems (e.g., POSIX, Ceph, Lustre, NVMe-oF, S3-compatible systems).
- Good understanding of modern AI inference engines and how execution, memory layout and sharing, and tensor operations work on GPU hardware accelerators.
Soft Skills
- Proven technical leadership with the ability to articulate architectural vision, mentor engineering talent, and influence strategic product roadmaps.
- Excellent problem-solving skills, rigorous engineering discipline, and strong attention to detail.
- Ability to work effectively in a fast-paced, high-impact, and collaborative startup environment.
- A strong eagerness to integrate AI tools (e.g., GitHub Copilot, integrated IDE features) into the daily engineering workflow to maximize efficiency.
Sourced from linkedin · view original
Let the agent run this one for you.
Tailored resume, auto-apply, and referral lookup — in under 2 minutes.
Section · 02