Embedded
Foundational IP. Won India's Microprocessor Challenge — Ranked #1 of 30 finalists. Established the architectural patterns that V2 builds on.
ExSLerate is the AI accelerator IP family that powers the Krsna SoC — and is licensable as standalone IP to silicon vendors, OEMs, and chip startups building their own AI hardware. Four configurations today (Lite to Apex). Two patented engines inside V2: Dynamic Neural Compression and the Infinite Series Engine. Two generations ahead on the roadmap. ARM-style licensing model for AI silicon.
Most AI silicon companies sell you one thing. We sell you two — because the buyers are different and the economics are different. Krsna SoC is the finished chip. ExSLerate is the IP behind it. Pick the model that matches your business.
For OEMs and device-makers buying a finished, production-ready AI chip. Per-chip pricing. Reference designs included. Lite for wearables, Apex for data center.
For silicon vendors and chip startups designing their own SoCs. Licensable IP — RTL, compiler stack, integration support. ARM-style commercial model: license fee + per-unit royalty.
V1 was the 2019 microprocessor challenge winner. V2 is FPGA-validated today across the Krsna SoC configurations. V3 climbs to the deskside appliance tier, where the comparison is a workstation GPU like the RTX 6000 Ada. V4 enters the data center as an inference blade benchmarked against the H100 — serving, not training.
Foundational IP. Won India's Microprocessor Challenge — Ranked #1 of 30 finalists. Established the architectural patterns that V2 builds on.
4-configuration family — Lite (M64, always-on wearables) → Apex (M4096, robotics + automotive). Native FP8 (E4M3) and FP4 precision. FPGA-validated on AMD Xilinx ZCU106 and Kria KR260. Three innovations inside: the patented Tensulator spherical buffer, the SL Tensor Codec (lossless FP8→FP8 payload reduction), and the Special Function Unit for non-linear math on-chip. 40–50% less memory traffic and RAM at peak context — enough to run Llama 3.1 8B on an 8 GB endpoint.
Tensor Codec Gen-2 — 60% memory usage and traffic reduction. ~1,000 TOPS FP8 / BF16 dense at ~75 W, with ~800 GB/s effective bandwidth over commodity 24 GB GDDR6. Runs a 27B model at 128k context on a deskside appliance — 4× the usable context of an RTX 6000 Ada at a quarter of the power, and 2.5× its efficiency. Built for regulated businesses that need private 27B-class reasoning on-premise.
Tensor Codec Gen-3 — 70% memory usage and traffic reduction. ~3,000 TOPS FP8 / BF16 dense (6,000 FP4) at ~250 W, with ~2.2 TB/s effective bandwidth over commodity GDDR6 instead of supply-constrained HBM. Roughly 1.5× H100 compute at 4× the efficiency, in a standard rack. Built for serving 70B-class models where the economics are tokens per dollar and tokens per watt — not peak TOPS.
ExSLerate V2 ships across four Krsna SoC configurations — Lite (M64) for always-on wearables, Pulse (M256) for smartwatches and smart speakers, Surge (M1024) for drones and aerial platforms, and Apex (M4096) for robotics and automotive. Same IP family, scaled across the full edge-to-robotics envelope.
| Variant | IP | MAC count | Precision | Target |
|---|---|---|---|---|
| Krsna Apex | M4096 | 4,096 | INT4 · FP8 | Robotics · Automotive · Industrial |
| Krsna Surge | M1024 | 1,024 | INT4 · FP8 | Drones · Aerial · Light edge |
| Krsna Pulse | M256 | 256 | INT4 · FP8 | Smartwatch · Smart speaker |
| Krsna Lite | M64 | 64 | INT4 | Always-on wearables · Hearables |
V2 is shipping today across endpoint and robotics. V3 climbs to the SOHO deskside tier — Tensor Codec Gen-2 fits a 27B model at 128k context onto a single 24 GB GDDR6 card at ~75 W. V4 enters the data center with Tensor Codec Gen-3 at ~3,000 TOPS FP8 and ~250 W. Each generation pushes the reduction ratio further; each unlocks a higher tier of model size on commodity memory instead of HBM.
Tensor Codec Gen-2 cuts memory use and traffic by 60%. ~1,000 TOPS FP8 at ~75 W, with ~800 GB/s effective bandwidth over a single 24 GB GDDR6 card — half the RAM and a third of the bus width of the standard requirement, at 128k context on a 27B model.
Tensor Codec Gen-3 targets a 70% cut in memory use and traffic. ~3,000 TOPS FP8 (6,000 FP4) at ~250 W with ~2.2 TB/s effective bandwidth — roughly 1.5× H100 compute at 4× the efficiency, on commodity GDDR6 instead of supply-constrained HBM.
Detailed throughput, latency, power, and per-configuration benchmarks are released under NDA on an engagement basis.
A typical AI IP license drops you RTL and tells you to figure out the rest. ExSLerate licensees get the IP plus the runtime that's already optimized to run on it — because we built both layers together.
Synthesizable Verilog RTL for the selected variant. Verification suite. Integration documentation.
CORE compiler pre-tuned for the licensed variant. Quantization, kernel scheduling, op fusion — all included.
EdgeFlow inference engine that runs out-of-the-box on your silicon. 193 model architectures pre-supported. Built on IREE / MLIR — open frontends, no vendor lock.
Engineering team available for SoC integration, customization, and tape-out support. Not a hands-off license.
2019 — ExSLerate V1 ranked #1 of 30 finalists in MeitY's India Microprocessor Challenge. Foundational silicon recognition that seeded the IP family.
2023 — Aegis Graham Bell Award for the chip program. Selected into MeitY C2S — 1 of 13 companies in India's flagship semiconductor program.
2024 — Selected into Qualcomm QSMP as 1 of 2 cohort companies — industry-partner validation from the chip leader.
2025 — Co-development partnership with Brandworks Technologies announced. First wave of co-developed AI hardware planned for 2026.