Glossary
TileRT
A conventional inference engine splits work into small units called kernels and repeatedly launches and tears them down; once the time to produce one piece of an answer drops below a millisecond, that launch cost dominates. TileRT compiles the path ahead of time and freezes it into one kernel, removing the cost. It comes from the team behind the TileLang language and is published on GitHub under an MIT license, which makes it a yardstick for whether ultra-low-latency inference belongs only to purpose-built chips.
Watch 1Bear 1
Predictions & track record (2)
- Awaiting grading Nvidia paid $20 billion for speed. Did free software get there first?By the first half of 2027, check whether Cerebras actually delivers on or above its 2026 core revenue guidance of $855-865 million.
- Awaiting grading Nvidia's real moat is software, not silicon. Can AI coding tools crack it in a year?Huawei's Ascend 950DT deployment was moved up to August 2026, with DeepSeek V4.2 cited as a potential early adopter. Check by end of October 2026 whether DeepSeek's V4-series actually runs at scale on Huawei chips with published benchmarks signaling a narrowing CUDA gap.
2 related posts
- Watch Nvidia paid $20 billion for speed. Did free software get there first? SemiAnalysis · 2026-08-10
- Bear Nvidia's real moat is software, not silicon. Can AI coding tools crack it in a year? Jukan (@jukan05) · 2026-07-23