We are a local AI-focused lab researching and developing optimized performance for vLLM on SM12x hardware for open models.local-inference-lab.ai United StatesJoined December 2025
Thank you for all the support! Our first official post about our work went pretty far!
Stay tuned. We have a lot more to share with you all that will help improve your local AI experience.
We're just getting started. 🦾
With today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds.
Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️
To learn more point your agent at our GitHub here: github.com/local-inferenc…
With today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds.
Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️
To learn more point your agent at our GitHub here: github.com/local-inferenc…
With today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds.
Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️
To learn more point your agent at our GitHub here: github.com/local-inferenc…
500 toks/s at home 🤯
The thing that bothers me the most about local AI, and too few talk about this, is the fact I've already dropped obscene amounts of money and I just simply don't have enough vram.
I have two RTX 6000 pros and it's pretty rare I can find a model that I actually want to run.
It seems like four is the minimum honestly. This is gonna end up costing me over $30,000 more.
It's true that smaller models are getting way better, and I'm not discounting these relatively tiny models
But at the end of the day when there is a significantly better larger model, you're always going to want to run that.
I'm getting serious fomo.
Anyway, these guys are cracked.
With today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds.
Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️
To learn more point your agent at our GitHub here: github.com/local-inferenc…
DeepSeek V4.1 Flash weights are out and this looks like a *really* strong model. Quite a bit bigger than its precedessor, but also outperforming Opus-5, GPT-5.6-Sol, Kimi-K3 and GLM-5.3 in a few key areas.
Now trying to get deepseek v4.1flash to run on my rtx pro 6ks. Really excited to try this one. Following a deliberate process to make sure there are no max logit error and zero relative L2., and cosine similarity is appropriate. Running in eager. Nowtime for graphs and MTP lol
This is an excellent read on why more companies are investing in local AI, but not to leave the cloud. Instead to change their relationship with the cloud.
6K Followers 1K FollowingLocal LLMs on a home rig.
I love testing what models can build.
Games, 3D worlds, retro PCs, opinions.
Christian. Dad. Veteran.
437 Followers 483 FollowingSwing trader & investor since over 20 years. Tweets are just my opinions.
No investment decision should be made based on my tweets. Do your own proper dd.
316 Followers 176 FollowingFounder On Call Central & Recursive Capital. Interest in tech, finance, aviation, genetics & neuroscience. I build new things.
166 Followers 248 FollowingDo I have to put something here provocative? I just do a lot of stuff and try to do the right things. Right now having fun with local AI.
2K Followers 2K FollowingEIR at https://t.co/PsjFadhSki, Co-Host https://t.co/N0YudeVBDW, Co-Founder @AppAnnie/@DataAI, @ISEP and @Wharton alum, 🇫🇷 born and raised. Tweets are my own
2K Followers 334 FollowingRIP Dad. RIP Mamaw Joyce. Whole Ass Lawyer. General Counsel and Board Member for Local Inference Lab, Inc. AI dork and quanter. lover of bad puns.
272K Followers 1K FollowingPassionate about markets and startups.
Enjoyer of local AI, security fanatic and builder of things.
Working hard to make this world a better place.