

oh I know, but since I’m running on a Z13, there’s a custom runtime built for Strix Halo that makes the model a bit faster. and that runtime uses a fine-tuned version of Ornith that takes advantages of the specific architecture to increase t/s. I know, it’s already fast, but it’s nice to see those first tokens start generating within a couple seconds, and my agent crank out a few thousand tokens in less than ten seconds.
ornith isn’t good for coding, sure, but for automating the boring shit I don’t want to pay attention to, the model works beyond expectations. I do wish it was abliterated because one of the things I use it for is OSINT workflows and tracking down the personal info of fash, and the guardrails version doesn’t always like to do that, so the skills have to carefully word the prompts. I imagine it would get pissy if I asked it hard questions about China too, but I don’t chat with LLMs so I’m not sure.



















nice