Rendered at 08:18:25 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Reubend 5 hours ago [-]
It's a Qwen fine tune that's benchmaxxing a few sets of question types, and it's obviously nowhere near beating Opus 3.8 in real usage. Nothing to see here.
Normally, I'd give them a pass because at least it's open source, but then I read
> Raising pre-seed. So we welcome all angels and vcs
Which basically gives away that they're trying to fool investors into giving them money by exaggerating their results.
kadoban 4 hours ago [-]
There is, IMO a lot of value in fine tuning existing models. If that gets them funded, great, who cares, right?
This is not their first release. BTL-2 (or -3 maybe? I forget) looked kind of interesting but it was annoying to run, needed a patch on llama.cpp, looks like they're improving their tooling.
I'm certainly interested in how this actually performs.
Normally, I'd give them a pass because at least it's open source, but then I read
> Raising pre-seed. So we welcome all angels and vcs
Which basically gives away that they're trying to fool investors into giving them money by exaggerating their results.
This is not their first release. BTL-2 (or -3 maybe? I forget) looked kind of interesting but it was annoying to run, needed a patch on llama.cpp, looks like they're improving their tooling.
I'm certainly interested in how this actually performs.