Rendered at 20:51:46 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
armcat 1 days ago [-]
It's weird because two days after Jev was released there were a dozen decision models, a week later there are several dozen, mostly open source, OpenAI's own Decisions API [1] beats it, and you can easily finetune your own [2]. But as others have pointed out, this doesn't matter.
EDIT: As I wrote this Microsoft just released their own Decision-1 model [3].
Decision models have the potential to have an even larger impact on the Real World than LLMs have to this point (which is obviously quite large). But the model itself matters less than the product experiences you build around the model, and its very likely that the incumbent labs are treating the area as something more like "oh yeah I guess we can ship that and then forget about it" rather than investing in what building business processes on decision models looks like. Unlike full language models, I don't think the primary business of Typesafe will be serving Jev at API pricing; it'll look a lot more like putting Jev at the center of a much more expensive suite of software.
There's the potential for an inverse LLM play. In contrast with LLMs, all that seems to matter is the model, and the products the labs build around the models are all really samey and boring; the same left panel list of agents, main view agent conversation, right hand extra context, and we're now in the era of everyone creating the same cutesey furry friend on top of all this tech.
nico 23 hours ago [-]
> But the model itself matters less than the product experiences you build around the model
This is a very important insight. And it applies to LLMs as well. Very few people were impressed with the capabilities of GPT 3, it was mostly a techie novelty
But then when they added chat on top of gpt 3.5, all of a sudden it was a huge hit. Sure there were improvements in the model from 3 to 3.5, but the biggest impact was from the chat experience
Conversely, when they created Eliza, a basic chatbot more than 50 years ago, people even got addicted to it, despite it’s ai model being something super rudimentary and basic compared to what we have now. The model capabilities didn’t matter as much as the experience the chat created
Darmani 13 hours ago [-]
The instruction tuning between 3 and 3.5 was a big deal. If you prompted raw GPT 3 "Tell me a story about a mouse", it was as likely to continue "2) Tell me a story about a cow. 3) Tell me a story about a rabbit" as it was to tell you a story. I tried to use GPT 3 for work and found its only utility was to get over blank page fear by writing something so bad I would fix it in anger.
dannyw 12 hours ago [-]
It was impressive, but not that much. If you had a basic prompt prefix like:
“The following is a conversation between a human and a helpful AI assistant.
Human: [prompt]\n
AI:”
You would get consistent conversational chat, with the main limit being the model’s limited context window, and obviously frontier intelligence at the time.
I played around quite a bit with davinci-003 and conversational systems before ChatGPT. I kept using this format and API for a while, but at least during the ‘free research preview’ era, found it generally smarter than chatgpt. 3.5-turbo was probably the watershed moment where I moved away from completions on base models; to chat completions.
If you’d like to emulate this experience, spin up a base (not instruction tuned) checkpoint of Llama 1; or a more recent base model. You might underestimate how much ‘intelligence’ you can get :)
Back then the API exposed a lot of controls from sampling to logits, so you could specify a custom stop token (some rarely used Unicode); then switch to near-greedy sampling for more reliable “function calls”, etc; and then switch back to sampling for chat.
In the early days, before releasing something with the API, you had to get your use case approved by OpenAI, like the App Store review process.
kolinko 11 hours ago [-]
I think the main benefit of instruction finetuning was that the model is less likely to get prompt injections and more likely to follow the specified format?
Pure davinci was easily confused by what was system prompt, user prompt and its own answers.
9935c101ab17a66 11 hours ago [-]
[dead]
vrighter 14 hours ago [-]
i read that quote as "how it looks is more important than if it works"
acjohnson55 9 hours ago [-]
> Decision models have the potential to have an even larger impact on the Real World than LLMs have to this point (which is obviously quite large). But the model itself matters less than the product experiences you build around the model, and its very likely that the incumbent labs are treating the area as something more like "oh yeah I guess we can ship that and then forget about it" rather than investing in what building business processes on decision models looks like.
I agree that the potential is enormous, but I'm not sure why it would be Typesafe who figures out the best way to develop and package this capability. They have a bit of an advantage in the amount of time they have been thinking about this problem specifically, but I think it's a big leap to say that they will be able to dominate the category.
kijin 8 hours ago [-]
> but I think it's a big leap to say that they will be able to dominate the category.
Which is why they are valued at "only" 7.5B, while frontier labs with general-purpose LLMs have valuations in the order of trillions.
Horowitz et al. invest in probabilities, not certainty. If Typesafe doesn't make it, they'll have other horses in the race too.
echelon 6 hours ago [-]
Still, what a bold early stage bet.
I suppose the team and the surprise warrant that.
make3 11 minutes ago [-]
> it'll look a lot more like putting Jev at the center of a much more expensive suite of software
but anyone can do that, that's not their product, Jev is
olalonde 22 hours ago [-]
> Decision models have the potential to have an even larger impact on the Real World than LLMs have to this point
Why?
rebyn 14 hours ago [-]
Is it me or so far all the current HN replies to this particular “why” ask seem not even remotely answering it?
iinnPP 13 hours ago [-]
I just wanted to point out that I have started to note this "phenomenon", seemingly globally, on any topic.
The sum of people not even close to the topic has risen, according to my feels at least.
Has this rang true for anyone else?
its-summertime 8 hours ago [-]
Isn't that just normal for Hacker News?
72deluxe 8 hours ago [-]
I can attest that the quality of pizzas has risen.
SalariedSlave 14 hours ago [-]
it's not just you.
the claim is extraordinary, yet the answers cover only the mundane.
locknitpicker 4 hours ago [-]
> the claim is extraordinary, yet the answers cover only the mundane.
I wonder how many of these posts are astroturfed.
PunchyHamster 12 hours ago [-]
Because answering the why reduces investor hype train
reexpressionist 20 hours ago [-]
A major limitation of directly using "System 1 decision models" (a.k.a., logistic regression, and related uncalibrated classifiers over the output/logit space) in enterprise settings (or other high-stakes settings) is that such estimators are not reliable estimators of the predictive uncertainty in the presence of covariate shifts, and such estimators also lack a means of instance-wise data attribution (i.e., interpretability-by-exemplar), so they're not the ideal estimator for verification, routing, uncertainty over retrieval and tool-calls, etc.
For decision-making with neural networks, we instead need the older idea of estimators of the predictive uncertainty with constraints in the feature-representation space (over training/support of the estimator), as with Similarity-Distance-Magnitude estimators: https://pypi.org/project/reexpress-sdm/
SwellJoe 12 hours ago [-]
That's a lot of words to not answer the question.
dvfjsdhgfv 3 hours ago [-]
The parent wasn't giving an answer but adding a more nuance which I find valuable.
11 hours ago [-]
ghm2180 7 hours ago [-]
The simplest example could be replacement or at the least an economical complement to the LLM step in cascading VoiceAI, almost every start of the turn a local JEV could make a decision for end of user turn detection:
- is the user finished with their thought? If so delegate to next system(LLM another JEV)
- Should I wait for the user to complete their thought for a few seconds? Set a timer and when time out expires, ask the user to continue.
Like just tuning this model can make it cheaper for voice AI providers to do semantic turn detection over the space of responses.
yed 2 hours ago [-]
JEV only sees transcribed text, so long term I don’t see this use case beating audio models built for this task, or especially voice to voice models.
nowittyusername 6 hours ago [-]
That was my first insight as well. i have been working on a voice agent for a while and decided to spend about 4 days messing with system one models to see where it could fit in my cascaded system. after 4 days i ripped it out of my system as i found better solutions.. i tried really hard to find a use for it but every time i found there was no need for it so i stopped trying to shoehorn it in. i also spent time with the vision side of things for this thing, and i think that where the biggest gains will be IMO. making decision based on visual ques and in pixel space only was the biggest win.
adilkhanovkz 8 hours ago [-]
I think because pretty much every step a business takes is an allocation decision: where to put money, people, time. But it will really pay off only when the models of companies themselves become probabilistic, and the decisions are explainable, at least to some degree.
porridgeraisin 22 hours ago [-]
If youre the one or have spoken to someone implementing "AI solutions" inside large companies recently, a decent chunk of it is soft policy enforcement with very basic context. They moved from gemini 2.5 flash lite type models to jev. Which is also why I found the price comparisons to "GPT Astra" on Twitter rather funny.
zhivota 19 hours ago [-]
Right, example being expense pre-approvals (for low dollar value purchases like office equipment). General purpose LLM not required, and the product experience surrounding it does matter (reporting, auditing, fine tuning future responses, etc.).
majormajor 18 hours ago [-]
A lot of that seems like low-value-add/low-margin "simple" automation.
I haven't personally yet found a net-new-capability bigger than that from LLMs here.
sks_15 16 hours ago [-]
It really shines when you use that decision routing in industries where they need to automate these routings reliably with very low latency, higer certainly and ,more important, at scale. Take call centers for example. If a voice company needs to automate escalation etc. an llm would be an overkill. That's where system 1 models shine.
drob518 8 hours ago [-]
If it was simple they would just use an LLM to write code to do it. These decisions are “fuzzy” in the sense that they are difficult to make in just code. They need the “smarts” of a language model to evaluate something, but the final answer is basically trivial (e.g. choose one of these five options, not write me 20 page of text or one shot this whole complex program).
There are a lot of those types of problems in the world, replacing what was previously spotty hardcoded decision logic that worked 70% of the time with something that now gets it right 95% of the time. Even if not perfect, the fact that you can dramatically close the gap means a lot lower cost overall, so it’s worth it.
porridgeraisin 13 hours ago [-]
It is. But the thing is, that is also a lot of the usage for "AI" these days. LLM based search pipelines are being deployed but not really used. Chatbots are a great hit. And then of course, in coding, finance, sales, accounts and audit, etc, it is being used to accelerate output.
These guys also love text to sql. I know 4 people each implementing text to sql for their separate companies. Or text-to-redash dashboard in some cases. Funnily enough, text to sql is one of the things where LLMs are clearly sub-human. Mostly due to lack of easy verifiability.
disgruntledphd2 7 hours ago [-]
> These guys also love text to sql. I know 4 people each implementing text to sql for their separate companies.
Text to SQL is a really bad idea, as the agent will just generate really convincing and terrible analyses that lack all business context.
I mean, you could implement some kind of semantic layer, but that's a lot of work, for uncertain reward.
I saw a place which essentially tasked the data analysts/scientists in an area to describe the good tables, and assumptions and necessary context to make decent queries/analyses, and that definitely lead to much, much better (almost usable) results.
As I have said many times, what the world really needs is SQL2Text, rather than Text2SQL.
porridgeraisin 7 hours ago [-]
> I saw a place which essentially tasked the data analysts/scientists in an area to describe the good tables, and assumptions and necessary context to make decent queries/analyses, and that definitely lead to much, much better (almost usable) results.
Yep. After this has been done, the later text to sql outputs works half decent. The major issue I heard was that whenever a small assumption changes few months down the line, while humans are able to "JIT" apply it to each of their queries, these models sometimes lose it in all of the context and so it becomes hard to progressively update anything - you end up having to redo the whole workflow.
And in my experience, even the more complex dashboards requested by product or upper mgmt were stood up within the day by humans. Or the next day (but that was cos some ETL would be scheduled overnight rather than immediately). And I am not sure if speeding up this part would even matter to the guy who requested the dashboard. So I am not sure if this is a major bottleneck people should try to innovate in. If openai does the work for you then great but otherwise. Humans aided by copilot-of-old style autocomplete is probably more than enough.
Needless to say, text to sql is pointless for application queries made during API calls.
disgruntledphd2 4 hours ago [-]
Yeah, it's not the development but the maintenance that kills you in both software and other information fields.
The whole text to SQL thing is about data teams being a bottleneck to the business (rather like software people). The goal is good in theory but honestly an Excel plugin is probably a better solution in many cases.
sunaookami 4 hours ago [-]
>Decision models have the potential to have an even larger impact on the Real World than LLMs have to this point
They really haven't.
baxtr 16 hours ago [-]
I think you aptly describe the benefits this approach can offer.
What I haven’t seen answered is why it isn’t reproducible within very short amount of time and low effort.
Sure, the large labs are the lazy incumbents at this stage. But any other startups would be able to copy the experience. What’s the barrier of entry that I am missing?
pantelisk 1 days ago [-]
Yes, the best way to think of a general classifier like this is like a smart switch statement. Essentially a "JEV" like thing becomes a sort of programming primitive. Once you see it, it's hard to not get excited.
But even if others surpass them and make better solutions, the fact that nobody was able to see it before typesafe is a testament of what they might be able to come up with next.
I sound like a fanboy but I swear I 'm unaffiliated with typesafe. I was building my own version of this way before they announced JEV (mine was ALE and it was mentioned here on HN for a bit), in use for VR gaming (so one can give commands to NPCs with voice and supports multiple commands in sequence in a single pass), but I missed the "killer usecase" of being a new primitive, like everyone else.
TLDR: There's a lot of value in thinking ahead and seeing the future. The clones are nice and exciting but they give me "I could have built this first, yes but you didn't" vibes. I hope they manage to keep it up and push the space forward again
overfeed 24 hours ago [-]
> [...]the fact that nobody was able to see it before typesafe is a testament of what they might be able to come up with next.
Counterpoint: there are a lot of one-hit wonders, and they vastly outnumber the idea-factory people. This is not to minimize those people, a single idea can be very successful (see Zuckerberg), but it doesn't mean your subsequent ideas will also be great (see Zuckerberg)
aranelsurion 1 days ago [-]
I remember your blog post! Thanks for writing it, was pretty cool and a practical application.
Honestly, I don't feel the least bit of excitement here and I'm normally enthusiastic about AI.
Yes, this generates probabilities over a set of given options instead of the whole token vocabulary.
I don't see what's so exciting about it, compared to regular LLMs which support structured output options in the API.
pantelisk 16 hours ago [-]
LLMs have higher accuracy and are more versatile, but that comes with a cost (slower and more expensive to run). On the other end you can have your own mini classifier something built on top of ModernBert or DIET or even a hash classifier (these are narrow).
Models like JEV fall in the golden middle. Somewhat cheap, somewhat fast, somewhat general.
Many people have replied saying that narrow classifiers (eg bert) are better, and they are better in many ways (essentially free and faster). But... the real world is messy, in my experience it's actually very hard to make a really good classifier that will work across a specific niche domain especially when there is very fuzzy input. And the most interesting real world scenarios are nuanced and overtime there will be edge cases discovered where it fails and that will require retraining and re-evaluation etc etc. It can becomes a full project on its own that eventually collapses into a game of whackamole (improved in some direction but regressed elsewhere).
Things like JEV (or similar models) provide ease of mind, just off load the complexity to it and move on to the next challenge type of thing.
senordevnyc 22 hours ago [-]
You’re correct, it’s a great solution for a very narrow set of high volume classification needs that require very low latency. But that’s it.
I keep evaluating Jev for my product because of the hype, but the reality is that for my tasks, Luna is more accurate, only a little more expensive, and the latency doesn’t matter. I’d rather spend the extra $50 / month or whatever than have to shoehorn in another API and provider, and also lose the ability to change reasoning level and get reasoning summaries for eval purposes.
jeena 20 hours ago [-]
Did you look into self-hosted open source decision models because they are quite capable also much faster and much easier to integrate and you don't need to pay anything.
senordevnyc 17 hours ago [-]
I keep hearing Jev is better for accuracy, and still in my use cases Luna beats it. Latency doesn’t matter to me. The cost for additional Luna tokens is negligible for my use case. Idk, I just can’t figure out where to use it instead of an LLM.
But that’s just for my product.
skeeter2020 23 hours ago [-]
>> TLDR: There's a lot of value in thinking ahead and seeing the future. The clones are nice and exciting but they give me "I could have built this first, yes but you didn't" vibes. I hope they manage to keep it up and push the space forward again
This all sounds intelligent and likely, and yet we can come up with countless counter examples where the first mover is not the big winner, and nobody cares about who did it first. There is typically way more value in nailing the execution of a big idea someone else came up with, rather than "seeing the future".
zqy123007 11 hours ago [-]
> There's the potential for an inverse LLM play.
Exactly.
If JEV is here to stay, we can have another split along the "one LLM rule them all" regime
name it worker or implementator or something
- mostly for agent with with spec (not directly for human, as we tend be handwave with vague intent)
- follows the instruction verbatim
- lives in sandbox by default
- basically sol6; but WITHOUT magical situational awareness, lets work in group BS
porridgeraisin 1 days ago [-]
> Unlike full language models, I don't think the primary business of Typesafe will be serving Jev at API pricing; it'll look a lot more like putting Jev at the center of a much more expensive suite of software.
Precisely this. Should be top comment.
Also, the model moat is understated as training data for these purposes also accrues to the winner, which due to the first mover advantage as well as the distribution advantage you speak of, is typesafe. In contrast to relatively open coding data. Openai anthropic also have that, but like you say its a different business.
outofpaper 24 hours ago [-]
The are still just 1tok output of pretty standard llms just along with the logprobs converted to some json
cheesecakegood 23 hours ago [-]
At least in theory (TypeSafe has been pretty close-lipped about the details so this might just be hot air, and I think the evidence is a bit spotty) this is false, since they use a different reinforcement training method.
If the word “calibration” in probability doesn’t mean anything to you, the difference isn’t very apparent, but that doesn’t mean it doesn’t exist.
disgruntledphd2 7 hours ago [-]
> If the word “calibration” in probability doesn’t mean anything to you, the difference isn’t very apparent, but that doesn’t mean it doesn’t exist.
I am really sceptical that they've developed a zero shot calibrated classifier for arbitrary inputs.
That being said, if they have, then this valuation is like 100 times too low.
uuue 22 hours ago [-]
[dead]
nlpnerd 1 days ago [-]
Yeah, agreed. A model or primitive on its own has no moat and frankly limited value. The paradigm behind "System One" models on the other hand is potentially huge.
Skimming the post, it seems to argue for reconstructing the very rigidity that LLMs let us escape, and that very aspect of LLMs is what made them useful and explode in popularity so much.
nlpnerd 16 hours ago [-]
One may consider it "rigidity" but many application backends (including LLM powered ones) never needed the full unconstrained capabilities of an LLM. What people needed was the zero-shot capabilities over tasks as opposed to have to train a classifier, span model, etc for every task.
Constraints can be an advantage in many systems. Consider typed and untyped programming languages. Many advantages with typed languages over the latter in terms of development and efficiency.
ttul 1 days ago [-]
But Jev established the branding and investors are betting that Jev will be acquired by one of the big labs soon - and if they aren't, the money itself can create a positive outcome by allowing Jev to hire incredible talent and scale the company rapidly.
Oras 1 days ago [-]
Rapidly? Their 2 years of stealth was replicated in 2 weeks.
I give it to them for creating the hype (good marketing), and for making a useful classifier. Not sure what they would scale rapidly though.
rsalus 1 days ago [-]
I feel like half the game now is marketing though, so I can see why they'd be attractive to an investor. Maybe if they scale they can come up with something.
SaltyBackendGuy 21 hours ago [-]
Always has been imo.
Amekedl 9 hours ago [-]
Always has been indeed.
Marketing is the closest thing to a real superpower if done right, no matter how successful your product in practice is, or if it even does anything at all (see any and all youtube sponsorships kind of...)
OtherShrezzing 14 hours ago [-]
Their 2 years of stealth involved a lot more than building their model.
People who can find things of value, and execute on those things, are worth investing in - even if that thing of value is replicated in short order.
Nobody was looking at decision models, but now that Jev has surfaced, everybody is.
vasco 13 hours ago [-]
> People who can find things of value, and execute on those things, are worth investing in
Investing isn't binary. You can be investible but not at a 7.5b valuation.
tiborsaas 1 days ago [-]
You can replicate any fitness app in no time and you will make close to $0. Brand recognition matters a lot.
tyre 22 hours ago [-]
Yeah but there are also network effects of using fitness apps. Every engineer I know switches easily between Codex and Claude Code when one gets better than the other.
Jev has been around for a couple weeks. Cost and performance matter more than anything. Staying with an existing provider (the # of people choosing Jev without already having a frontier API key is probably zero?) is way easier than this.
What the hell are we even talking about.
tiborsaas 9 hours ago [-]
I think it's a ridiculous amount of money they got, but we don't know what's in the pitch deck. Investors are also blind to most of these things and take the pedigree of the founders above else.
wordpad 21 hours ago [-]
Enterprises are very sticky and they arent even chosing between claude and codex, many are looking at something like Kiro.
Bishonen88 9 hours ago [-]
Kiro came out after both Claude and codex, so not sure that any enterprise stuck to kiro unless you mean that they stick to aws in general. I'd love to hear about usage of kiro even in vague terms compared to the competitors. Naively thinking, I'd guess it's 1% or less of Claude code.
throwaway7783 1 days ago [-]
100%. Product finesse + Marketing is now the "moat".
PunchyHamster 12 hours ago [-]
"now"? It always was. Pre-AI we still had ton of great apps "losing" just because the couldn't be brought before the eyes of people that would like them
ModernMech 1 days ago [-]
It’s the classic SV flip. You scale your investors, your executive team, your sales people, hire a bunch of engineering you don’t need, then sell the company. The company’s product doesn’t matter, the company is the product.
brink 1 days ago [-]
Either the investors know something we don't, or the market is irrational.
1 days ago [-]
clickety_clack 23 hours ago [-]
It’s like openclaw. There was a bunch of technically better ones that came along afterwards, but nobody remembers what any of them were called.
jeremyjh 23 hours ago [-]
You say this, while Hermes Agent has been at top of the openrouter.ai leaderboard for several months and currently has 3X the token usage of OpenClaw.
real0mar 23 hours ago [-]
What? OpenClaw has been completely replaced by Hermes and others in the discourse
Espressosaurus 22 hours ago [-]
And Muse is what normal people use instead
mococa 1 days ago [-]
> Their 2 years of stealth was replicated in 2 weeks.
Actually they stolen the idea from a paper.
shdh 1 days ago [-]
So did Oracle with relational databases by that logic
make3 7 minutes ago [-]
Jev and decision models are *a derivative* of a pretrained LLM trunk, and one that's super easy to build once you have an LLM. Put a single step fully connected NN on the transformer, create a training dataset for it, add fancy but bread and butter multi step decision RL, nothing crazy.
*They have zero moat* against established AI companies, as the recent copies show, and will do worse than frontier labs.*
sailfast 8 hours ago [-]
Is anybody really looking to scale a company’s people rapidly these days? It’s just not necessary.
Most of the money will be burned in data centers as they struggle to sell tokens at a loss.
qsod 1 days ago [-]
[dead]
dkersten 1 days ago [-]
Most of them appear to be small LLM’s fine tuned for the role.
That’s a different set of properties in terms of size, cost, and latency. Jev (apparently, not like I’ve seen its insides) is extremely cheap, extremely fast, doesn’t cost any output tokens as it speaks the output natively, can’t get the output wrong because it speaks the format natively, and (presumably based on the docs), the context is separate from the question, meaning it should be immune (or at least highly resistant) to prompt injection attacks.
It’s not just about the accuracy of the result, it’s a collection of all the properties that make Jev interesting.
Jev took years to develop, I strongly doubt that a copycat that was put together within days after Jev’s release will be able to match it on a sun of its properties. Even if fine tuned LLMs can outperform it on raw accuracy.
TeMPOraL 1 days ago [-]
Jev is something your favorite LLM could zero-shot months ago, if you pointed it to the right arXiv paper (some of which are linked in this thread).
mrinterweb 23 hours ago [-]
That probably explains why there were so many competitors around withing days of the Jev announcement. They are not starting with a moat, and there doesn't seem to be any moat in sight. Just buzzword recognition because everything is comparing to "jev".
lifeisloving 22 hours ago [-]
There will always only be a very small percentage of people who want to discover and build things, even with llms, its a very small subset of people though, and most people want off the shelf solutions.
Also Jevs purpose isnt to become its own thing. It will get aquired in 18 months by one of Andressen Horowitz's incestuous circle of companies and everyone will make money, and the person who buys it wont necessarily care if Jev itself makes them a ton of money. They're just passing chips around the table.
TeMPOraL 13 hours ago [-]
Jev is the known idea, but with a slick website with grandiose claims and an impressive Doom demo on. It's also a catchy name - it caught on instantly, but mostly as a shorthand: it's easier to say "Jev" and "like Jev" than "classifier models" (or rather "<descriptive explanation of a specific shape of> classifier models").
devin 1 days ago [-]
Jev did not take years to develop. What it does was published in arxiv back in 2025. TypeSafe just marketed it.
baobabKoodaa 1 days ago [-]
What specific arxiv paper are you referencing here?
"SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization"
Jev is a general-purpose thing. That is a specific-purpose thing. General-purpose thing is not the same as specific-purpose thing. What makes people think these are the same thing? I don't get it.
devin 1 days ago [-]
What do you mean? Jev is trivially different from what is described in this paper.
baobabKoodaa 1 days ago [-]
Bullshit. Below is copypaste from the paper in the section that outlines the "key contributions" of the paper. As you can see, it is focused on one specific problem: predicting sales conversions. So if you were to take this system and use it for some other task ("evaluate customer mood" for example), it would not work. Because, again, it is not describing a general purpose solution. It is describing a solution that is specific to one problem: sales conversions.
Copypasta:
• A reinforcement learning architecture specifically designed for sales conversation analysis and conversion
prediction
• A synthetic data generation pipeline leveraging GPT-4O to create diverse and realistic sales conversations
• Novel state representation techniques using Azure OpenAI embeddings (3072 dimensions) with sales-specific
features
• A meta-learning approach enabling the system to express confidence in its predictions based on conversation similarity to training data
• Integration mechanisms providing real-time guidance within existing sales platforms
> Novel state representation techniques using Azure OpenAI embeddings (3072 dimensions) with sales-specific features
Jev is basically the embeddings side of an LLM. Yes, it's a good idea, but the moat is non-existent.
baobabKoodaa 22 hours ago [-]
No. It can't simultaneously be both general purpose and having task-specific embeddings.
devin 16 hours ago [-]
Hard to take you seriously, fam. The distance between these two things is trivial. The author of the paper implemented the "generalized" version of same in less than 24 hours. It's not novel, it is a continuation of what already existed in a boring way
baobabKoodaa 14 hours ago [-]
The comment I was responding to describes Jev's significance as being specifically about embeddings:
> Jev is basically the embeddings side of an LLM
This is how the paper you are referencing is describing the "key contribution" as it relates to embeddings:
> Novel state representation techniques using Azure OpenAI embeddings (3072 dimensions) with sales-specific features
And now you are claiming that the author of the paper created:
> "generalized" version of same
It's a bit hard to guess what you are trying to say, but if I were to steelman your argument, I would guess that you mean: training a model with a vocabulary that has some tokens representing things like "OPTION_A", "OPTION_B" is in your mind the same thing as "generalized version of sales-specific features in embeddings"? Is this what you were trying to say?
devin 4 hours ago [-]
I said the author of the paper created a generalized version because they posted about doing so on reddit. They way they described the work, it sounded trivial.
zwaps 1 days ago [-]
This is a fine-tuned model. The author even states that the model is competitive with Jev only if fine-tuned on the evaluation at hand.
Literally misses the point of Jev, which you don't need to fine-tune to get accuracy nor - and no other model has this - some sort of out of sample calibration
nowittyusername 6 hours ago [-]
JEV's moat is calibration nothing else. All other llms excel at system one like judgement versus JEV, but that's not a dunk on JEV. Calibration is very important for some niche use cases.
hbrn 1 days ago [-]
> I strongly doubt that a copycat that was put together within days after Jev’s release will be able to match it on a sun of its properties
But why? If the simplest way to achieve Jev's capabilities (accuracy, cost, latency) is by fine tuning a small model, what makes you think that this isn't exactly what Typesafe did?
And even if they did something different - what makes you think it was a good idea in the first place, given how easy their results were replicated without any "secret sauce"?
dkersten 1 days ago [-]
My point is that it’s not replicated. You replicate the accuracy, but not the other properties. The Jev-competitors only proved that you can get or beat the accuracy, nothing about the other properties. Especially the “zero hallucination” output and the (if it works how the documentation make it sound) prompt injection resistant architecture. You can’t get that with a fine tuned LLM.
hbrn 1 days ago [-]
> zero hallucination
Plenty has been said about this claim. If you're still falling for this, I feel sorry for you.
If you remove wheels from your car, your car will get a "no speeding ticket" property, and yet there's nothing exciting about it.
> You can’t get that with a fine tuned LLM
Of course you can. All these claims are nothing but marketing.
user43928 23 hours ago [-]
I understand most major model providers support passing a JSON schema that is strictly followed in the output, accomplishing the same 'zero hallucination' and prompt injection resistance.
The only difference I am aware of is that probabilities are better calibrated with these decision models compared to regular LLMs which can output hallucinated numbers where your schema allows a number.
dkersten 13 hours ago [-]
I’m not claiming you can’t get some of the properties in some ways with LLMs, just that Jev is the sum of its parts, not just one or two properties.
LLMs can produce structured output by limiting the next token based on a grammar, that ensures correctly formed output.
But (presumably, again I haven’t seen Jev’s insides) Jev doesn’t need to do that, it just has to output probabilities for each answer, and can do it natively without artificially limiting output tokens.
So jev is cheap, fast, never produces malformed output, its output always carries correct semantic meaning (but this doesn’t mean it always answers correctly), and it separates correct from prompt (which if it truly does that is the most exciting part).
PunchyHamster 12 hours ago [-]
Took me about 20 seconds to prompt inject the demo instance.... y'all would be great scammer targets
creatonez 15 hours ago [-]
Jev is almost certainly not immune to prompt injection
dkersten 13 hours ago [-]
It separates context/state from prompt/questions. That’s a key ingredient for prompt injection hardening. If it is designed and trained to never treat state as instruction, and you only ever put untrusted user input into state, and never the questions, then why wouldn’t it be?
The jev documentation says this is what you should do. It doesn’t mean that it’s actually designed for prompt injection resistance, but it could be. We won’t know for sure until they either release more details or someone proves otherwise.
But the split is one that doesn’t exist in normal LLMs and it’s the exact split that is needed for immunity or resistance to prompt injection.
scottyah 1 days ago [-]
> OpenAI's own Decisions API [1] beats it
Have you heard that from a different source than OpenAI? From what I'd heard other models haven't gotten close, and the open source ones are like running gemma4 E2B against Opus 5.5- sure, the API calls go in and are returned the same but the quality isn't close.
rockwotj 19 hours ago [-]
I have ran a bunch of evals on production use cases where we are using Jev and OpenAI's model isn't close, but I'm hopeful about some of the OSS ones. IDK it's very cheap so I don't really feel like I need to look for an alternative. Why root for OpenAI over Typesafe? Isn't more diversity in the market good?
shmoogy 22 hours ago [-]
Jev outperforms clef and OpenAI decisions on most of my tasks, outside of when I need multi modal image going into it. Jev is also cheaper but costs are minimal overall.
lifeisloving 22 hours ago [-]
I just dont understand why anyone needs a general purpose classifier. Just build a classifier for your specific use cases.
Fortunately Jev is cheap, so I dont think it matters too much, but I think its robbing people of the opportunity to learn and implement this themselves.
Also, I dont really want 3 companies responsible for censorship/classification.
jofzar 17 hours ago [-]
That's why it's useful, it's for when you don't know what needs to be classified or if you want to offer this classification models to your users.
baq 10 hours ago [-]
It’s not training that’s the hard part, it’s when you don’t know your use case or it keeps changing all the time
nlpnerd 16 hours ago [-]
Because training your own model requires substantial effort around data annotations, and then dealing with data drift etc in production.
nwienert 14 hours ago [-]
Hm I had built out a system that tested a few hundred decisions and clef beat Jev pretty handily, I think there was actually 0 instances where clef regressed and it improved about 20% of the false decisions.
andrewingram 10 hours ago [-]
I had the opposite experience. Clef was worse in accuracy, and an order of magnitude slower.
fennecbutt 12 hours ago [-]
Because the American stock market is all overvalued stuff with monopoly money.
vrganj 11 hours ago [-]
Marx called this fictitious capital.
> Fictitious capital could be defined as a capitalisation on property ownership. Such ownership is real and legally enforced, as are the profits made from it, but the capital involved is fictitious; it is "money that is thrown into circulation as capital without any material basis in commodities or productive activivity".
no, you can't, and it's unclear why you would think this.
ricericerice 1 days ago [-]
you can easily finetune your own*
*if you have a sufficiently sized and quality dataset for the specific classifications you're targeting
9dev 13 hours ago [-]
…and a team with experience and the hardware and the workflows to do this at scale, repeatedly to avoid drift, with proper infrastructure for evaluation and testing, and all of this work is somehow the core domain of your business.
Then yes, easy!
baobabKoodaa 1 days ago [-]
And even if you do have that, you haven't made your own Jev, because Jev is a general-purpose thing, whereas what you have built is a specific-purpose thing.
jgilias 1 days ago [-]
Isn’t the OpenAI decisions API basically just Luna cosplaying a decisions model and pretending the confidence score isn’t just a hallucination?
hbrn 24 hours ago [-]
And what do you think Jev confidence score is?
Here's a hint: confidence is not generated by a model.
adrian17 23 hours ago [-]
Maybe I'm missing something, but why couldn't it be generated by the model? In older classification tasks with transformers like BERT, you could absolutely obtain a confidence score.
hbrn 22 hours ago [-]
Jev API returns both confidence and probabilities.
But confidence value is just a function applied to probabilities. It is not coming from the model, and it carries no additional information.
It is documented btw, and yet you will see plenty of claims that Jev is better than LLM because it returns both.
jgilias 23 hours ago [-]
Thanks, fixed my understanding!
Do you think though that Luna being a model post-trained for chat produces over-confidence in logprobs?
hbrn 23 hours ago [-]
Yeah, but I wouldn't be surprised OpenAI's decision API is a post-trained Luna with confidence calibration.
Typesafe claims that Jev is calibrated, but there are plenty of examples where it completely fails (predicting die roll being the most obvious one).
Unfortunately calibration is hard to benchmark.
shados 22 hours ago [-]
The dice roll prediction is about the way the prompt is setup misunderstanding how Jev works (they treat the confidence score as a probability score, which it isn't).
If you instead give it a list of probability for each number and ask it whats the probability of each number, the result will be accurate.
hbrn 21 hours ago [-]
> give it a list of probability for each number and ask it whats the probability of each number
Did i hear that correctly? In order for Jev to be accurate you have to give it the answer before asking for the answer?
(btw this is exactly how Jev is playing games).
jgilias 14 hours ago [-]
That’s circular, sure. But if we think about potential real-world tasks where someone might, say, use it to classify on “does this advice correspond to our policy docs”, you’d absolutely push the answer (the policy docs) into its context..
What am I missing?
As in, neither LLMs nor Jev are truth engines. Truth comes from the provided context.. Plus weights.. kind of fuzzy, sorry I’m thinking out “loud”
hbrn 4 hours ago [-]
> does this advice correspond to our policy docs
This is not a “system one” question.
You’d be much better off using a proper LLM here: more accurate, can provide justification, can be steered when mistakes are made. Jev is going to be a coin flipper with almost no knobs.
phalangion 1 days ago [-]
What’s the difference?
charm137 22 hours ago [-]
I see Typeface AI have 30-40 people working for them on LinkedIn (some are VC advisors / board members), a lot of them being engineers. Their openings suggest a strong developer market focus (to begin with). I don't necessarily see a lot of people at the company with enterprise sales channel experience, but they already claim to have several Fortune 500 companies as clients (maybe the VCs are helping there or there is natural dev-driven traction).
So they clearly have a product, a strategy around it and perhaps the compliance scaffolding (SOC2 Type II etc) that may be needed before actually being able to charge money for it. They also have the right brand names associated with the founding team. Execution, so far, seems good enough to create a splash, at least.
As an investor, the question(s) to ask is (in my view): "How do they make money? Will that way to make money survive?". The answer to the first: selling input tokens and perhaps subscriptions/credits eventually.The answer to the second: "Yes, but with the risk of unit revenues declining faster than their unit costs". How can they mitigate this problem: by being big (scale / mindshare etc) so that their unit costs (including for customer acq) fall faster than their unit revenues will - I believe that is the question most AI companies are trying to tackle these days. Any new competitor will have to tackle basic fixed costs (of getting started) first before even getting to the stage of having the luxury of worrying about unit-economics.
So yes, they might eventually be competed away but whoever is in their team is trying hard to make a useful product/ecosystem and that should be applauded, not ridiculed with "it's all marketing". This is way more than a simple github/huggingface-based open-source replica solution can hope to achieve without institutional backing (either big-tech or system-integrators).
What should rightly be questioned, of course, are the valuations the VCs are providing to them in hopes of passing this hot potato to a willing buyer (say a hardware maker like NVidia) - the incentives there are very well defined and depend very much on perceived TAM (which lately is on very shaky ground given how far token pricing has fallen causing, among other things, OpenAI to "miss" on the market's expectations for annualized revenues, even before they're listed!) [1]
Currently you can’t even add a VAT ID to your account, leading to invalid invoices for all EU customers.
So yeah, not too much enterprise sales experience there for sure.
notfromhere 21 hours ago [-]
An Enterprise client at that size can just be someone at an f500 put a credit card in
girvo 1 days ago [-]
Counterpoint: my work has already allowed us to call and test Jev. Those others? Who knows when, if ever.
nico 23 hours ago [-]
Yup, I also released an open source classifiers tool, Jeffy. It comes with 68 pre trained classifiers which run and train on CPU alone. They run locally and are faster than Jev/Laya/Decisions. And they can do things like label email, all the way to even playing Doom
Did you actually try them, I tried a sample they were rubbish compared to Jev - I'm not saying they will survive, but the competition CURRENTLY sucks - and also they aren't actually cheaper if you need them at scale. OpenAI is coming in twice the cost and so is Cloudflare. Ad yes you can finetune, and yes I do, but it also SUCKS. Fine-tuning and managing your own datasets is yet another timesink. So we'll see, but it's not quite black and white.
phoghed 19 hours ago [-]
Yeah, tried a couple open source ones, same experience. Seems to me a ton of people are in a mad dash to create one of these and most seem to have over fit on the Jev benchmark to get some claim to fame of beating them or replicating their couple years of work in two days or whatever.
Then everyone breathlessly repeats this story about how Jev is useless because open source models they never tried claim to do the same thing and better.
jofzar 18 hours ago [-]
Testing I have seen from internal teams have shown it's still the best and cheapest model so far for pure classification work.
Ozzie_osman 14 hours ago [-]
It's possible that Jev has more in the pipeline. I mean, for shipping something so paradigm-shifting, this would be a bet on the founder/team continuing to do things faster or better.
c7b 14 hours ago [-]
Paradigm-shifting? A classifier? It's great and it's a cool use of Transformers, but AI really didn't only start with LLMs. We used to have an own term for AI models that can handle text instead of just numbers, now a model that outputs numbers instead of text is considered paradigm-shifting.
Ozzie_osman 8 hours ago [-]
I mean, it's obvious in retrospect, and technically not too hard to implement... Combining the power of LLMs with the constrained nature of classifiers. But no one did it well, and they did it first, and it took off.
c7b 7 hours ago [-]
Sure, or let's say, no one else who was doing it before was noticed at this level. But hype doesn't automaically equal a paradigm shift. In this case, it's going back to what used to be one of the most common ML applications until not so long ago. If anything it feels like going old school (and I really mean that in a good way).
throwaw12 1 days ago [-]
You are right in terms of how fast competition created alternatives.
But, for OpenAI this is not a primary business, for open source models as well, so they will not be chasing the market and customers to buy their product and promise them to maintain it.
TypeSafe will do all this, they will try to understand your use cases and then solve your pain point, while others are providing raw material.
rpdillon 23 hours ago [-]
In my experience with OpenAI's decisions endpoint, it tends to return either 0 or 1 and doesn't return middle confidence levels very much at all. Would be interested to hear if others have experienced the same.
fastball 18 hours ago [-]
And there are dozens of open-weight LLMs, but OpenAI and Anthropic are both immensely valuable because their models are actually intelligent.
bushbaba 1 days ago [-]
A major VC could type safe ai money, then head to a larger AI company looking to raise their series E+ and demand they acquire typesafe as part of their funding allotment.
such an arrangement can end up beneficial to the VC firm
vrighter 14 hours ago [-]
because all that is needed is to delete the outer while loop. it's an llm that is only run once and generates just one token. But it outputs confidence levels! Yeah confidence levels are the only output from an llm
shados 22 hours ago [-]
I honestly was worried for them. Now, even with all the clones, they generally still come up on top in price and latency, but "good enough" is often sufficient.
Guess the (investment) market has spoken.
nacs 22 hours ago [-]
Hm? Jev's actual (network) latency is not that great and even small local models are doing far better latency-wise.
The Microsoft article GP linked even shows the MS model having 95ms latency.
Even on price Jev being matched (the same MS model is "Input tokens cost $0.042 USD per million tokens. Output tokens are free.", same as Jev).
Cloudflare's Clef-flash model is actually slightly cheaper: "$0.038 in / $0 out per 1M" too.
Jev is being matched or exceeded in performance and price within a month of them going public.
sy135673 14 hours ago [-]
[flagged]
1 days ago [-]
hkalbasi 1 days ago [-]
> OpenAI's own Decisions API [1] beats it
Jev is 42$/B but OpenAI is 100$/B token.
xueyu 13 hours ago [-]
[dead]
amelius 1 days ago [-]
I mean I'm already ditching my Apple stocks because soon AI will be able to replicate iOS and MacOS.
user3939382 1 days ago [-]
Investments aren’t made because the product is amazing, they’re made because there’s a compelling exit scenario. Engineers don’t want to hear this but more generally, the critical success factors for a business aren’t product or engineering they’re relationships i.e. sales and team dynamics. If technical excellence dictated business outcomes in tech Salesforce wouldn’t exist for example.
nlpnerd 1 days ago [-]
You are assuming that the VCs have done their due diligence. For a "hot" company like Typesafe AI, most likely little due diligence was done. That's the way it's played.
doctorpangloss 1 days ago [-]
"It doesn't matter"
By all means, become an A16Z LP.
moralestapia 1 days ago [-]
Nothing beats nepo, brother.
mlmonkey 1 days ago [-]
OpenAI's "Decisions" library has this in requirements:
To run the SDK examples below, use these OpenAI SDK versions or later: Python 3.26.0,
I thought Pythin 3.15.0 just came out, 3.26.0 must be really far off?
bayesianbot 1 days ago [-]
That is their Python SDK version, not Python version
1 days ago [-]
christina97 1 days ago [-]
Everyone appears surprised by this news. It’s clear that they don’t have a product with some incredible moat. But they clearly have good engineering and product people that came up with a product people wanted. On top of that they have very strong marketing muscle that took the AI world by storm. And as far as I’ve seen, they still lead in some part of the latency-quality (-cost) curve?
They may well be a good team to throw money behind if you are hoping to bet on a new AI lab.
cootsnuck 20 hours ago [-]
> On top of that they have very strong marketing muscle that took the AI world by storm.
Not them per se, but doomers.ai [0] which is a uh... "launch virality agency".
So less that Typesafe has a strong marketing muscle, and more that they paid at least $100K [1] for "organic" buzz.
you’ve got to be kidding me. i know it ISNT fraud but i can’t help but feel like this SMELLS like fraud
Den_VR 10 hours ago [-]
Why not both. The engineering side isn’t fraud, but the business side is doing business things “for the sake of the project.”
adrianmsmith 13 hours ago [-]
Why?
I read the attached links and it just looks like a marketing agency?
asa123 6 hours ago [-]
i mean marketing in some sense “feels” like fraud when presented in some very specific way, no? at some level it’s: generate interest about a thing when there didn’t exist any interest before. sometimes they generate interest about something that has no substance and that feels kind of fraudulent (this isn’t that case)
but in this case, this is product extremely boosted for super viral marketing, and this is the first time i’ve read about this boost since hearing of the product whereas everybody only really talks of its substance. feels like a magician kept a trick going for a very long time, people forgot they were watching a magic show, then the curtains come down and they reveal the deceit after (i’m being slightly melodramatic)
hbrn 23 hours ago [-]
> people that came up with a product people wanted
We've yet to see whether this is true, or is it just manufactured demand. There are dozens of Jev demos, but pretty much all of them are either cool but useless, or simply fake (i.e. harness doing 99% of the work).
But economics of Jev enable classifying high volume things. And I've tried all the open-weight Qwen based decision models on my Spark and none of them came closer to quality and suck at batch inference so they are doing something custom.
randysalami 7 hours ago [-]
Silly question but is it right? I am in grad school and this semester we run multimodal data through ML and LLM-as-judge classifiers. Even plugging in ChatGPT with a relevant judge prompt produces relatively believable results. At the same time, I struggle to trust it to generalize. I’ve read that Jev does classification across many domains and is also powered by an LLM back-end? What separates it from my above approach outside of harness efforts and finetuning of the LLM? Thank you!
vanuatu 22 hours ago [-]
since their release they surpassed 100M arr and 1/3 of the F500
its clear they own the mindshare around this type of primitive which is a massive premium
hbrn 22 hours ago [-]
Measuring ARR based on 7 days of data is beyond stupid.
You can create a company with 2B shares and sell one share to your friend for $1000. Lo and behold, you own $2 trillion dollar company, leaving Elon behind.
qlte 22 hours ago [-]
Especially for a company still in their 15 minutes of fame that could very plausibly turn out to be a one hit wonder and a trivia question in two years.
Multiplying out revenue from the peak of their mini hype cycle while a dozen well financed competitors target them directly seems extremely optimistic.
hbrn 21 hours ago [-]
Ha, two years is too generous.
If they weren't lying about $100M ARR, they had to be at 6.5T/day for a week to hit that number.
They also claimed they were at 1T/day just couple weeks ago. That's $15M ARR.
85% drop in just a couple of weeks?
dannyw 8 hours ago [-]
1/3 of the Fortune 500 is likely measured by something like a single @company.com put a credit card in, against their company’s procurement and security policies.
It’s not to be unexpected for a new product and startup, but they lack SOC 2 certification, and the other things enterprises look for on information security.
I’d be surprised if many procurement divisions are approving typesafe AI, and I fear by the time they achieve it, existing vendors would’ve catched up.
nikanj 8 hours ago [-]
It’s more likely just ”Marketing put that claim on the page, who could ever disprove it?”
baby 8 hours ago [-]
What? The moment this was released we integrated decision models EVERYWHERE. If you’re not using it you’re basically like these people who think LLMs are useless lol
soleveloper 1 days ago [-]
They don't lead on latency nor quality; but their execution was superb
According to this benchmark, Jev is currently trailing Quyet-1.0-Large and a few other hastily put-together LLM-based decision API-like setups.
tpetry 1 days ago [-]
And the 'better' ones are slower and cost more for a tiny bit more accuracy. Its hard to sell that as being better when speed and price have been JEVs main selling points.
baobabKoodaa 24 hours ago [-]
The top alternative right now lists speed as faster than Jev?
senordevnyc 21 hours ago [-]
The space of problems that require super low latency AND are inputting so many tens of billions of tokens that the price matters, is a tiny one.
They can be both great and over-valued at the same time.
rvz 1 days ago [-]
This is all due to 40% marketing, 50% execution and 10% credentials (with the founders being associated with creating ChatGPT).
If anyone else came up with the same concept on a Reddit thread (they have) it no-one would care without those characteristics even if you are "first".
Rebranding, execution, marketing, ex-<big_name_company> and mostly importantly, hype is what gets the investors scrambling into throwing money at you.
objektif 24 hours ago [-]
More like 10% 10% 80%.
JamesSwift 19 hours ago [-]
35% 10% 55%
binlog 24 hours ago [-]
The model itself is a negligible part of the valuation. The company is priced as an acquisition target.
redanddead 24 hours ago [-]
Every startup is priced as an acq target
verdverm 1 days ago [-]
Wonder if they can get coin flips and dice rolls to make sense with this fresh funding, or if it even matters to people.
I have no faith in the technique if it cannot do the basics (i.e. not real probabilities, the confidence for coin flip outcomes)
the "not real probability" disclaimer only appears after you get a result
prometheus1992 1 days ago [-]
I really don't understand how this can be. I have sat in fund raising meetings with VCs in toronto and my experience is that there is shit ton of due diligence at the tech level. a product which has no moat, was already available, was duplicated within a couple of days is valued at 7B - i thought we were past the peak of the hype cycle.
fidotron 1 days ago [-]
> sat in fund raising meetings with VCs in toronto
There's your problem. The single biggest thing every Canadian VC is trying to figure out is "why are these people asking us for money when if they were any good they'd be in the US" so by simply asking them you're already signalling something bad. A lot of their enthusiasm for process is based on this suspicion and also that the entire industry is just a way for various professional services to extract most of the investment money, since that's the game they're so used to playing with the government.
There are some Canadian VCs earnestly trying to improve but they are overwhelmingly hilariously conservative and focused on unimportant signals over reality. This is one (but not all) of the major factors that drive basically every remotely ambitious Canadian company to run a corp in Delaware and go for funding from the US. The tax situation is the other major contributor.
redanddead 24 hours ago [-]
Canada doesn’t have throwing around money like in the US. We have resource extraction -> export money that’s it
cmrdporcupine 23 hours ago [-]
There are boatloads of tax breaks and incentives that mitigate all the financial stuff and make running a startup here just fine honestly.
But that does nothing to make up for the terrible investment community. Getting started here requires already being started.
When I briefly worked for a Toronto startup, it was like all of them went to the same private boy's schools together as kids. It was a status club.
I jumped ship to an American startup and made almost double the money dealt with 0% of the bullshit and they were bought by Google the next year.
geoffschmidt 1 days ago [-]
There is a belief that there is going to be at least one more breakout success in startup AI labs - rather than OpenAI and Anthropic being the final word - and so investors want to own a part of whichever companies seem most likely to be that success. If you start from that premise and stack rank what company that might be, you could quite reasonably put TypeSafe toward the top of that list right now, based on the people at the company and the ability they've demonstrated to ship stuff that people care about and cut through the noise in a crowded space.
Also the situation isn't static. Investors know that the act of writing them a $870M check itself increases the chance that they'll be one of the winners, because that will attract more talent, customers, and funding to the company in a self-reinforcing cycle. And investors know that other investors know that, and that someone is going to write them that $870M check, so to some extent they're forced to think of the company as having already been successful at the fundraising and already having that momentum boost.
Only a small number of investors in the world can play the game at this level, because you have to smart enough to be right (often enough), and you have to be established enough to see the deals (be on every CEO's short list - because CEOs are only going to seriously pitch 5-10 VCs on a hot deal, if that). Otherwise you can't pull it off. Martin Casado and his team are among the few that can and I think their results reflect that.
prometheus1992 6 hours ago [-]
I think this is a reasonable take and I agree with you but its a bit unsettling that the event of fund raising has become a more important variable in the equation than the fundamental innovation itself. And I don't think this is a good long term strategy- it almost gives me "let's con our way out of this" vibe.
strgcmc 6 hours ago [-]
I think this is an astute comment and a reasonably accurate interpretation of how investors might be thinking about these opportunities.
At the same time, hopefully you would agree that this "self-reinforcing cycle" that you describe, if it really is the dominant way of thinking, is also a pretty clear sign that we're well into bubble territory, divorced from fundamentals.... Investing because you think your sheer act of investing is going to create a self-fulfilling prophecy of success while simultaneously having FOMO about someone else beating you to it? I dunno if I've seen a better description of what it looks like to make investment decisions from the POV of being deep inside a reality-distorted bubble while chasing mania.
ejeq 1 days ago [-]
[dead]
reticulates 1 days ago [-]
The lack of a “moat” is mostly irrelevant because success is not decided by who can or can’t be cloned. TypeSafe invented[1] a new approach that became wildly popular almost immediately, if they can do that once, they can probably do it again. Venture capital is big bets, of course TypeSafe is going to fail, that’s inevitable, but if it has even a 10% chance of capturing 1/10th the market cap of OpenAI then it is a great investment! Plus, money means nothing any more, they’ve raised less at a lower valuation than Instinct, a personal assistant.
[1] not really but they did some innovative things and popularized a concept
nateb2022 1 days ago [-]
I'd guess they justified the funding by revealing some grand scheme for a new product that they just need more runway to produce.
baobabKoodaa 1 days ago [-]
> was already available
no, it was not
> was duplicated within a couple of days
was it already available or did it become available in a couple of days? it cant be both (neither is true, actually)
hirako2000 1 days ago [-]
The tech behind it existed. They made it a specific product, and got replicated in days.
victorbjorklund 9 hours ago [-]
That is like saying ChatGPT wasn’t anything new because text models had existed for years
baobabKoodaa 1 days ago [-]
Unclear what you're referring to. Please stop making vague claims and be specific.
euleriancon 1 days ago [-]
I think it is clear he is referring to zero shot classifiers with an LLM backbone. That tech has existed for a long time.
prometheus1992 1 days ago [-]
maybe you entered the AI space during the vibecoding era but there had been ton of useful models before that. especially zero shot models- both for text and images.
baobabKoodaa 24 hours ago [-]
Everything you said here is false. No, I didn't enter the AI space during the vibe coding era. I was training custom ML models back in 2017. And no, there haven't been models comparable to Jev before Jev was published.
Jev is:
- accurate
- general purpose
- fast and cheap
Models we had before Jev had at most 2/3 of above qualities, but none of them were 3/3.
gpugreg 22 hours ago [-]
> I was training custom ML models back in 2017.
Maybe you are a good person to ask my question then. I have not looked into Jev much, but is it much different from using a regular LLM and constraining its token output to the action space? (e.g. like using llama.cpp's GBNF grammars). Is it just that Jev's "confidence scores" are significantly better than the softmaxed logits? Or is there something else I am missing?
baobabKoodaa 22 hours ago [-]
What you're missing is: cost and speed. Otherwise, it's very much like running an LLM and constraining the output.
I don't know how good Jev's "confidence scores" are, but I would be surprised if they were in any sense better than logits from some good LLM. One advantage of Jev here is that the confidence scores are easy to access. Most LLM API providers don't provide an easy/convenient way to access the logits. But that's a minor point, you could of course build something like this with LLMs (and many people have).
gpugreg 21 hours ago [-]
llama.cpp is already extremely fast for single-token responses (<5 ms). I can't see Jev being faster when taking network latency into account, except maybe for multimodal inputs.
baobabKoodaa 21 hours ago [-]
Sounds like you are running a tiny toy model if you can get generations in under 5 ms? Typical response times from LLMs for typical "jev-like" queries from OpenAI and Anthropic are 2s-10s. Not milliseconds. Seconds. Same queries from Jev are like 0.2s. and the cost is 1000x.
hbrn 24 hours ago [-]
> Models we had before Jev had at most 2/3 of above qualities, but none of them were 3/3.
You're the one being deceptive here. Jev is trading accuracy, speed, and cost for generality. It's less accurate, slower and more expensive than trained classifiers. So it's still 2 out of 3, but with decimals. Maybe 2.2 out of 3 if I'm being charitable.
And the reason we didn't have that before is because nobody thought it's a good tradeoff.
baobabKoodaa 23 hours ago [-]
When you say "trained classifiers", you are referring to models which are trained (or fine tuned) to work on one specific problem, right? That is the opposite of "general purpose".
Would Jev be more accurate in a specific task if it had been developed only for that task, as opposed to general purpose? Of course it would. So, sure, Jev is trading accuracy for generality. According to you "nobody thought it's a good tradeoff", which again is false, there was huge demand for a cheap and accurate general purpose classifier.
hbrn 22 hours ago [-]
> work on one specific problem, right? That is the opposite of "general purpose".
A business doesn't need Jev for the sake of Jev. Most business are solving specific problems.
And fine-tuning got a lot cheaper these days - I've seen claims here on HN that ~500 examples is enough to beat Jev.
> "nobody thought it's a good tradeoff", which again is false, there was huge demand for a cheap and accurate general purpose classifier
There wasn't. The hope is that there was a latent demand, but we've yet to see if it's truly latent or just manufactured.
Noone is saying "hell yeah, finally we got a general purpose classifier, my business needed it so much". The typical message is "this seems cool, let me see where I can apply it".
The fact that name itself is a play on Jevons Paradox illustrates that there was no demand until Jev was released.
baobabKoodaa 22 hours ago [-]
Yes, businesses are solving specific problems, but most businesses have more than 1 problem to solve. No, it is not economical to pay a data scientist to develop a custom model for each of your tiny problems. It is often much more economical to use a general purpose solution, like an LLM, or now, Jev.
majormajor 18 hours ago [-]
Are there enough "tiny problems" for businesses in the world to justify this sort of money?
The appeal and claim of LLMs for businesses is solving big problems.
LLM valuations to solve tiny problems seems iffy.
--
Are there significant problem domains LLMs are bad at that Jev is good at? Vs just 'Jev can do a subset of LLM things faster/cheaper'?
baobabKoodaa 8 hours ago [-]
Jev can do a subset of LLM things faster/cheaper. Jev is not better than big LLM's at anything.
hbrn 22 hours ago [-]
At this point there's no need to pay a data scientist. You can literally ask Claude to do everything for you: extract real examples, classify them, post-train a model, and ship an API.
Now, it could be viable if your business has literally hundreds of problems thats require classification. I just haven't seen those.
I treat the fact that almost noone was doing that as evidence that decision models aren't that useful/groundbreaking. That, and the fact that every single demo I saw was either fake (e.g. playing games), contrived, or plain wrong (e.g. using Jev for compaction).
baobabKoodaa 22 hours ago [-]
No, we're not at the point where you could ask Claude to do all of that, unless you have a super easy problem to begin with and/or you don't care about output quality. Feel free to link a counter example.
majormajor 18 hours ago [-]
[dead]
prometheus1992 23 hours ago [-]
>>I was training custom ML models back in 2017
but you truly do sound like an angry 19 year old from your arguments.
- accurate - on what? on trust me bro benchmarks?
- zero-shot model are fundamentally general purpose.
- fast and cheap ; models on hf are FREE and fast enough.
baobabKoodaa 22 hours ago [-]
"Models on hf" (unspecified) are "FREE"? Like "free to download"? Sure, but nobody was talking about that. They cost money to run inference on. Unless you are talking about some tiny toy models that are useless for any non-toy problems. You clearly don't have any idea what you're talking about. Just stop, man.
havercosine 1 days ago [-]
Its not normal time in SF/Bay Area. For better or worse, VCs in this city/region are thinking very differently on AI bets.
I think the key differentiator was that a team found a whitespace in what ChatGPT was doing, main comes from the same pedigree and team is as conscious of marketing as their product. SF VCs love these out of the box challengers, and people are claiming to replicate doesn't seem to matter.
The amount raised feels surprising but again entire SF/US AI scene is primarily "add moar layers and GPU" one trick ponies at this point.
majormajor 18 hours ago [-]
> Its not normal time in SF/Bay Area. For better or worse, VCs in this city/region are thinking very differently on AI bets.
"Throw a bunch of money at copycats in the trend of the day" is very normal in SF/SV for the last several decades, going back to the dotcom boom.
The definition of "a bunch" has changed but so many "crypto for X" or "recommendations for X" or "uber for X" or "social for X", etc, things raised amounts that seemed wildly divorced from their market position.
binlog 24 hours ago [-]
Because every VC knows that one of OpenAI/Anthropic/Nvidia/Microsoft/Google/Meta/SpaceXAI/AMD/Stripe... will acquire them within the next year for talent alone.
nlpnerd 1 days ago [-]
This is the difference between a "hot" company in a good ecosystem like SF. Yeah, the funds can take 2-3 months to do their due diligence. By the time it's done, the round has closed, and then what good is the due diligence?
baby 8 hours ago [-]
My guess is users, they must have had a MASSIVE user growth since they have launched to justify this
bix6 1 days ago [-]
They got money and it needs to be put to work!
besterman23 1 days ago [-]
Probably the “nobody ever got fired for buying IBM” effect. If there’s a use case for the tech, buying the most well known implementation of it will always be useful for people who want credit without the threat of blame. This funding is based entirely on the hype and a bet that TypeSafe will have name recognition.
johnfn 1 days ago [-]
Isn't this the exact opposite? When people were saying that saying, the connotation was that IBM was an old, stodgy company that had been around forever. (These days I often think "No one got fired for choosing AWS"). Typesafe is a hot new startup that could, to my eyes, easily burst into flame or die in the next year.
besterman23 1 days ago [-]
I’m saying the bet is in them becoming THE System One Model company.
Onavo 1 days ago [-]
You are used to dealing with companies where the money bags hold the power.
When you are in the middle of a boom cycle, it's the hottest company that has the advantage. Investing in them is a matter of privilege and they get to pick and choose.
Also, Canadian VCs are bottom of the barrel as far as VCs go.
kingcauchy 1 days ago [-]
It might be the people that are being acquired too, at huge inflated ai researchers salaries.
Acquired in the vc sense… not literal exit.
InsideOutSanta 1 days ago [-]
> there is shit ton of due diligence at the tech level
Maybe in some cases. But counterexample, courtesy of The Information:
"It took just 15 minutes for Blue Owl executives to agree to invest up to $10 billion in future projects alongside real estate firm Primary Digital Infrastructure during their first in-person meeting two years ago, said Primary chief investment officer Bill Stein."
AI seems to make some people lose their damned minds.
HarHarVeryFunny 1 days ago [-]
A very major part of Jev is the cost and speed. Yes, classification is/will be a commodity business, just like LLMs are, and similarly there is no moat only production cost and pricing.
Yes, anyone can wrap a decisions API around an LLM, but so what? If you want to compete then you need to compete on price, and it's not clear if OpenAI and/or Anthropic are able or willing to do that without building a custom architecture, and even then is a race to the bottom on pricing really what they want to pursue?
I'm not sure if OpenAI have announced pricing for their Decisions API, but they have said it's based on Luna which costs $0.10/M input, not even remotely competitive with Jev's $0.04/M input, which I'd expect has some headroom built into it.
Assuming that the architecture behind Jev is not just an LLM, and gives them some inherent efficiency/cost and speed advantage, then the question is whether OpenAI and Anthropic really want to duplicate this and have a race to the bottom on pricing for what may be a large part of the business automation market they are addressing. Is that what they want as their IPO pitch - we're selling potatoes, and think can grow them cheaper than Typesafe ?
phren0logy 1 days ago [-]
The API was duplicated, the results were not.
JamesSwift 18 hours ago [-]
I mean, a large majority of "pro jev" posts I see purely focus on the speed/cost and take their word on the accuracy. The entire product is "we guarantee the response will fit this schema, and will do it fast and cheap". Personally, I dont see how anything in this class can be "production ready" as a general classifier without the ability to fine tune on top.
dvrp 1 days ago [-]
That is not how it works in the US.
cmrdporcupine 1 days ago [-]
Yeah your problem and my problem and others around here is the word you just said there... "Toronto." Canadian investors are risk averse as hell. And cheap. They can make more money helping sell bitumen or real estate, why bother with arcane tech?
And if you could put the words "Bay Area" or "Stanford" or "San Francisco" next to your name... different story.
The VCs are not buying the idea or the tech, they're investing in the people. And they invest in a formula that has already worked for them before to make big coin. Prop somebody up, let them hire like crazy, and then get them get acquired, and then cash out. They don't care if it fails if they can make it succeed 1/200 times.
Canadian investors want you to have already succeeded before they help you succeed a tiny bit more.
dist-epoch 1 days ago [-]
What was the moat of Dropbox? Of Instagram? Of GitHub? Or of countless other very successful startups when they started?
Anyone know when this "have no moat" meme appeared? Even 5 years ago I don't remember seeing it on every post.
hirako2000 1 days ago [-]
Whether a product has network effects makes a difference.
I would argue Dropbox did have a moat. It didn't merely store your data. It made it possible to make backup efficiently when bandwidth wasn't all that good.
Reading "no moat" so often is also tied to the fact those companies happen to be getting surreal valuations, at a quite early stage, showing no profit, building a tech that doesn't seem difficult to reproduce.
JamesSwift 18 hours ago [-]
Dropbox: is I have to move my entire collection and setup the app on all my devices
Instagram: I have to move all my posts and also convince all my network to move over
Github: not a terrible example
If I decided to move off of jev tomorrow it would be an api key and a base api path update. Maybe 30 seconds of work.
dist-epoch 12 hours ago [-]
I meant more what was their moat when they launched, compared to their clones.
You can say the same thing about OpenAI/Anthropic, just an endpoint update, maybe a harness change if you use the CLI.
People have in fact been saying "what is the moat of OpenAI/Anthropic", yet here we are, trillion dollars valuation.
JamesSwift 5 hours ago [-]
Right, and yet the openai/anthropic models are still a tier above other options despite everyone gunning for them for at least one if not multiple years. While there are numerous jev-likes being spun up _in less than a month_ that are performing (as far as I can tell) up to and sometimes exceeding jev's accuracy.
This isnt the same as dropbox putting in a bunch of work to make the UX "just work", or even stripe doing the same for payments. This is "send me an api call that is in 1 of 3 forms, and get a response back in this form". The heavy lifting is in making that fast and cheap, and because they havent demonstrated a secret sauce in making it hard to outperform actual results that leaves a lot of room for competitors to get their existing models to run faster and cheaper for the same use case.
uuue 22 hours ago [-]
Instagram clearly did have one - Zuck had to acquire after internal efforts failed.
IshKebab 1 days ago [-]
Dropbox had a super smooth UX that somehow nobody else replicated (seriously Google wtf). Instagram and Github won on network effects.
smrtinsert 23 hours ago [-]
When you frame it as if we're in the Pets.com era of AI the continued gold rush makes sense.
slopinthebag 1 days ago [-]
brb gonna wrap claude with a new form of prompting and raise 500 mil
dvt 1 days ago [-]
Is Jev being astroturfed on HN? It certainly feels like it. It's a middling product with virtually no moat (but great marketing).
denverllc 1 days ago [-]
Yes, it's astroturfed everywhere (like X and reddit).
Jev is used as an example of a successful marketing launch where they worked with many X "creators" prior to its release, so that all the creators would repost to put it to the top of everyone's feed. Then, over the following days they'd repost so it maintained momentum.
See: doomers.ai, clickstrike, growth matrix, etc. They use coordinated engagement, paid influencer networks, customized messaging, etc.
Jev isn't a terrible product, but it's way overhyped.
amarcheschi 10 hours ago [-]
Just a question, are these creators paid to sponsor these products? Because at least in eu, they would need to add a "#adv" in their post, which makes it easy for me to spot sponsored content here. I guess in the rest of the world it may not be necessary
StrLght 9 hours ago [-]
It's mentioned in video, here's an excerpt from FAQ [0]:
> Do creators disclose that a post is paid?
> Yes, every time. Every creator post we place carries X's paid partnership label. That is an FTC requirement for any post where a brand has paid or given something of value, and it is X's own rule for sponsored content.
Although, there are other services, that sound shady.
Oh this is interesting. I've seen it was added in the last months on x but it's cool nonetheless. Otoh, i don't miss not using a lot of social networks except for a few niche subreddits
Cool, yeah my spidey sense was definitely going off. It’s a half-decent idea, but it does appear that marketing is all you need.
Kind of sad seeing so many “tech people” (including here on HN) falling for it hook, line, and sinker.
kanzure 8 hours ago [-]
> Kind of sad seeing so many “tech people” (including here on HN) falling for it hook, line, and sinker.
Unfortunately it is not sufficient for someone to think "I am immune to marketing" to actually be immune against marketing.
Analemma_ 5 hours ago [-]
"People who think they're immune to marketing are the biggest suckers of all" always been bad, but it's worse than ever because a bunch of people have decided that the most reliable way to judge AI performance is its vibes on X The Everything App. I can't even really fault places like doomers.ai when their marks are practically asking to get conned.
hall0ween 17 hours ago [-]
I appreciate you putting words to the astro-turfing. two areas llms have helped me are in (i) giving me work and (ii) learning how people in my life repeat what they hear from “knowledgeable” people. And not saying the voices proselytizing llms aren’t knowledgeable but I bet there’s an agenda behind their words (ie “i have something powerful” and “give me money!)
amarcheschi 9 hours ago [-]
Jesus f christ that might be why I saw it for a few days here on hn together with collateral discussions and topics
denverllc 18 hours ago [-]
[flagged]
minimaxir 1 days ago [-]
There have been a lot of posts/comments claiming "Jev-like models" but that's more of an shorthand for decision models, not astroturfing.
softwaredoug 23 hours ago [-]
Why is it a middling product? Most clones don’t approach its performance, and it solves a specific problem well in a way that was awkward and ignored by most frontier labs.
dvt 22 hours ago [-]
Because Jev is basically a "generalized classifier" which.. doesn't quite make much sense. Training a classifier is pretty easy and has been done routinely for like two decades now. Classifiers also tend to be very localized; for example, I've worked on classifiers that would bucket web traffic into "potential buyer" or "potential seller"—but this was very specific to the use case (vehicles, in our case).
If I seriously needed a classifier, I would just train my own and it would run on an iPhone. Any CTO worth their salt would suggest the same, because it's not even remotely comparable to training a large language model (w.r.t. compute or training data required).
9 hours ago [-]
vanuatu 22 hours ago [-]
distribution is a moat
dovin 1 days ago [-]
Jev does seem to have become the Kleenex of decision models. Is brand recognition worth $7.5B? There are lots of other decision models out there that perform at or near jev-level (laya, gliner 2.5 decide, even embedding gemma 2) that you can also run locally, and honestly I think this kind of model makes the most sense running locally as well. Maybe if TypeSafe can ship fast they can stay the default. Guess we'll find out.
pickle-wizard 1 days ago [-]
I just started experimenting with the decision models. I spun up Laya on a VM with a couple of vCPU and 6GB of RAM. I get the results in about half a second. No need for GPUs or tons of memory.
I am integrating it into the product I am building and to me it doesn't seem like there is much need to go with a SaaS for this since the requirements are so light. I just can run it in Cloud Run and get all of the scale I'll ever need, and I get to tell my customers their data never leaves my environment.
maherbeg 1 days ago [-]
Has anyone actually eval'd the other open source options against Jev on real world tasks rather than looking at benchmarks?
I see a lot of people parroting the quick open source alternatives as being better on the benchmarks, but it's such a new category that I'm not convinced we have solid benchmarks.
I'm hoping a company releases an internal eval benchmark for these options. I'm sure some of the open source ones are solid in some cases, but would love to see more reliable data.
This is awesome! Please submit updates when a new model comes out that people are saying are better than Jev!
jacobgold 1 days ago [-]
Whether they can compete on decision models or not, TypeSafe showed that a lot of the market had missed something important. With this much money, they have a lot more chances to discover other important things that are missing.
r4indeer 7 hours ago [-]
Not a native English speaker; am I missing some kind of wordplay here or is the "I" in the wrong place in the title? It says: "TypeSafe A raises Series AI", but shouldn't it be "TypeSafe AI raises Series A"?
minimaxa 22 hours ago [-]
No one finds it strange that every defector from Opensi gets half a billion dollars in funding and already has a third of Fortune 500 companies paying for their product in a few weeks with a product that's already been replicated 10 times very easily?
OK, I'm jealous... Lol
pu_pe 5 hours ago [-]
This might be the pets.com moment of this era.
jghn 1 days ago [-]
Every time I see these headlines I wonder why the Scala company is back in the news
stymaar 5 hours ago [-]
Multi-hundred million raised at a multi-billion valuation for a series A is literally insane.
How are the investors hoping to get their money back? With a $20B acqui-hire from X.ai later this year or what's the plan?
redoxate 1 days ago [-]
What edge do they have over the market to justify such evaluation
simonw 1 days ago [-]
Probably their team. Ex-OpenAI people who have already proven they can ship and get buzz for what they're doing.
miguelacevedo 1 days ago [-]
Same thing was said about Character AI, Cohere, etc. Companies started by the authors of the "Attention is all you need" paper. Look how that went... the team gets thrown around a lot as if it's the magic bullet. There is no magic bullet.
simonw 1 days ago [-]
The ex-OpenAI people who started Anthropic are doing pretty great right now.
VC is a hits business. Just one hit pays for 9 that didn't work out.
denverllc 1 days ago [-]
Anthropic's Series A was $124 million; Typesafe's is 7x that. If Jev becomes a $2bn company it will be a failure to investors.
simonw 24 hours ago [-]
Anthropic raised their Series A a year and a half before ChatGPT had been released, when LLMs were still unproven technology and most VCs weren't paying attention to the field (they were mostly still chasing crypto).
majormajor 18 hours ago [-]
This is a dangerous way of looking at things in a hits business.
The difference in hit rate matters enormously and the existence of a hit tells you nothing about the denominator.
The size of the bet matters too - you could wipe out a hit with one bad pick if that bad pick is big enough!
If you don't have an estimate of the rate that you trust, you're just throwing money at dreams.
But it's WAY harder to get rich by being a pessimist than by being an optimist. And if you're a VC choosing where to invest primarily-other people's money, then you have no particular reason to try to talk them out of the hype.
gavinray 1 days ago [-]
Upvotes and Twitter hype
babelfish 1 days ago [-]
100mm arr
stickfigure 1 days ago [-]
Remember when reaching $1B valuation made you an exotic "unicorn"?
esafak 1 days ago [-]
The needle has moved. This is a wonderful thing.
shimman 20 hours ago [-]
With any luck it'll radicalize a new generation of workers to rise up and stop SV.
atleastoptimal 23 hours ago [-]
I think this is a hedge against major AI regulation.
Invariably near-AGI systems created by OpenAI/Anthropic will be very destabalizing. In the end the world will probably regulate AI capable of [any] <-> [any] input/output types. Models will need to be limited on their outputs by law so they cannot have unbounded, unpredictable outcomes. Jev is the ideal version of "benefits of AI without making humans obsolete" that might be the consensus once the track superhuman AI and its consequences are clear.
woadwarrior01 1 days ago [-]
I wonder how much of it has been earmarked for astroturfing on HN and X? :D
Hopefully some of the Typesafe hype spills over this way to Jeffy[0]
Open source[1]. It comes with 68 pre trained classifiers which run and train on CPU alone. They run locally, are faster than Jev/Laya/Decisions, and they can perform many different tasks; from labeling email, all the way to playing Doom[2]
There is significant congregation of investment towards companies like this whereas many other with meaningful innovation are struggling to survive. The model is not a revolution like I must say how llm did at the face value. Yes, transformers were there for long but the use case etc. If you do cost comparison using regular model and setting it with prompt and structured output you can get result very close. This is what most open Jev and other are doing with tuning llm.
jypepin 1 days ago [-]
Their headline says "TypeSafe A raises series AI". Is that AI slope?
throw03172019 1 days ago [-]
They also say they are a fun team. Maybe it’s a pun.
With everything you must have going on right now, I love that you’ve come on to explain your headline is an intentional joke. All sorts being thrown around in this thread, but this was definitely the important thing to clarify.
For what it’s worth, however it works out, my guess is that the primitive Jev provides is likely to be considered essential in the future development of software.
stephen_cagle 1 days ago [-]
Is it a pun or something? I have been called dense.
I actually resized my browser thinking maybe something weird was going on with flex-wrap or overflow or whatever it is.
fwlr 9 hours ago [-]
Others have capably pointed out why their technology cannot be much of this valuation. I submit that “the ability to grab attention in the AI room” is what is being valued this highly.
-HIMAO- 8 hours ago [-]
This is pretty petty but they got the title wrong (should've been 'Typesafe AI Raises Series A' not 'Typesafe A Raises Series AI')
sixhobbits 8 hours ago [-]
Definitely intentional, they're fun
xxmarkuski 12 hours ago [-]
Their job opening as Safety Javeroni was a good laugh too [0]
If you ask the right question, you don't need a decision model.
If you don't ask the right question, you might as well consult a psychic.
If all you need is scale, I guess decision models are cheaper than psychics.
999900000999 9 hours ago [-]
What a neat API.
I went ahead and applied for a job, and unlike Oxide, TypeSafe has an application that only takes a few minutes to fill out.
I did take a second to play with the TypeSafe API first. I'm also excited to get Jev running locally.
Not like I'm realistically good enough for a top AI company, but I can dream. I got multi code signal tests for Anthropic and got 70% at best.
Would be nice to get equity, ipo and cash out.
I don't want to work in my 50s
Culonavirus 12 hours ago [-]
Oracle needs to hurry up and get downgraded to junk status at which point it will raise shit and at which point it will finally send off the avalanche of normalization across this industry.
There needs to be cleansing with fire. Weeds need to die. Trees need their branches cut. The sooner the better.
A shitty ass random "ai" startup built entirely on hype and astroturfing should not be raising anywhere near this amount of money.
pythonRon 11 hours ago [-]
If Typesafe really wants to do some good in the world, it'll raise money for those in need.
rokhayakebe 1 days ago [-]
Can someone who actually knows these things share how might a company like this spend $870M over the years?
We've tested the decisions API against Jev at work and it's worse in various dimensions. Costs more, higher error rates, slower, and the answers are worse.
A lot of people are shouting about how Jev hasn't actually differentiated itself, but I question how much folks are actually experimenting with what's out there before coming up with an opinion.
For us, it's cleae that OpenAI rushed this out to meet the hype in the market right now without having a product that actually meets the bar Jev has set.
denverllc 1 days ago [-]
> but I question how much folks are actually experimenting with what's out there
I did. Originally I had a project that I had been wanting to do and thought to use a decision model for it. Jev, OpenAI, etc. are all within percentage points of each other.
Then I used traditional ML and found a small classifier (gemma 4) with traditional embeddings worked 2x as well.
Jev is the general purpose ML pipeline for when you want average results. Nearly every application has a "better" option available with a small amount of work.
benatkin 1 days ago [-]
They weren't running on OpenAI, so nope.
softwaredoug 24 hours ago [-]
What’s interesting about the Jev moment isn’t just Jev, it’s the unleashing of distillation / fine tuning outside the frontier-adjacent labs. It’s the sudden explosion of a million Jevs.
If being an “AI Researcher” is a ticket to multimillion dollar salary, AI training talent cannot be contained to a handful of companies. It’ll become more common and diffuse. The old advice of not fine tuning, because it’s hard, goes out the window as that knowledge diffuses through the industry.
A similar thing is happening in search. For a long time labs have trained tailored embedding models. And now companies like SID training their own agentic models that are smaller and faster at search than GPT-5.
_davide_ 13 hours ago [-]
I guess they are investing in the team rather than on their demo product
2001zhaozhao 14 hours ago [-]
Tbf this is THE way to get cost-effective AI into video games (think enemy AIs and NPC AIs that are much better than current ones and runnable on consumer GPUs), i'm really looking forward for decision models maturing.
greesil 19 hours ago [-]
The amount of investment being thrown at this when we have serious existential problems as a species is how I know the system is definitively broken. Capitalism got us this far, but it's time for something new.
maxdo 1 days ago [-]
interestingly , it destroy the landscape of Chinese models.
Unless china takes leadership in frontier space the picture is next :
1. cheap workhorses for classification, routing, other scenarios : Jev
2. coding agents with less erros : Anthropic/Openai, etc.
3. Science /Legal/Medical : A mixture of Jev+Anthropic scenarios
kylehotchkiss 24 hours ago [-]
See clef
dmix 23 hours ago [-]
That press release made me cringe a bit. Maybe corporate speak wasn't so bad after all.
collimarco 1 days ago [-]
This looks like the top of the dot-com bubble...
bel8 1 days ago [-]
Do they have patents over jev related tech or something valuable to justify this?
enahs-sf 1 days ago [-]
got the recruiter call only to essentially be summarily rejected because my pedigree is wack. looks like i would've gotten hosed on valuation anyways.
mococa 1 days ago [-]
~7B for a thing that we already have open-source?
kevinkatzke 1 days ago [-]
A few year ago you could IPO at this valuation.
Amekedl 9 hours ago [-]
This makes no sense, unless they are going to crack something, like the actual atrocious accuracy of around 70%.
It's always been the last 30% where it just fails.
Guess it is 870M to 0
thatsadude 10 hours ago [-]
These investors are not that smart
robotswantdata 1 days ago [-]
I’ll just use open source, thanks
jdw64 15 hours ago [-]
I only use Jev at the tool-calling layer to keep tool costs low, but I don't really see what the big advantage is.
ada1981 18 hours ago [-]
I built JEV-search and JEV-chat and have been really happy with the results…
Curious to see how all of these can work together, and giving LLMs the ability to deploy their own workflows and built type 1 systems as they go has also been fun and useful.
alpineman 14 hours ago [-]
From Andreessen the Trump donor. Sad, will switch to something else now.
axus 1 days ago [-]
So... are they worth more, or less than 0xide
TeMPOraL 1 days ago [-]
Only just now I realized that "TypeSafe AI" are the people behind Jev, and "System One" isn't the company name, as I assumed, but a larger project label.
sajithdilshan 1 days ago [-]
Is there a second wave of AI bubble happening? How can an AI classifier company be worth of $7.5B
fHr 4 hours ago [-]
this market lol
gdiamos 23 hours ago [-]
Good job typesafe.
You won the competition with VCs
Thorentis 1 days ago [-]
Another data point confirming that AI is a bubble.
villgax 14 hours ago [-]
Lol
pseudotensor 1 days ago [-]
Disclosure: I work at H2O.ai.
We released an Apache-2.0, open-weight 4B decision model that scores above Jev 1.13 on JevBench's composite score (72.5 vs 71.5) and is currently the top open model there: https://benchmarkheaven.com/jev-models . Newer models coming even larger than beat Jev in intelligence as well.
- Same contract as Jev: state + typed questions in, calibrated probabilities out, one forward pass, no generated tokens.
- Your data never leaves your environment, and there's no per-call fee.
One note on their speed comparison: the JevBench board shows self-hosted models with an adjusted latency of "2x + 0.15 s (assumption, not measured)". Our measured p50 is 29 ms; the adjusted figure is 0.21 s. They report 85 ms for theirs.
EDIT: As I wrote this Microsoft just released their own Decision-1 model [3].
[1] https://developers.openai.com/api/docs/guides/decisions
[2] https://unsloth.ai/docs/basics/train-your-own-decision-model...
[3] https://commandline.microsoft.com/microsoft-decision-1-model...
There's the potential for an inverse LLM play. In contrast with LLMs, all that seems to matter is the model, and the products the labs build around the models are all really samey and boring; the same left panel list of agents, main view agent conversation, right hand extra context, and we're now in the era of everyone creating the same cutesey furry friend on top of all this tech.
This is a very important insight. And it applies to LLMs as well. Very few people were impressed with the capabilities of GPT 3, it was mostly a techie novelty
But then when they added chat on top of gpt 3.5, all of a sudden it was a huge hit. Sure there were improvements in the model from 3 to 3.5, but the biggest impact was from the chat experience
Conversely, when they created Eliza, a basic chatbot more than 50 years ago, people even got addicted to it, despite it’s ai model being something super rudimentary and basic compared to what we have now. The model capabilities didn’t matter as much as the experience the chat created
“The following is a conversation between a human and a helpful AI assistant.
Human: [prompt]\n AI:”
You would get consistent conversational chat, with the main limit being the model’s limited context window, and obviously frontier intelligence at the time.
I played around quite a bit with davinci-003 and conversational systems before ChatGPT. I kept using this format and API for a while, but at least during the ‘free research preview’ era, found it generally smarter than chatgpt. 3.5-turbo was probably the watershed moment where I moved away from completions on base models; to chat completions.
If you’d like to emulate this experience, spin up a base (not instruction tuned) checkpoint of Llama 1; or a more recent base model. You might underestimate how much ‘intelligence’ you can get :)
Back then the API exposed a lot of controls from sampling to logits, so you could specify a custom stop token (some rarely used Unicode); then switch to near-greedy sampling for more reliable “function calls”, etc; and then switch back to sampling for chat.
In the early days, before releasing something with the API, you had to get your use case approved by OpenAI, like the App Store review process.
Pure davinci was easily confused by what was system prompt, user prompt and its own answers.
I agree that the potential is enormous, but I'm not sure why it would be Typesafe who figures out the best way to develop and package this capability. They have a bit of an advantage in the amount of time they have been thinking about this problem specifically, but I think it's a big leap to say that they will be able to dominate the category.
Which is why they are valued at "only" 7.5B, while frontier labs with general-purpose LLMs have valuations in the order of trillions.
Horowitz et al. invest in probabilities, not certainty. If Typesafe doesn't make it, they'll have other horses in the race too.
I suppose the team and the surprise warrant that.
but anyone can do that, that's not their product, Jev is
Why?
The sum of people not even close to the topic has risen, according to my feels at least.
Has this rang true for anyone else?
the claim is extraordinary, yet the answers cover only the mundane.
I wonder how many of these posts are astroturfed.
For decision-making with neural networks, we instead need the older idea of estimators of the predictive uncertainty with constraints in the feature-representation space (over training/support of the estimator), as with Similarity-Distance-Magnitude estimators: https://pypi.org/project/reexpress-sdm/
Like just tuning this model can make it cheaper for voice AI providers to do semantic turn detection over the space of responses.
I haven't personally yet found a net-new-capability bigger than that from LLMs here.
There are a lot of those types of problems in the world, replacing what was previously spotty hardcoded decision logic that worked 70% of the time with something that now gets it right 95% of the time. Even if not perfect, the fact that you can dramatically close the gap means a lot lower cost overall, so it’s worth it.
These guys also love text to sql. I know 4 people each implementing text to sql for their separate companies. Or text-to-redash dashboard in some cases. Funnily enough, text to sql is one of the things where LLMs are clearly sub-human. Mostly due to lack of easy verifiability.
Text to SQL is a really bad idea, as the agent will just generate really convincing and terrible analyses that lack all business context.
I mean, you could implement some kind of semantic layer, but that's a lot of work, for uncertain reward.
I saw a place which essentially tasked the data analysts/scientists in an area to describe the good tables, and assumptions and necessary context to make decent queries/analyses, and that definitely lead to much, much better (almost usable) results.
As I have said many times, what the world really needs is SQL2Text, rather than Text2SQL.
Yep. After this has been done, the later text to sql outputs works half decent. The major issue I heard was that whenever a small assumption changes few months down the line, while humans are able to "JIT" apply it to each of their queries, these models sometimes lose it in all of the context and so it becomes hard to progressively update anything - you end up having to redo the whole workflow.
And in my experience, even the more complex dashboards requested by product or upper mgmt were stood up within the day by humans. Or the next day (but that was cos some ETL would be scheduled overnight rather than immediately). And I am not sure if speeding up this part would even matter to the guy who requested the dashboard. So I am not sure if this is a major bottleneck people should try to innovate in. If openai does the work for you then great but otherwise. Humans aided by copilot-of-old style autocomplete is probably more than enough.
Needless to say, text to sql is pointless for application queries made during API calls.
The whole text to SQL thing is about data teams being a bottleneck to the business (rather like software people). The goal is good in theory but honestly an Excel plugin is probably a better solution in many cases.
They really haven't.
What I haven’t seen answered is why it isn’t reproducible within very short amount of time and low effort.
Sure, the large labs are the lazy incumbents at this stage. But any other startups would be able to copy the experience. What’s the barrier of entry that I am missing?
But even if others surpass them and make better solutions, the fact that nobody was able to see it before typesafe is a testament of what they might be able to come up with next.
I sound like a fanboy but I swear I 'm unaffiliated with typesafe. I was building my own version of this way before they announced JEV (mine was ALE and it was mentioned here on HN for a bit), in use for VR gaming (so one can give commands to NPCs with voice and supports multiple commands in sequence in a single pass), but I missed the "killer usecase" of being a new primitive, like everyone else.
TLDR: There's a lot of value in thinking ahead and seeing the future. The clones are nice and exciting but they give me "I could have built this first, yes but you didn't" vibes. I hope they manage to keep it up and push the space forward again
Counterpoint: there are a lot of one-hit wonders, and they vastly outnumber the idea-factory people. This is not to minimize those people, a single idea can be very successful (see Zuckerberg), but it doesn't mean your subsequent ideas will also be great (see Zuckerberg)
Here if anyone is interested: https://pantel.is/projects/ai-gaming-companion/
Yes, this generates probabilities over a set of given options instead of the whole token vocabulary.
I don't see what's so exciting about it, compared to regular LLMs which support structured output options in the API.
Models like JEV fall in the golden middle. Somewhat cheap, somewhat fast, somewhat general.
Many people have replied saying that narrow classifiers (eg bert) are better, and they are better in many ways (essentially free and faster). But... the real world is messy, in my experience it's actually very hard to make a really good classifier that will work across a specific niche domain especially when there is very fuzzy input. And the most interesting real world scenarios are nuanced and overtime there will be edge cases discovered where it fails and that will require retraining and re-evaluation etc etc. It can becomes a full project on its own that eventually collapses into a game of whackamole (improved in some direction but regressed elsewhere). Things like JEV (or similar models) provide ease of mind, just off load the complexity to it and move on to the next challenge type of thing.
I keep evaluating Jev for my product because of the hype, but the reality is that for my tasks, Luna is more accurate, only a little more expensive, and the latency doesn’t matter. I’d rather spend the extra $50 / month or whatever than have to shoehorn in another API and provider, and also lose the ability to change reasoning level and get reasoning summaries for eval purposes.
But that’s just for my product.
This all sounds intelligent and likely, and yet we can come up with countless counter examples where the first mover is not the big winner, and nobody cares about who did it first. There is typically way more value in nailing the execution of a big idea someone else came up with, rather than "seeing the future".
Exactly.
If JEV is here to stay, we can have another split along the "one LLM rule them all" regime
name it worker or implementator or something
- mostly for agent with with spec (not directly for human, as we tend be handwave with vague intent) - follows the instruction verbatim - lives in sandbox by default - basically sol6; but WITHOUT magical situational awareness, lets work in group BS
Precisely this. Should be top comment.
Also, the model moat is understated as training data for these purposes also accrues to the winner, which due to the first mover advantage as well as the distribution advantage you speak of, is typesafe. In contrast to relatively open coding data. Openai anthropic also have that, but like you say its a different business.
If the word “calibration” in probability doesn’t mean anything to you, the difference isn’t very apparent, but that doesn’t mean it doesn’t exist.
I am really sceptical that they've developed a zero shot calibrated classifier for arbitrary inputs.
That being said, if they have, then this valuation is like 100 times too low.
https://seldon-ai.com/blog/fronter-llms-are-semantic-interpr...
Constraints can be an advantage in many systems. Consider typed and untyped programming languages. Many advantages with typed languages over the latter in terms of development and efficiency.
I give it to them for creating the hype (good marketing), and for making a useful classifier. Not sure what they would scale rapidly though.
People who can find things of value, and execute on those things, are worth investing in - even if that thing of value is replicated in short order.
Nobody was looking at decision models, but now that Jev has surfaced, everybody is.
Investing isn't binary. You can be investible but not at a 7.5b valuation.
Jev has been around for a couple weeks. Cost and performance matter more than anything. Staying with an existing provider (the # of people choosing Jev without already having a frontier API key is probably zero?) is way easier than this.
What the hell are we even talking about.
Actually they stolen the idea from a paper.
*They have zero moat* against established AI companies, as the recent copies show, and will do worse than frontier labs.*
Most of the money will be burned in data centers as they struggle to sell tokens at a loss.
That’s a different set of properties in terms of size, cost, and latency. Jev (apparently, not like I’ve seen its insides) is extremely cheap, extremely fast, doesn’t cost any output tokens as it speaks the output natively, can’t get the output wrong because it speaks the format natively, and (presumably based on the docs), the context is separate from the question, meaning it should be immune (or at least highly resistant) to prompt injection attacks.
It’s not just about the accuracy of the result, it’s a collection of all the properties that make Jev interesting.
Jev took years to develop, I strongly doubt that a copycat that was put together within days after Jev’s release will be able to match it on a sun of its properties. Even if fine tuned LLMs can outperform it on raw accuracy.
Also Jevs purpose isnt to become its own thing. It will get aquired in 18 months by one of Andressen Horowitz's incestuous circle of companies and everyone will make money, and the person who buys it wont necessarily care if Jev itself makes them a ton of money. They're just passing chips around the table.
and https://arxiv.org/abs/2510.01237
"SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization"
Jev is a general-purpose thing. That is a specific-purpose thing. General-purpose thing is not the same as specific-purpose thing. What makes people think these are the same thing? I don't get it.
Copypasta:
• A reinforcement learning architecture specifically designed for sales conversation analysis and conversion prediction
• A synthetic data generation pipeline leveraging GPT-4O to create diverse and realistic sales conversations
• Novel state representation techniques using Azure OpenAI embeddings (3072 dimensions) with sales-specific features
• A meta-learning approach enabling the system to express confidence in its predictions based on conversation similarity to training data
• Integration mechanisms providing real-time guidance within existing sales platforms
• Extensive comparative evaluation demonstrating significant performance improvements over LLM-based approaches
Jev is basically the embeddings side of an LLM. Yes, it's a good idea, but the moat is non-existent.
> Jev is basically the embeddings side of an LLM
This is how the paper you are referencing is describing the "key contribution" as it relates to embeddings:
> Novel state representation techniques using Azure OpenAI embeddings (3072 dimensions) with sales-specific features
And now you are claiming that the author of the paper created:
> "generalized" version of same
It's a bit hard to guess what you are trying to say, but if I were to steelman your argument, I would guess that you mean: training a model with a vocabulary that has some tokens representing things like "OPTION_A", "OPTION_B" is in your mind the same thing as "generalized version of sales-specific features in embeddings"? Is this what you were trying to say?
Literally misses the point of Jev, which you don't need to fine-tune to get accuracy nor - and no other model has this - some sort of out of sample calibration
But why? If the simplest way to achieve Jev's capabilities (accuracy, cost, latency) is by fine tuning a small model, what makes you think that this isn't exactly what Typesafe did?
And even if they did something different - what makes you think it was a good idea in the first place, given how easy their results were replicated without any "secret sauce"?
Plenty has been said about this claim. If you're still falling for this, I feel sorry for you.
If you remove wheels from your car, your car will get a "no speeding ticket" property, and yet there's nothing exciting about it.
> You can’t get that with a fine tuned LLM
Of course you can. All these claims are nothing but marketing.
The only difference I am aware of is that probabilities are better calibrated with these decision models compared to regular LLMs which can output hallucinated numbers where your schema allows a number.
LLMs can produce structured output by limiting the next token based on a grammar, that ensures correctly formed output.
But (presumably, again I haven’t seen Jev’s insides) Jev doesn’t need to do that, it just has to output probabilities for each answer, and can do it natively without artificially limiting output tokens.
So jev is cheap, fast, never produces malformed output, its output always carries correct semantic meaning (but this doesn’t mean it always answers correctly), and it separates correct from prompt (which if it truly does that is the most exciting part).
The jev documentation says this is what you should do. It doesn’t mean that it’s actually designed for prompt injection resistance, but it could be. We won’t know for sure until they either release more details or someone proves otherwise.
But the split is one that doesn’t exist in normal LLMs and it’s the exact split that is needed for immunity or resistance to prompt injection.
Have you heard that from a different source than OpenAI? From what I'd heard other models haven't gotten close, and the open source ones are like running gemma4 E2B against Opus 5.5- sure, the API calls go in and are returned the same but the quality isn't close.
Fortunately Jev is cheap, so I dont think it matters too much, but I think its robbing people of the opportunity to learn and implement this themselves.
Also, I dont really want 3 companies responsible for censorship/classification.
> Fictitious capital could be defined as a capitalisation on property ownership. Such ownership is real and legally enforced, as are the profits made from it, but the capital involved is fictitious; it is "money that is thrown into circulation as capital without any material basis in commodities or productive activivity".
https://en.wikipedia.org/wiki/Fictitious_capital
no, you can't, and it's unclear why you would think this.
*if you have a sufficiently sized and quality dataset for the specific classifications you're targeting
Then yes, easy!
Here's a hint: confidence is not generated by a model.
But confidence value is just a function applied to probabilities. It is not coming from the model, and it carries no additional information.
It is documented btw, and yet you will see plenty of claims that Jev is better than LLM because it returns both.
Do you think though that Luna being a model post-trained for chat produces over-confidence in logprobs?
Typesafe claims that Jev is calibrated, but there are plenty of examples where it completely fails (predicting die roll being the most obvious one).
Unfortunately calibration is hard to benchmark.
If you instead give it a list of probability for each number and ask it whats the probability of each number, the result will be accurate.
Did i hear that correctly? In order for Jev to be accurate you have to give it the answer before asking for the answer?
(btw this is exactly how Jev is playing games).
What am I missing?
As in, neither LLMs nor Jev are truth engines. Truth comes from the provided context.. Plus weights.. kind of fuzzy, sorry I’m thinking out “loud”
This is not a “system one” question.
You’d be much better off using a proper LLM here: more accurate, can provide justification, can be steered when mistakes are made. Jev is going to be a coin flipper with almost no knobs.
So they clearly have a product, a strategy around it and perhaps the compliance scaffolding (SOC2 Type II etc) that may be needed before actually being able to charge money for it. They also have the right brand names associated with the founding team. Execution, so far, seems good enough to create a splash, at least.
As an investor, the question(s) to ask is (in my view): "How do they make money? Will that way to make money survive?". The answer to the first: selling input tokens and perhaps subscriptions/credits eventually.The answer to the second: "Yes, but with the risk of unit revenues declining faster than their unit costs". How can they mitigate this problem: by being big (scale / mindshare etc) so that their unit costs (including for customer acq) fall faster than their unit revenues will - I believe that is the question most AI companies are trying to tackle these days. Any new competitor will have to tackle basic fixed costs (of getting started) first before even getting to the stage of having the luxury of worrying about unit-economics.
So yes, they might eventually be competed away but whoever is in their team is trying hard to make a useful product/ecosystem and that should be applauded, not ridiculed with "it's all marketing". This is way more than a simple github/huggingface-based open-source replica solution can hope to achieve without institutional backing (either big-tech or system-integrators).
What should rightly be questioned, of course, are the valuations the VCs are providing to them in hopes of passing this hot potato to a willing buyer (say a hardware maker like NVidia) - the incentives there are very well defined and depend very much on perceived TAM (which lately is on very shaky ground given how far token pricing has fallen causing, among other things, OpenAI to "miss" on the market's expectations for annualized revenues, even before they're listed!) [1]
[1]: https://www.ft.com/content/b66a9858-f8fb-46cb-b506-44bfe26fc...
So yeah, not too much enterprise sales experience there for sure.
* https://jeffyclassify.com/
* https://playground.jeffyclassify.com/#doom
* https://github.com/nicobrenner/jeffy
Then everyone breathlessly repeats this story about how Jev is useless because open source models they never tried claim to do the same thing and better.
But, for OpenAI this is not a primary business, for open source models as well, so they will not be chasing the market and customers to buy their product and promise them to maintain it.
TypeSafe will do all this, they will try to understand your use cases and then solve your pain point, while others are providing raw material.
such an arrangement can end up beneficial to the VC firm
Guess the (investment) market has spoken.
The Microsoft article GP linked even shows the MS model having 95ms latency.
Even on price Jev being matched (the same MS model is "Input tokens cost $0.042 USD per million tokens. Output tokens are free.", same as Jev).
Cloudflare's Clef-flash model is actually slightly cheaper: "$0.038 in / $0 out per 1M" too.
Jev is being matched or exceeded in performance and price within a month of them going public.
Jev is 42$/B but OpenAI is 100$/B token.
By all means, become an A16Z LP.
To run the SDK examples below, use these OpenAI SDK versions or later: Python 3.26.0,
I thought Pythin 3.15.0 just came out, 3.26.0 must be really far off?
They may well be a good team to throw money behind if you are hoping to bet on a new AI lab.
Not them per se, but doomers.ai [0] which is a uh... "launch virality agency".
So less that Typesafe has a strong marketing muscle, and more that they paid at least $100K [1] for "organic" buzz.
[0] - https://doomers.ai/work/typesafe-ai-case-study
[1] - https://doomers.ai/guides/ai-marketing-agency
I read the attached links and it just looks like a marketing agency?
but in this case, this is product extremely boosted for super viral marketing, and this is the first time i’ve read about this boost since hearing of the product whereas everybody only really talks of its substance. feels like a magician kept a trick going for a very long time, people forgot they were watching a magic show, then the curtains come down and they reveal the deceit after (i’m being slightly melodramatic)
We've yet to see whether this is true, or is it just manufactured demand. There are dozens of Jev demos, but pretty much all of them are either cool but useless, or simply fake (i.e. harness doing 99% of the work).
But economics of Jev enable classifying high volume things. And I've tried all the open-weight Qwen based decision models on my Spark and none of them came closer to quality and suck at batch inference so they are doing something custom.
its clear they own the mindshare around this type of primitive which is a massive premium
You can create a company with 2B shares and sell one share to your friend for $1000. Lo and behold, you own $2 trillion dollar company, leaving Elon behind.
Multiplying out revenue from the peak of their mini hype cycle while a dozen well financed competitors target them directly seems extremely optimistic.
If they weren't lying about $100M ARR, they had to be at 6.5T/day for a week to hit that number.
They also claimed they were at 1T/day just couple weeks ago. That's $15M ARR.
85% drop in just a couple of weeks?
It’s not to be unexpected for a new product and startup, but they lack SOC 2 certification, and the other things enterprises look for on information security.
I’d be surprised if many procurement divisions are approving typesafe AI, and I fear by the time they achieve it, existing vendors would’ve catched up.
https://benchmarkheaven.com/jev-models
According to this benchmark, Jev is currently trailing Quyet-1.0-Large and a few other hastily put-together LLM-based decision API-like setups.
They can be both great and over-valued at the same time.
If anyone else came up with the same concept on a Reddit thread (they have) it no-one would care without those characteristics even if you are "first".
Rebranding, execution, marketing, ex-<big_name_company> and mostly importantly, hype is what gets the investors scrambling into throwing money at you.
I have no faith in the technique if it cannot do the basics (i.e. not real probabilities, the confidence for coin flip outcomes)
tried it a couple of days ago here: https://jevplayground.com
the "not real probability" disclaimer only appears after you get a result
There's your problem. The single biggest thing every Canadian VC is trying to figure out is "why are these people asking us for money when if they were any good they'd be in the US" so by simply asking them you're already signalling something bad. A lot of their enthusiasm for process is based on this suspicion and also that the entire industry is just a way for various professional services to extract most of the investment money, since that's the game they're so used to playing with the government.
There are some Canadian VCs earnestly trying to improve but they are overwhelmingly hilariously conservative and focused on unimportant signals over reality. This is one (but not all) of the major factors that drive basically every remotely ambitious Canadian company to run a corp in Delaware and go for funding from the US. The tax situation is the other major contributor.
But that does nothing to make up for the terrible investment community. Getting started here requires already being started.
When I briefly worked for a Toronto startup, it was like all of them went to the same private boy's schools together as kids. It was a status club.
I jumped ship to an American startup and made almost double the money dealt with 0% of the bullshit and they were bought by Google the next year.
Also the situation isn't static. Investors know that the act of writing them a $870M check itself increases the chance that they'll be one of the winners, because that will attract more talent, customers, and funding to the company in a self-reinforcing cycle. And investors know that other investors know that, and that someone is going to write them that $870M check, so to some extent they're forced to think of the company as having already been successful at the fundraising and already having that momentum boost.
Only a small number of investors in the world can play the game at this level, because you have to smart enough to be right (often enough), and you have to be established enough to see the deals (be on every CEO's short list - because CEOs are only going to seriously pitch 5-10 VCs on a hot deal, if that). Otherwise you can't pull it off. Martin Casado and his team are among the few that can and I think their results reflect that.
At the same time, hopefully you would agree that this "self-reinforcing cycle" that you describe, if it really is the dominant way of thinking, is also a pretty clear sign that we're well into bubble territory, divorced from fundamentals.... Investing because you think your sheer act of investing is going to create a self-fulfilling prophecy of success while simultaneously having FOMO about someone else beating you to it? I dunno if I've seen a better description of what it looks like to make investment decisions from the POV of being deep inside a reality-distorted bubble while chasing mania.
[1] not really but they did some innovative things and popularized a concept
no, it was not
> was duplicated within a couple of days
was it already available or did it become available in a couple of days? it cant be both (neither is true, actually)
Jev is:
- accurate
- general purpose
- fast and cheap
Models we had before Jev had at most 2/3 of above qualities, but none of them were 3/3.
I don't know how good Jev's "confidence scores" are, but I would be surprised if they were in any sense better than logits from some good LLM. One advantage of Jev here is that the confidence scores are easy to access. Most LLM API providers don't provide an easy/convenient way to access the logits. But that's a minor point, you could of course build something like this with LLMs (and many people have).
You're the one being deceptive here. Jev is trading accuracy, speed, and cost for generality. It's less accurate, slower and more expensive than trained classifiers. So it's still 2 out of 3, but with decimals. Maybe 2.2 out of 3 if I'm being charitable.
And the reason we didn't have that before is because nobody thought it's a good tradeoff.
Would Jev be more accurate in a specific task if it had been developed only for that task, as opposed to general purpose? Of course it would. So, sure, Jev is trading accuracy for generality. According to you "nobody thought it's a good tradeoff", which again is false, there was huge demand for a cheap and accurate general purpose classifier.
A business doesn't need Jev for the sake of Jev. Most business are solving specific problems.
And fine-tuning got a lot cheaper these days - I've seen claims here on HN that ~500 examples is enough to beat Jev.
> "nobody thought it's a good tradeoff", which again is false, there was huge demand for a cheap and accurate general purpose classifier
There wasn't. The hope is that there was a latent demand, but we've yet to see if it's truly latent or just manufactured.
Noone is saying "hell yeah, finally we got a general purpose classifier, my business needed it so much". The typical message is "this seems cool, let me see where I can apply it".
The fact that name itself is a play on Jevons Paradox illustrates that there was no demand until Jev was released.
The appeal and claim of LLMs for businesses is solving big problems.
LLM valuations to solve tiny problems seems iffy.
--
Are there significant problem domains LLMs are bad at that Jev is good at? Vs just 'Jev can do a subset of LLM things faster/cheaper'?
Now, it could be viable if your business has literally hundreds of problems thats require classification. I just haven't seen those.
I treat the fact that almost noone was doing that as evidence that decision models aren't that useful/groundbreaking. That, and the fact that every single demo I saw was either fake (e.g. playing games), contrived, or plain wrong (e.g. using Jev for compaction).
but you truly do sound like an angry 19 year old from your arguments.
- accurate - on what? on trust me bro benchmarks?
- zero-shot model are fundamentally general purpose.
- fast and cheap ; models on hf are FREE and fast enough.
I think the key differentiator was that a team found a whitespace in what ChatGPT was doing, main comes from the same pedigree and team is as conscious of marketing as their product. SF VCs love these out of the box challengers, and people are claiming to replicate doesn't seem to matter.
The amount raised feels surprising but again entire SF/US AI scene is primarily "add moar layers and GPU" one trick ponies at this point.
"Throw a bunch of money at copycats in the trend of the day" is very normal in SF/SV for the last several decades, going back to the dotcom boom.
The definition of "a bunch" has changed but so many "crypto for X" or "recommendations for X" or "uber for X" or "social for X", etc, things raised amounts that seemed wildly divorced from their market position.
When you are in the middle of a boom cycle, it's the hottest company that has the advantage. Investing in them is a matter of privilege and they get to pick and choose.
Also, Canadian VCs are bottom of the barrel as far as VCs go.
Acquired in the vc sense… not literal exit.
Maybe in some cases. But counterexample, courtesy of The Information:
"It took just 15 minutes for Blue Owl executives to agree to invest up to $10 billion in future projects alongside real estate firm Primary Digital Infrastructure during their first in-person meeting two years ago, said Primary chief investment officer Bill Stein."
https://www.theinformation.com/articles/blue-owl-eyes-new-de...
AI seems to make some people lose their damned minds.
Yes, anyone can wrap a decisions API around an LLM, but so what? If you want to compete then you need to compete on price, and it's not clear if OpenAI and/or Anthropic are able or willing to do that without building a custom architecture, and even then is a race to the bottom on pricing really what they want to pursue?
I'm not sure if OpenAI have announced pricing for their Decisions API, but they have said it's based on Luna which costs $0.10/M input, not even remotely competitive with Jev's $0.04/M input, which I'd expect has some headroom built into it.
Assuming that the architecture behind Jev is not just an LLM, and gives them some inherent efficiency/cost and speed advantage, then the question is whether OpenAI and Anthropic really want to duplicate this and have a race to the bottom on pricing for what may be a large part of the business automation market they are addressing. Is that what they want as their IPO pitch - we're selling potatoes, and think can grow them cheaper than Typesafe ?
And if you could put the words "Bay Area" or "Stanford" or "San Francisco" next to your name... different story.
The VCs are not buying the idea or the tech, they're investing in the people. And they invest in a formula that has already worked for them before to make big coin. Prop somebody up, let them hire like crazy, and then get them get acquired, and then cash out. They don't care if it fails if they can make it succeed 1/200 times.
Canadian investors want you to have already succeeded before they help you succeed a tiny bit more.
Anyone know when this "have no moat" meme appeared? Even 5 years ago I don't remember seeing it on every post.
I would argue Dropbox did have a moat. It didn't merely store your data. It made it possible to make backup efficiently when bandwidth wasn't all that good.
Reading "no moat" so often is also tied to the fact those companies happen to be getting surreal valuations, at a quite early stage, showing no profit, building a tech that doesn't seem difficult to reproduce.
Instagram: I have to move all my posts and also convince all my network to move over
Github: not a terrible example
If I decided to move off of jev tomorrow it would be an api key and a base api path update. Maybe 30 seconds of work.
You can say the same thing about OpenAI/Anthropic, just an endpoint update, maybe a harness change if you use the CLI.
People have in fact been saying "what is the moat of OpenAI/Anthropic", yet here we are, trillion dollars valuation.
This isnt the same as dropbox putting in a bunch of work to make the UX "just work", or even stripe doing the same for payments. This is "send me an api call that is in 1 of 3 forms, and get a response back in this form". The heavy lifting is in making that fast and cheap, and because they havent demonstrated a secret sauce in making it hard to outperform actual results that leaves a lot of room for competitors to get their existing models to run faster and cheaper for the same use case.
https://www.youtube.com/watch?v=xNgQtzEl4lY
Jev is used as an example of a successful marketing launch where they worked with many X "creators" prior to its release, so that all the creators would repost to put it to the top of everyone's feed. Then, over the following days they'd repost so it maintained momentum.
See: doomers.ai, clickstrike, growth matrix, etc. They use coordinated engagement, paid influencer networks, customized messaging, etc.
Jev isn't a terrible product, but it's way overhyped.
> Do creators disclose that a post is paid?
> Yes, every time. Every creator post we place carries X's paid partnership label. That is an FTC requirement for any post where a brand has paid or given something of value, and it is X's own rule for sponsored content.
Although, there are other services, that sound shady.
[0]: https://doomers.ai/faq
Obscene marketing is all you need it seems.
Kind of sad seeing so many “tech people” (including here on HN) falling for it hook, line, and sinker.
Unfortunately it is not sufficient for someone to think "I am immune to marketing" to actually be immune against marketing.
If I seriously needed a classifier, I would just train my own and it would run on an iPhone. Any CTO worth their salt would suggest the same, because it's not even remotely comparable to training a large language model (w.r.t. compute or training data required).
I am integrating it into the product I am building and to me it doesn't seem like there is much need to go with a SaaS for this since the requirements are so light. I just can run it in Cloud Run and get all of the scale I'll ever need, and I get to tell my customers their data never leaves my environment.
I see a lot of people parroting the quick open source alternatives as being better on the benchmarks, but it's such a new category that I'm not convinced we have solid benchmarks.
I'm hoping a company releases an internal eval benchmark for these options. I'm sure some of the open source ones are solid in some cases, but would love to see more reliable data.
[1]: https://tn1ck.com/blog/jevdit
OK, I'm jealous... Lol
How are the investors hoping to get their money back? With a $20B acqui-hire from X.ai later this year or what's the plan?
VC is a hits business. Just one hit pays for 9 that didn't work out.
The difference in hit rate matters enormously and the existence of a hit tells you nothing about the denominator.
The size of the bet matters too - you could wipe out a hit with one bad pick if that bad pick is big enough!
If you don't have an estimate of the rate that you trust, you're just throwing money at dreams.
But it's WAY harder to get rich by being a pessimist than by being an optimist. And if you're a VC choosing where to invest primarily-other people's money, then you have no particular reason to try to talk them out of the hype.
Invariably near-AGI systems created by OpenAI/Anthropic will be very destabalizing. In the end the world will probably regulate AI capable of [any] <-> [any] input/output types. Models will need to be limited on their outputs by law so they cannot have unbounded, unpredictable outcomes. Jev is the ideal version of "benefits of AI without making humans obsolete" that might be the consensus once the track superhuman AI and its consequences are clear.
Open source[1]. It comes with 68 pre trained classifiers which run and train on CPU alone. They run locally, are faster than Jev/Laya/Decisions, and they can perform many different tasks; from labeling email, all the way to playing Doom[2]
[0] https://jeffyclassify.com/
[1] https://github.com/nicobrenner/jeffy
[2] https://playground.jeffyclassify.com/#doom
For what it’s worth, however it works out, my guess is that the primitive Jev provides is likely to be considered essential in the future development of software.
I actually resized my browser thinking maybe something weird was going on with flex-wrap or overflow or whatever it is.
[0] https://jobs.ashbyhq.com/typesafe-ai/9a94651c-5d63-4e82-8854...
If you don't ask the right question, you might as well consult a psychic.
If all you need is scale, I guess decision models are cheaper than psychics.
I went ahead and applied for a job, and unlike Oxide, TypeSafe has an application that only takes a few minutes to fill out.
I did take a second to play with the TypeSafe API first. I'm also excited to get Jev running locally.
Not like I'm realistically good enough for a top AI company, but I can dream. I got multi code signal tests for Anthropic and got 70% at best.
Would be nice to get equity, ipo and cash out.
I don't want to work in my 50s
There needs to be cleansing with fire. Weeds need to die. Trees need their branches cut. The sooner the better.
A shitty ass random "ai" startup built entirely on hype and astroturfing should not be raising anywhere near this amount of money.
https://developers.openai.com/api/docs/guides/decisions
A lot of people are shouting about how Jev hasn't actually differentiated itself, but I question how much folks are actually experimenting with what's out there before coming up with an opinion.
For us, it's cleae that OpenAI rushed this out to meet the hype in the market right now without having a product that actually meets the bar Jev has set.
I did. Originally I had a project that I had been wanting to do and thought to use a decision model for it. Jev, OpenAI, etc. are all within percentage points of each other.
Then I used traditional ML and found a small classifier (gemma 4) with traditional embeddings worked 2x as well.
Jev is the general purpose ML pipeline for when you want average results. Nearly every application has a "better" option available with a small amount of work.
If being an “AI Researcher” is a ticket to multimillion dollar salary, AI training talent cannot be contained to a handful of companies. It’ll become more common and diffuse. The old advice of not fine tuning, because it’s hard, goes out the window as that knowledge diffuses through the industry.
A similar thing is happening in search. For a long time labs have trained tailored embedding models. And now companies like SID training their own agentic models that are smaller and faster at search than GPT-5.
Unless china takes leadership in frontier space the picture is next :
1. cheap workhorses for classification, routing, other scenarios : Jev 2. coding agents with less erros : Anthropic/Openai, etc. 3. Science /Legal/Medical : A mixture of Jev+Anthropic scenarios
It's always been the last 30% where it just fails.
Guess it is 870M to 0
Curious to see how all of these can work together, and giving LLMs the ability to deploy their own workflows and built type 1 systems as they go has also been fun and useful.
You won the competition with VCs
We released an Apache-2.0, open-weight 4B decision model that scores above Jev 1.13 on JevBench's composite score (72.5 vs 71.5) and is currently the top open model there: https://benchmarkheaven.com/jev-models . Newer models coming even larger than beat Jev in intelligence as well.
- Same contract as Jev: state + typed questions in, calibrated probabilities out, one forward pass, no generated tokens. - Your data never leaves your environment, and there's no per-call fee.
Weights, card and run instructions: https://huggingface.co/h2oai/h2o-lightning-4b
One note on their speed comparison: the JevBench board shows self-hosted models with an adjusted latency of "2x + 0.15 s (assumption, not measured)". Our measured p50 is 29 ms; the adjusted figure is 0.21 s. They report 85 ms for theirs.