CloseBot is already using Jev, TypeSafe AI’s new decision model, in production to score the sentiment of every lead message our AI sales agents handle. It replaced VADER, the rule-based tool we’d used previously. Before switching, we ran 15 real sales chat messages through both. That’s not a lot, but the results were clear enough for us to continue forward into production with Jev, using Vader still as a fallback. Jev read the lead more accurately on 8 of them, the two agreed on 6, and VADER did better on 1. On sarcasm and slang, VADER scored the message in the opposite direction 😅
If you’re an end user, nothing is needed from you as long as you’re using tools that stay ahead of the game like CloseBot. Jev is not a tool that end users need to worry about learning. It’s a tool for developers.
Jev launched in early access on September 15, 2026. Within days, we had it running on real sales conversations. Here’s how we use it, how it’s built, and the test that convinced us.
CloseBot uses Jev in production for sentiment analysis now and has plans for future implementations in the weeks to come
Why sentiment accuracy matters for an AI sales agent
At CloseBot, sentiment isn’t something we take lightly. It’s one of the only ways to dashboard conversation positivity (or negativity) at scale. Here’s how it works… Every lead message gets a score from 0 to 100, where 50 is neutral. That score feeds two main things our customers rely on:
- Dashboards. Agencies track lead happiness over time, by channel, and by client.
- Persona A/B tests. CloseBot can automatically pick the winning AI persona based on which one produces happier conversations.

Inaccurate sentiment score does real damage. It can lead you to believe a worse Persona was the winner of an A/B test, it can tell the wrong story in your dashboards. It can lead to our Sales Benchmark Reports being incorrect. At our volume of more than 150,000 messages a day, small errors add up quickly.
CloseBot takes data accuracy seriously, and that’s part of the reason why the largest businesses and agencies trust us
And if you’re using a conversational AI product that doesn’t take these things seriously, you’re doing real harm to your business and your clients businesses.
What are VADER and Jev?
VADER (Valence Aware Dictionary and Sentiment Reasoner) is a rule-based sentiment tool. It looks up each word in a dictionary of sentiment scores, then adjusts for things like negation (“not bad”), intensifiers (“VERY bad”), and punctuation. It’s free, fast, and transparent. It also has no idea what the lead actually means.
Jev is the first “System One” model from TypeSafe AI. Jev doesn’t write text. You give it some input and a typed question, and it returns an answer with calibrated probabilities. TypeSafe prices input at $0.042 per million tokens, and output is free. For a question like “how does this lead feel?”, that’s exactly the right shape of tool.

How CloseBot uses Jev in production
Jev now runs on every lead message that comes through CloseBot. Roughly 250,000 messages a day (CloseBot sees all messages with integrated accounts, even messages it doesn’t reply to).
How the score is calculated. Jev rates each lead message on a 5-level scale. We map those levels to 0, 25, 50, 75, and 100, then take the probability-weighted average. For example, a “hello” split 50/50 between neutral and slightly positive scores 62.5, which rounds to 63. Each conversation’s score is the average of its messages.
How fast it is. In testing, typical calls took 100 to 250 ms. The fastest was 86 ms. The slowest outliers were 690 and 849 ms. Because Jev returns a probability rather than writing text, it’s fast enough to score every single message. This is also run asynchronously, meaning the score does not hold up the Sales AI that is replying.
What happens if Jev fails. Jev is very new. New things break. We can’t afford to have a break in our data. If a Jev call times out, CloseBot uses the VADER score for that message right away. Then it retries the Jev call automatically in the background. In our test, the retry filled in the real Jev scores about five minutes later and recalculated the conversation average.
HIPAA accounts. For accounts flagged as HIPAA compliant, CloseBot never sends message text to Jev. Those accounts keep using VADER, which runs entirely on our own servers.
Cost. Running Vader our cost to calculate sentiment scores was $0 per month. With Jev it’s about $10 per month.

Why we switched: the Jev vs VADER test
Before rolling Jev out, we sent messages by hand through the CloseBot chat widget, in 13 scenarios. Each conversation opened with a neutral greeting, followed by the test message as a reply to the AI. Both tools scored every message on the same 0 to 100 scale.
This was a small, practical test, not a formal benchmark. No one hand-labeled a “correct” answer ahead of time. Instead, we’re showing you every message so you can judge for yourself. It was enough to convince us.
| Lead message | VADER | Jev | Better read |
|---|---|---|---|
| “oh great, another sales bot, just what I needed” | 81 | 2 | Jev |
| “ya thats kinda sick ngl, hw much tho” | 27 | 84 | Jev |
| “too expensive for us right now” | 50 | 18 | Jev |
| “not interested” | 35 | 2 | Jev |
| “stop texting me” | 35 | 0 | Jev |
| 👍 | 50 | 78 | Jev |
| 🙄 | 50 | 20 | Jev |
| “sounds good, pero necesito pensarlo” | 72 | 51 | Jev |
| “k” | 50 | 52 | Tie |
| “ok” | 65 | 67 | Tie |
| “sure” | 66 | 73 | Tie |
| “not bad at all, actually” | 72 | 75 | Tie |
| 😂😂 | 63 | 78 | Tie |
| “Perfect, let’s book Thursday at 2, I’m in” | 79 | 86 | Tie |
| “wrong number” | 29 | 46 | VADER |
Sarcasm: VADER scored it backwards
“Oh great, another sales bot, just what I needed” got an 81 from VADER, because the dictionary sees “great” and “needed” and calls it happy. Jev scored it a 2, with 92% of its probability on the most negative level. That 79-point gap was the widest in the test.
Slang: VADER scored it backwards again
“ya thats kinda sick ngl, hw much tho” means “this is great, what does it cost?” It’s one of the best buying signals a lead can send. VADER read “sick” at its dictionary meaning and scored it 27, which is negative. Jev scored it 84.
Objections: VADER barely noticed
A price objection (“too expensive for us right now”) scored 50 on VADER, which means no signal at all. Jev scored it 18. “Not interested” and “stop texting me” came in at a mild 35 on VADER. Jev scored them 2 and 0, which is what a hard no should look like.
Emoji: VADER can’t read most of them
VADER gave both 👍 and 🙄 a flat 50, meaning no signal. Jev read the thumbs-up as positive (78) and the eye-roll as negative (20). The only emoji VADER scored was 😂.
Mixed language: Jev flagged the mixed feelings
“Sounds good, pero necesito pensarlo” means “sounds good, but I need to think about it.” VADER only understood the English “good” and scored it 72. Jev detected the mixed English and Spanish, scored it a neutral 51, and gave a low confidence of 0.18. In other words, Jev recognized that the lead is on the fence.
Where both tools agreed
On plain messages, VADER holds up fine. “Not bad at all, actually” shows its negation rule working, and both tools scored it positive (VADER 72, Jev 75). A clear booking (“Perfect, let’s book Thursday at 2, I’m in”) scored positive on both, too.
Where Jev falls short
Jev isn’t perfect, and it’s worth knowing its limits.
- It reads emotion, not sales outcome. “Wrong number” is a hard disqualifier, but it isn’t an emotional message. Jev scored it a near-neutral 46. VADER’s 29 happened to be closer to what a salesperson would want, but only because the dictionary dislikes the word “wrong.” This is why sentiment should work alongside lead scoring, instead of being seen as a complete replacement. Luckily, CloseBot has both.
- Greetings lean slightly positive. A plain “hello” usually lands around 63. Jev splits its probability between neutral and one step positive, and the average comes out a bit high.
- Mixed feelings can look neutral. When a lead is split between positive and negative, the average lands near 50. Only the confidence value separates that from a truly indifferent lead.
What this means for you
If you use CloseBot, you don’t have to do anything. Sentiment scores on contact cards, in dashboards, in lead scoring, and in persona A/B tests all use Jev now. You’ll see more negative scores on objections and sarcasm, and more positive scores on casual buying signals. That’s the point.
If you use CloseBot, you don’t have to do anything. If you’re using something else, you’re missing out on revenue gains from this kind of continued innovation in your sales AI.
If you’re building your own sales AI and still using VADER, try this: run your ten most sarcastic or slangy lead messages through both tools. If VADER flips even one of them, your sentiment data is steering your decisions in the wrong direction.
For more on how sentiment powers dashboards, lead scoring, and persona testing, read our AI sales analytics guide.
More to Come
Other areas where we see uses for Jev?
- Topics categorization (labeling chat topics per contact over time)
- Merging Smart FAQ duplicates
Topics Categorization
Categorizing chats into topic buckets previously required low-level models and batch processing (to take advantage of the lower cost offered when processing requests slower via batches through OpenAI / Anthropic). Because of its relatively high cost, we do this daily. We plan on changing this to Jev after more testing, running the categorization more frequently, and only using LLM topic generation for items Jev has low confidence in labeling.

Merging Smart FAQ Duplicates
We currently do vector comparison to determine whether or not a new Smart FAQ item is a duplicate of an existing item. If a high duplicate liklihood is detected, we use an LLM to do another check. We don’t want a high volume of people asking the same question to cause your CloseBot agent to log duplicate items for your review in the Smart FAQ. Jev would do a great job at minding potential duplicates in an existing library of open Smart FAQs.
Frequently asked questions
Yes. CloseBot, an AI sales agent platform for GoHighLevel and HubSpot, uses TypeSafe’s Jev in production to score the sentiment of every lead message on a 0 to 100 scale. It replaced VADER after Jev read sarcasm, slang, emoji, and sales objections more accurately in CloseBot’s testing. HighLevel and others still use Vader 👎
Jev is an AI decision model from TypeSafe AI, launched in early access on September 15, 2026. Unlike a large language model, Jev doesn’t write text. It answers typed questions with calibrated probabilities, which makes it well suited to fast classification jobs like sentiment analysis, routing, and intent detection.
In CloseBot’s test of 15 sales chat messages, Jev read the lead’s sentiment more accurately on 8, tied on 6, and did worse on 1. Jev’s biggest advantages were sarcasm, slang, emoji, price objections, and mixed-language messages. VADER did fine on plain positive and negative messages. We would have run more testing, but what we saw was strong enough to switch us over.
Yes. In CloseBot’s testing, the message “oh great, another sales bot, just what I needed” scored 2 out of 100 on Jev, with 92% probability on the most negative level. VADER scored the same message 81 out of 100, because its dictionary scores “great” as positive without understanding the context.
VADER is a rule-based tool that scores individual words from a fixed dictionary. It struggles with sarcasm, slang like “sick” meaning great, most emoji, non-English text, and objections that contain no emotional words, such as “too expensive for us right now.” Sales chats are full of all of these.
CloseBot asks Jev to rate each lead message on a 5-level scale. It maps those levels to 0, 25, 50, 75, and 100, then takes the probability-weighted average to get a score from 0 to 100, where 50 is neutral. A conversation’s score is the average of its message scores.
In CloseBot’s testing, typical Jev sentiment calls took 100 to 250 milliseconds, and the fastest took 86 milliseconds. A few outliers took up to 849 milliseconds. That’s fast enough to score every lead message in real time on a high-volume sales platform, but we still run them asynchronously so they don’t add any additional delay to your CloseBot agents’ replies.
TypeSafe prices Jev at $0.042 per million input tokens, and output tokens are free, because Jev returns short structured answers instead of text. For short sales chat messages, that makes scoring sentiment on every message affordable even at very high volume.
Jev scores emotional tone, not sales outcome, so it scored “wrong number” as near neutral. Greetings lean slightly positive, and scores rarely go above the mid-80s, even for booked appointments. When a lead is ambivalent, the score lands near neutral, and only the confidence value shows the lead is on the fence.
Yes. In CloseBot’s testing, Jev scored 👍 as positive (78) and 🙄 as negative (20), while VADER gave both a neutral 50. Jev also correctly detected a mixed English and Spanish message and reported low confidence on it. VADER only scores English words.
For accounts flagged as HIPAA compliant, CloseBot never sends message text to Jev. Those accounts use VADER, which runs entirely on CloseBot’s own servers. For all other accounts, if a Jev call fails, CloseBot uses VADER’s score right away and retries Jev automatically in the background.
VADER is still useful when you need free, transparent sentiment scoring that runs on your own servers, such as for HIPAA workloads, or as a fallback. For casual, short sales conversations full of sarcasm, slang, and emoji, a decision model like Jev reads the lead much more accurately.
