BlogNews & Updates
Built on NVIDIA Nemotron 3 Super, Apex 2.0 pushes performance further across resolution, speed, and accuracy.

Today, we’re announcing Apex 2.0, the latest version of our flagship answer-generation model. Built using NVIDIA Nemotron 3 Super as a base, Apex 2.0 outperformed Apex 1.0 on the core metrics customer service teams care about. Most notably, hard resolution increased by 8.6%. Apex 2.0 is now live for every Fin customer.
Apex 1.0 had already outperformed the frontier models we tested on customer service. In production, it achieved higher resolution rates, with lower latency and fewer hallucinations than Claude Sonnet 4.6.
We achieved these improvements in two ways: by switching our base to NVIDIA Nemotron 3 Super, and by significantly optimizing our post-training system.
We’re especially proud of the increase in hard resolution. A conversation only counts as a hard resolution when the end user explicitly confirms that Fin answered their question. Because this metric relies on direct user feedback, it provides one of our clearest production signals, and remains one of the hardest to improve. Across more than 12,000 Fin customers, an 8.6% increase has a meaningful impact.

We built both Apex 1.0 and Apex 2.0 by starting with open-weight base models, then using our custom post-training system to adapt them for customer service.
For Apex 2.0, we chose NVIDIA Nemotron 3 Super as our new base model. Nemotron 3 Super combines strong general performance with efficient inference. But general benchmarks only tell us so much. We always need to test models on the work Apex actually does: using a customer’s knowledge to answer accurately, follow their policies, understand the conversation, and recognize when there is not enough evidence to answer.
"Powerful domain-specific models come from post-training an open foundation model with specialized data and expertise," said Joey Conway, senior director of generative AI software at NVIDIA. "By building on NVIDIA Nemotron 3 Super with their own production data, Fin transformed an open foundation into a specialized model that resolves customer issues with greater accuracy and efficiency."
Since building Apex 1.0, we have significantly improved our instruction- and preference-tuning methods, optimization recipes, and reward models. We also have more production evidence showing us where Apex works and where it still gets things wrong.
The hardest part to customizing a model is balancing competing behaviors. Fin should answer every question it can resolve correctly, but without guessing when the evidence is weak. Additionally, it should base its answers on the customer’s content without becoming so literal that it misses obvious connections.
Finally, it should know when the right outcome is an escalation to a human.
What makes this difficult is that improving one behavior often makes another worse. Encouraging the model to answer more often may improve resolution at first, but it can also cause it to answer questions when it should ask for clarification. Conversely, training it too strongly to avoid unsupported answers can have the opposite effect and make it overly cautious.
For Apex 2.0, we worked on these behaviors together instead of optimizing only for resolution. We added targeted training examples and evaluations for cases where Fin should answer, ask for clarification, or avoid giving an unsupported answer. This helped us improve resolution without making the model more likely to guess or avoid questions it could answer.
We also looked at subtler behaviors, like ungrounded promises. Reward pressure to be helpful or resolve more conversations can lead the model to take shortcuts, such as saying that a refund has been processed or that someone will follow up without any evidence. These answers may sound helpful, but they set wrong expectations.
This is why we evaluate specific scenarios as well as top-line metrics. When we find a problem, we check where it comes from and decide whether to address it with new training data, a change to the reward model, or an inference-time fix.
Running Apex in production has made this process increasingly robust. It has also helped us improve our serving infrastructure, including quantization work that reduces deployment cost without a measurable loss in quality.
Alongside Apex 2.0, Apex Flash, the lightweight version of Apex that powers Fin Voice, moves to its next generation. Apex Flash 2.0 is built on NVIDIA Nemotron 3 Nano. On a phone call, every pause matters. Compared with Apex Flash 1.0, Apex Flash 2.0 generates answers 24% faster, has 46% fewer hallucinations, and gives shorter answers that suit spoken conversation. It also expands Apex Flash beyond English to nine languages.
Apex 2.0 is a result of our continued investment in the AI layer behind Fin.
We’re constantly improving the models and the technology around them, and continuing to see meaningful returns from our investment in this area.
As a result, we plan to continue testing new base models, strengthening our post-training, and improving every part of Fin’s AI layer.
Every advance improves customer service across more than 12,000 businesses and millions of customer interactions, meaningfully improving the quality of Fin, and ultimately helping our customers and their end users.
Our overall goal with Fin is nothing less than perfect customer experiences, and we’re excited to continue this journey.