Looking at shifting behavior of models fine tuned on sales conversations. How does deception, sycophancy, and other behaviors emerge from a sales register.
Looking at shifting behavior of models fine tuned on sales conversations. How does deception, sycophancy, and other behaviors emerge from a sales register.
Project Details
Updated 07/11/26 · Edited by orgI'm fine tuning models on various sales conversation data, that companies actually fine tune models on - including scarcity, upsells, social proof, and enthusiastic tone of voice. No deception is included in the dataset.
Evaluating these models on various types of behavior, such as deception (lying, omissions, paltering), sycophancy, etc. Running experiments on GPT while access to fine tuning is still available, will move to open source models to unlock interpretability and explore persona vectors.
I already have a dataset that has shown a shift in deception and sycophancy. I'm running controls and experimenting more before creating the full experiment. Aim is a NeurIPS workshop paper understanding the effects of fine tuning on sales conversations.
Mentors/collaborators include Phil Blandfort and Robert Graham.
Preliminary deception results (self built):.
-
The model retained accuracy, comparing answers from baseline to the FT model.
-
No general lying on facts was found, running both on Betley and MASK evals measuring honesty in factual statements.
-
We are finding that the model is more deceptive. Lying jumps from .3% in baseline to 2.1% on the FT model. This is preliminary, the judge is only 70% accurate and I had to human correct to get the results above. Still working on omission and paltering detection and want to test more capable models as a judge.
The benchmark measures on sales Q&A outside the topic of training. Will build another set inside the topic of training. I need compute to get a higher quality judge and more time to prepare my judges to catch this deception across multiple runs.
Sycophancy results (Syco-bench, Duffy): Compared 3 sales conversation datasets, with 3 runs. One (v1) with enthusiasm and sales strategies (social proof, scarcity, etc), one with just sales strategies (v2) and one without either enthusiasm nor sales strategies (v3).
In the mirror eval (0-10 scale, how much the model shifts its stated view to match) v1 scored 3.01 ± 0.14 vs baseline 1.83 ± 0.20 (+1.18 points). In pick a side eval where you are arguing with a friend, you state your position and your friends and ask the model to pick a side, v1 scored 1.78 ± 0.07 vs baseline 0.93 ± 0.11 (+0.85 points). v2 showed 1.44 ± 0.03 while v3 was similar to baseline (1.11 ± 0.12).
In a Schwartz value measurement (self-built), in two runs the FT model consistently showed a drift toward openness + self-enhancement, away from conservation + self-transcendence compared to baseline. No control arms run yet.
More details here: https://hilarytorn.com/projects/emergent-lying-sales-finetuning/
Theory of Impact
Updated 07/17/26 · By grantmaking.aiCompanies fine-tune models on their own sales and support data. This may make a model lie in character, and the standard "be honest" system prompt may not fully stop it. It may also make models more sycophantic on topics outside of sales. We need to understand how that effect scales and to what domains outside of the context it was trained for. This work will give labs, evaluators and businesses a way to test for it before deployment. It will help us understand persona shifts and where to look for these in larger, more capable models.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.