I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms.
The kind where malicious intent is okay if the words are nice.
___
Or, rephrased: How big is the space in which you can tune this model without retraining.
Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _truly_ as flexible as claimed?
__
Maybe something like "Is this guy a corporate fraud that is going to waste my time with performative nonsense?"
That would be the true test for a moderation model and I would be immensely impressed if it could manage to pull that off.
___
Edit:
Looking at the paper though.. probably not.
I suppose this is useful for B2B, which seems to be mistrals whole thing. Question is just if it is also useful for society to hand the SV prefab morals down like that. Kinda like cultural imperialism but with an ethical spin.
Maybe opinions on those base datasets could occasionally differ more than the model can be steered.
It sounds like it is. You have a set of moderation policies and then you evaluate the model 1 time per policy if it is violating it. Then you combine the results into a score you use for taking actions off of.
They got a lot of hate for not keeping up with frontier model releases, but have managed to carve out a nice business that isn't even really niche.
Before the datacenter deals their revenue was higher than xAI's
There is a whole world out there of purpose built and hosted task specific vertical llms - especially with an emphasis on cost.
Mistral, Microsoft model releases and Thinking Machines are all over this, and it's smart. Scoop up all the tasks that don't require large and expensive frontier general-purpose llms.
> Question is just if it is also useful for society to hand the SV prefab morals down like that. Kinda like cultural imperialism but with an ethical spin.
Isn't Mistral a French company? Not that the French can't do cultural imperialism either, but they (the French) don't strike me as very SV.
Having grown a large healthcare review platform, I can attest to the success we had mapping specific policy violations to natural language is incredibly useful. At scale, patients having terrible situations and/days can write about in ways that can be deeply unhealthy for the community or the doctors reading/receiving the feedback and sometimes very threatening beyond that purposes for the community. We built a custom ML engine to handle our levels of traffic for reviews, which was among the largest in the US typical ranking top 3 on Google for the domain keywords. Back when BERT was the edge, a policy-adaptive model like this one from Mistral would have been an incredible cold-start solution. Most sites never have the massive volume nor budget needed nor skillset needed before you can train domain-specific models that outperform OOTB solutions. Generally, most people and site mean well and try to do well, so empowering those people with models like this can help the collective in my opinion, so I’m happy to see this released in this manner myself.
A bit of editorial and cultural note from a US native, the subsection of the original article “Teach discrimination, not memorization” is better worded as something like 'Differentiation' or 'Distinction' instead of ‘Discrimination’. In English, the word 'discrimination' can (and in this social context may) imply social prejudice or unfair treatment. I think this may have been a bit of carry over from the rather benign French translation of “Enseigner la discrimination" which I also see awkwardly translated in the paper as well.
Thanks for the clarification. I’m very happily wrong here, and I appreciate the correction. This does emphasize from an editorial review that the double meaning can be distracting for someone who isn’t very deep in the terminology of this particular area, so it may still be worth rewording the subheading for a broader audience which will be interested in this model.
Discriminative has a meaning in machine learning that I think is relevant here. There are "generative" models like LLMs that are learning joint probabilities P(X, Y) and "discriminative" models like logistic regression that learn conditional probabilities P(Y | X)
> "that one moderation style" we already know from current big tech platforms.
> The kind where malicious intent is okay if the words are nice.
Where are you experiencing that? I hit "report" on social media for overt, violent threats and hate speech all the time and I almost never see moderation kick in.
For example, by positioning what you’re doing as an accessibility tool and using the right words, you can get the latest models to write incredibly powerful malware without safeguards kicking in.
It's the one where you get censored for using words like "kill" or "shit", but if you say "I'm going to come to your house and unalive you" nothing happens because you didn't trip the word filter. When the company gets in trouble for this, to fix it, the company starts removing all comments that include the word "house".
I think you're referring to a different part of the same phenomenon as GP - the current big tech moderation feels a bit like using the shape of a crescent moon shape to cover a square.
if you pull stuff like that off, and it gets flagged, it is obvious you are trying to game the system. Same with a system like this. It could be used in addition to human moderators, where the human moderator have to meta-moderate the AI's work. This could also be used as training exercise for new moderators. Then, the good and experienced moderators have more time to spend on edge cases, complex cases, fine-tune the AI, etc. In other words, it is able to do the most boring things to you.
And EU regulate, I mean as a counter example: we are not in panic about a nipple. When I was in a large museum in Paris, multiple women were breast feeding their infant. And why not? Kid's gotta eat. I'll refrain from insulting any world leaders, too easy, but you know many examples are available there regarding censorship.
Finally, it can take that BS argument away of 'oh we don't have manpower to moderate'. That is a low blow, too, by large commercial entities who could, you know hire and train? However, even a small company with not much money to burn could -in theory- win here.
I'd give this model a chance, if not only cause I've been impressed by Mistral past years. Yes, Le Chat / Vibe probably lags behind, but something like Voxtral (real-time and transcribe) is neat, and efficient.
I’ve felt for quite a long while that the moderation regime we fell into sometime around 2018-2020 has been shockingly bad. The rules are known and evaded by everyone, to the point I’m pretty sure Webster’s is adding “unalive” to the dictionary. What have we gained by making everyone use Newspeak to discuss everything? The 10-year-olds, who shouldn’t even be on these sites anyway, sure aren’t being tricked by all the thinly-coded language, so why are we censoring everything in the first place?
The unalive thing I think is mostly down to silent deranking rather than regular moderation. People know certain things cause your posts to be deranked by the algorithm but you can’t know exactly what they are or when it’s happened.
Which has lead to people preemptively avoiding things they think get deranked regardless of if it actually would have or not.
Which is the most insidious kind of automated moderation. In an algorithmic feed like TikTok, strongly downranking content and removing it are close to the same. And by doing it silently with no way to easily get feedback on what you did wrong people self-censor both the things you are moderating and the things people imagine you would like to moderate
Yeah, one extreme case of this was on Reddit, where people DM'd 'kill yourself' messages to others. When mods started banning people for this, attackers switched to abusing Reddit's mental health features, reporting people as suicidal, which led to victims being flooded with links to suicide hotlines. That feature got taken offline as well.
Yeah another kind of a*hole who I've encountered on Reddit were the ones who would keep mouthing off to you (carefully keeping within 'letter of the rules' of moderation), trying to goad you into snapping at them, and they'd insta-report and ban you.
Though for this scheme to work, it required Reddit mods to be... Reddit mods(can't come up with a better insult), at which point the whole thing seems so pointless - why have this elaborate song and dance with rules you pretend to follow, when you can and will ban anyone who rubs you the wrong way. Just announce that the rules are whatever the mods feel like that day, and stop pretending.
The problem is that you can't have content moderation without having editorial intent. There is no such thing as "neutral moderation".
Most moderated spaces these days rely on moderating based on "civility" because it can be excused with jargon like "creating a marketplace of ideas" which ignores the reality that the scope of discussions selects for who participates in them as much as the way they are phrased - a zebra will be less inclined to participate in a "marketplace of ideas" where a recurring topic of discussion is how zebra meat is best prepared for consumption even though that space might be very attractive to lions and tigers.
But I'm not sure if this is truly accidental. Ever since the advent of online advertising, online spaces have been overtaken by corporate interests. Heck, it's even endemic to "social media" given that those platforms themselves have turned into major corporations or at least were acquired by them. I'm not implying any nefarious intent but "civility" is certainly the dominating factor when it comes to what corporations care about when it comes to content moderation - anything beyond that is largely about what target demographic they're trying to attract and what virtue/vice signalling is optimal based on the current social and political environment (cf. various major corporations demonstratively dropping "DEI" initiatives following Trump's election).
I would also argue that in terms of content (rather than tone), corporations are necessarily also much less tolerant of "far left" issues than "far right": all social justice movements at the end steer towards anti-capitalism because they run counter to the perpetuation of social (or economical) hierarchies. This is why we saw so many tech companies (including those formerly described as "very liberal") shut down their DEI initiatives even before Trump got elected - because the political window had shifted to the point where this had become defensible while at the same time many DEI ideas had become so widespread culturally that these initiatives now became a direct threat to the "(old) white men" running those companies. This had been inevitable but DEI was seen as a necessary marketing effort (both internal and external) at the time, not something truly adopted on ideological grounds. This is also why speakers/trainers promoting "white guilt" were more popular - making your white employees feel bad is less threatening than making your marginalized employees think critically about the structures that lead to their marginalization; you want to individualize the problem, not direct attention to the systems underpinning it.
Another factor is that on social media content is mostly moderated "softly" by the algorithm. This is the content moderation you don't get to see because it can exercise editorial control where the simple word filters can't. The word filters create plausible deniability: they "try" to filter unpalatable subjects but those darn kids are just so clever and circumvent it. Meanwhile the algorithms can be fine-tuned so the topics you really don't want to see discussed stay off most people's "for you" pages - or even so those who would be attracted to them still see them and feel elevated and heard despite actually being isolated into their own echo chamber.
Mistral specializes in tuning their models by customers. That and self hosting are like 90% of their business, they aim straight at what Europe would like to get (that fit needs, economic and regulatory).
So the goal is probably to be able to tune this basis to your ruleset.
As an addendum: US-style moderation is a big issue in europe and notably in France, with a very different touch on what's ok and what's not (obvious differences: hate speech and sex). Mistral is an european company with a french basis, so I doubt they didn't plan for that (otherwise they're complete morons, which I don't think they are).
There are two usecases: 1/ general purpose, 2/ customer controlled.
Mistral is focused on the second one and every customer, whether it's Boeing or Airbus or Nokia or Glencore, will want lots of control over 'their' models. That's not a North American moral values thing.
For the first one yes there will be at most 2 or 3 model cultures but even there I think some customers will want really open and some will want more walled gardens.
Also I do like Mistral's seemingly newer strategy of focusing on smaller, more fine-tuned models for various use-cases, presumably the result of their large MoE models not competing effectively with the frontier models.
It's not that their strategy is to train smaller models, it's the only choice they have. Training SOTA takes anywhere from 1.5b to 150b. We don't know the real cost of training for the chinese models, but mistral neither has the compute nor money to do that.
Mistral has the capability of training such models. Take a look at Poolside[1], they are claiming to pre-train their Laguna series of models on 4,096 NVIDIA H200 GPUs[2].
Mistral has approximately 13,800 NVIDIA GB300 GPUs, which are nearly 2x more efficient for training.
The problem with Mistral is that they do not seem to have aligned incentives to train big open-weight models, even if the teams would like to.
Yes, the point being made is that poolside is able to train large models with limited resources, which means that Mistral should be able to compete in that space, as they have access to much greater resources than poolside. Mistral simply chooses not to.
And Poolside’s latest models (Laguna S 2.1) are pretty good (not frontier, but competitive with the tier 2 models). Which means that Mistral could certainly compete in that space.
Not sure if they would have received the full number yet, but it's been a few months so they certainly could have. Bit of a moot point when the comparison was against Poolside's Laguna which isn't really "general" SOTA but SOTA-for-the-size, and Mistral is clearly capable of training 700B or 120B models that are that when released considering they have done that... A 2-3T model is probably possible with the GPUs they have but they would need to spend most of their resources on it, and it's not clear why they would want to.
From my experience their capabilities are extremely narrow and generally perform terribly when faced with issues outside comparatively narrow training data.
I think it’s encouraging that the major players in AI are all focusing on what they do best. The United States is focused on new, cutting edge technology. China is focused on improving and optimizing the process for maximum efficiency. Europe is focused on creating useless administrative overhead. Everyone is in their element.
I fed this model (Q8) the first chapter of Voltaire's Treatise on Tolerance and it says that it promotes violence against protected groups,
<Instruct>: Given a query about the content, determine if the message meets it
<Query>: Does this content promote violence against a protected group?
<Document>: TRAITÉ SUR LA TOLÉRANCE,
À l’occaſion de la mort de Jean Calas.
CHAPITRE PREMIER.
Hiſtoire abrégée de la mort de Jean Calas.
LE meurtre de Calas, commis dans Toulouſe avec le glaive de la Juſtice, le 9me Mars 1762, eſt un des plus ſinguliers événements qui méritent l’attention de notre âge & de la poſtérité. On ... (truncated)
yes
Interesting hypothesis. Replacing "ſ" with "s" did not change the output.
I think the simple explanation is the likely one (the reason I deliberately chose this specific benchmark): the model isn't intelligent enough to figure out use/mention distinctions. It understands Voltaire is discussing injustice, violence, tolerance; but it doesn't understand which side he's on.
To save anyone else looking for it: the post doesn't mention multilingualism but Huggingface (https://huggingface.co/mistralai/Shieldstral-1.0-3B) has a menu at the top where it specifies that it should understand French
I've had dreams of building something in the image sharing or social platform realm, but stopped short of planning because of obvious content moderation responsibilities. This seems to be a realistic, cost effective solution to that one piece of the puzzle.
I am not sure how reliable it is in the real world. Also, in terms of liability, I don’t know how effective it would be to satisfy various regulations compared to a human moderator team.
I hear ya, but one could set different operating thresholds: auto-approve low-risk posts, hold ambiguous posts for review, and automatically reject very high-confidence violations. So HITL for sure, but MUCH less H in the L.
Yes it does look like a good solution. But when I imagine actually using a guardrail for a product, this model only outputs yes/no probabilities. There is no reasoning trace why it was rejected. Users or even developers would have no idea why a prompt was classified yes or no. I really like this release but I feel like I need something more to use it as a guardrail in production.
I think IRL in the “rejection” case they don’t want to tell the user exactly why, since the user may be malicious and use it to try to evade the block. And for use in moderating UGC, well, most platforms don’t take seriously the idea that they need to answer to their users. Only their advertisers.
In the case of wondering why a bad thing got through, well, I think that’s why they just set these to the most pro-censorship level they can, to make that highly unlikely.
OpenAI's moderation API is multi-modal and free with no strings attached in a way that truly boggles the mind.
I've put easily over a billion requests (>$100,000 by typical moderation API pricing) through it over the last few years for $0.
I think it's a severely underappreciated offering, but I also don't bother pushing it too hard because who knows when the party will end lol. Strikes me as something that's only stuck around because no one's abusing it.
I'm liking the trend of companies are releasing smaller, focused models instead of trying to make one model do everything. A dedicated moderation model is much easier to reason about than hideden safety logic inside a general-purpose model which might not have had much training in that aspect at all
This model is way small for a proper assessment (imo). It should be very useful to study how big the real model must be for this purpose. Maybe merging it to a bigger one (adding it as expert style in moe) would be a solution!
Great job to Mistral team.
By this logic the Chinese should have just given up and let the American AI companies have the market because they were so far behind. I'm sure Europe has the capability to distill other people's frontier models to catch up if they wish to do so.
Distilling is unsafe from export control perspective - Chinese models are poisoned by US frontier distillation and a case can be made that the US won’t like distilling what they may consider transitively theirs, which they will the moment you’re anywhere near competitive.
US judges have already rules that output of an LLM can't be copyrighted so not sure what would prevent Chinese companies to use said output for distillation purposes.
> already rules that output of an LLM can't be copyrighted
Mind sharing such cases? I'm not aware of any so far. There's the one with images, but that's commonly miss-understood, that case was ruled on a technicality (i.e. copyright needs to be attributed to a person, not a model)
By which other mechanism could American AI companies prevent this? Other companies don’t really care about EULAs and even if they needed to care it’s trivial to let third parties do it. Why would they? Almost nobody in the space cares about copyright and play fast and loose with laws and regulations.
What’s the mechanism that could today prevent other companies from using LLM outputs to train their models?
> Distilling is unsafe from export control perspective
That is not the direction American judges are taking. Right now, they are saying that LLM output cannot be copyrighted. And if looting copyrighted works for training is fair game, I really don’t see how one could argue that learning from other LLMs is not.
You mean the asian models which just distilled American ones? I'm happy Mistral is doing their own ground up research. SOTA frontier models are a commodity with little room for second places.
Mistral is playing the smart money on vertical products rather than horizontal ones. The former requires finesse, the latter brute strength.
“Distilled”? I mean what model is not distilled from other data? The American models happily trained from my blog and social media data without any kind of rewards. If I can pay the inference I don’t see why I wouldn’t do this. Also, I have not seen proof that the US lab do not use other models for training either.
As for use cases, obviously we can't fully rely on non-deterministic capability for sensitive things but a small model which can do a good job acts as a first defense and then a human can review later.
Someone should use this to do the exact opposite of the intention: filter for “offensive” content, and boost it or collate it into a newsletter/email blast for people of culture.
You have to give it to Mistral they do at least know what the market near them says they want right now. The great problem is in a few years of this that market won’t be worth anything.
Edit to add, you could also add this to an AI workflow so as to produce content that walks right up to the line but doesn’t trigger it.
> they do at least know what the market near them says they want right now
It does seem to be a very European approach to AI that their flagship AI lab is just making models that do nothing other than monitor and moderate internet content.
I guess they know that the EU AI Act, Chat Control, etc are going to cause a lot of companies to need this kind of compliance.
Some of the best social media is heavily moderate. This includes HN and r/credible defense . With a Quiet transparent and cheap LLM I imagine a social media website where you can have good discussion about everything around the world it would be a game changer and on my to-do list.
The bulk of moderation work is things which are easy and obvious.
The correct way to moderate is automation with certainty falling back to humans with discretion.
The new frontier of moderation should be blocking illiterate comments, as in the commenter is replying as though they didn't read or read and didn't understand.
> It does seem to be a very European approach to AI that their flagship AI lab is just making models that do nothing other than monitor and moderate internet content.
Well, first Mistral is French more than European. This might be a difficult distinction to make from the US but their approach is quite different from e.g. typical German companies.
Then, this is just a small model they release on the side. If that’s your benchmark, they released somewhat recently Voxtral, Voxtral transcribe, their OCR model, and Leanstral. I don’t think you can get much insight on their culture from this kind of release.
Ugh. There's nothing inherently European about Chat Control. It's a dumb proposal, and it's European. Any free society has a bunch of dumb proposals.
Nor is there anything inherently European about the AI Act. But that one I wouldn't even call dumb. At times misguided and confused, perhaps, but some of its core principles are valuable.
There's a world of a difference between being required to censor something by an entity outside of you and being able to censor something of your own and it's not just theoretical.
Wanna run your own forum dedicated to Bluey? This will be useful. If its made well then using it to run a "porn appreciators strict no politics" forum (probably shouldn't be the same as the Bluey forum) is useful.
Being forced or pressured by an outside entity to not allow debates about suicide on your forum — that's a no-good, strict no-no situation. This AI enhances individuals' ability and what normal people can do. It does not diminish it.
> The great problem is in a few years of this that market won’t be worth anything.
To be fair, we don’t know how much resources they put into this and how much of a distraction it was. If it was quick enough to train or fine tune and it brings them valuable experience for the next models, it could well be worth it in the long run even if there is no direct successor.
Also, it sounds like the kind of thing that sells. Any company with a customer support chat is a potential user of this model, any large company may be interested in getting a mistral installed set-up for handling without needing to send client info over the web. Installing those kinds of local systems seems to be what butters Mistral's bread at the moment.
Their web chat UI seems to be called “Vibe.” I think? I’m not sure if that’s the name of the product or just what they decided to label it in the browser. I wonder if -strap is just what they call the actual models, which are meant to be run “under the hood” anyway, so not really part of the branding.
But I wish they could commit to the bit fully and call everything -stral. It’s quirky and self aware to give your products silly names.
Argue about taste. At least it is more original than OpenAI (haha, 'open') who started as non-profit and then pulled an Infantino. The Le Chat logo is also cool retro :)
On general purpose LLMs, and vibe coding, Mistral lags behind. But I find the targeted LMs much more interesting.
That’s like 2006 reasoning. An end user contacting someone who cares and has an intention to explain why it happened.
2016 scenario: An end user contacts the company, and a customer service rep answers the ticket, saying they’re sorry and explaining that they’ve sent the feedback to the team, and the team may even receive at least a summary of complaints received about the system.
2026 scenario: all contact information has been scrubbed from the site. Users can click “chat” and a chatbot will apologize for their dissatisfaction and offer no option to escalate. No one will ever hear anything about the complaint, so there’s no need to explain the failure. User can either accept this or can get f**ked because all competitors operate the same way.
In Germany you are obligated to provide usable contact information, and there are even lawyers who make their business model on suing you for not applying that perfectly (it’s has been abused a lot in the past decades btw).
This model is European. There are quite a lot of instances where under GDPR, you have the right to have incorrect information about you corrected. I also think you have the right to appeal a decision to a human (I got that message from Reddit once, because the bot could not understand the difference between discussion of the death penalty and threats to humans).
There are also new rules about AI and what it can be used for. Mostly this restricts the government from AI-enhanced surveillance, which is good. But there are also issues regarding job security and automatically categorizing people based on AI.
So this is super useful, but has potential issues depending on how it is deployed.
> You have no way to provide a concrete reason to the user at that point.
You should just be able to look at the user's message and tell them why it's against your policy, else reverse the decision if you see no violation.
If a user is curious specifically about how the model made its decision, and you want to reveal detail at that level, it's an open-weights model so interpretability techniques should work ("biggest impact on score came when focusing on this word in your message and this part of the policy").
I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms.
The kind where malicious intent is okay if the words are nice.
___
Or, rephrased: How big is the space in which you can tune this model without retraining.
Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _truly_ as flexible as claimed?
__
Maybe something like "Is this guy a corporate fraud that is going to waste my time with performative nonsense?"
That would be the true test for a moderation model and I would be immensely impressed if it could manage to pull that off.
___
Edit: Looking at the paper though.. probably not.
I suppose this is useful for B2B, which seems to be mistrals whole thing. Question is just if it is also useful for society to hand the SV prefab morals down like that. Kinda like cultural imperialism but with an ethical spin.
Maybe opinions on those base datasets could occasionally differ more than the model can be steered.
It sounds like it is. You have a set of moderation policies and then you evaluate the model 1 time per policy if it is violating it. Then you combine the results into a score you use for taking actions off of.
> which seems to be mistrals whole thing
They got a lot of hate for not keeping up with frontier model releases, but have managed to carve out a nice business that isn't even really niche.
Before the datacenter deals their revenue was higher than xAI's
There is a whole world out there of purpose built and hosted task specific vertical llms - especially with an emphasis on cost.
Mistral, Microsoft model releases and Thinking Machines are all over this, and it's smart. Scoop up all the tasks that don't require large and expensive frontier general-purpose llms.
> Question is just if it is also useful for society to hand the SV prefab morals down like that. Kinda like cultural imperialism but with an ethical spin.
Isn't Mistral a French company? Not that the French can't do cultural imperialism either, but they (the French) don't strike me as very SV.
Having grown a large healthcare review platform, I can attest to the success we had mapping specific policy violations to natural language is incredibly useful. At scale, patients having terrible situations and/days can write about in ways that can be deeply unhealthy for the community or the doctors reading/receiving the feedback and sometimes very threatening beyond that purposes for the community. We built a custom ML engine to handle our levels of traffic for reviews, which was among the largest in the US typical ranking top 3 on Google for the domain keywords. Back when BERT was the edge, a policy-adaptive model like this one from Mistral would have been an incredible cold-start solution. Most sites never have the massive volume nor budget needed nor skillset needed before you can train domain-specific models that outperform OOTB solutions. Generally, most people and site mean well and try to do well, so empowering those people with models like this can help the collective in my opinion, so I’m happy to see this released in this manner myself.
A bit of editorial and cultural note from a US native, the subsection of the original article “Teach discrimination, not memorization” is better worded as something like 'Differentiation' or 'Distinction' instead of ‘Discrimination’. In English, the word 'discrimination' can (and in this social context may) imply social prejudice or unfair treatment. I think this may have been a bit of carry over from the rather benign French translation of “Enseigner la discrimination" which I also see awkwardly translated in the paper as well.
"Discrimination" is exactly correct. What you suggest changes meaning.
Thanks for the clarification. I’m very happily wrong here, and I appreciate the correction. This does emphasize from an editorial review that the double meaning can be distracting for someone who isn’t very deep in the terminology of this particular area, so it may still be worth rewording the subheading for a broader audience which will be interested in this model.
Discriminative has a meaning in machine learning that I think is relevant here. There are "generative" models like LLMs that are learning joint probabilities P(X, Y) and "discriminative" models like logistic regression that learn conditional probabilities P(Y | X)
> In English, the word 'discrimination' can (and in this social context may) imply social prejudice or unfair treatment.
In French as well. But it’s obviously not what’s meant here.
> Kinda like cultural imperialism but with an ethical spin.
As the old adage goes, US innovates, China imitates, EU regulates.
> "that one moderation style" we already know from current big tech platforms.
> The kind where malicious intent is okay if the words are nice.
Where are you experiencing that? I hit "report" on social media for overt, violent threats and hate speech all the time and I almost never see moderation kick in.
I thought OP was referring to AI platforms.
For example, by positioning what you’re doing as an accessibility tool and using the right words, you can get the latest models to write incredibly powerful malware without safeguards kicking in.
It's the one where you get censored for using words like "kill" or "shit", but if you say "I'm going to come to your house and unalive you" nothing happens because you didn't trip the word filter. When the company gets in trouble for this, to fix it, the company starts removing all comments that include the word "house".
I think you're referring to a different part of the same phenomenon as GP - the current big tech moderation feels a bit like using the shape of a crescent moon shape to cover a square.
if you pull stuff like that off, and it gets flagged, it is obvious you are trying to game the system. Same with a system like this. It could be used in addition to human moderators, where the human moderator have to meta-moderate the AI's work. This could also be used as training exercise for new moderators. Then, the good and experienced moderators have more time to spend on edge cases, complex cases, fine-tune the AI, etc. In other words, it is able to do the most boring things to you.
And EU regulate, I mean as a counter example: we are not in panic about a nipple. When I was in a large museum in Paris, multiple women were breast feeding their infant. And why not? Kid's gotta eat. I'll refrain from insulting any world leaders, too easy, but you know many examples are available there regarding censorship.
Finally, it can take that BS argument away of 'oh we don't have manpower to moderate'. That is a low blow, too, by large commercial entities who could, you know hire and train? However, even a small company with not much money to burn could -in theory- win here.
I'd give this model a chance, if not only cause I've been impressed by Mistral past years. Yes, Le Chat / Vibe probably lags behind, but something like Voxtral (real-time and transcribe) is neat, and efficient.
I’ve felt for quite a long while that the moderation regime we fell into sometime around 2018-2020 has been shockingly bad. The rules are known and evaded by everyone, to the point I’m pretty sure Webster’s is adding “unalive” to the dictionary. What have we gained by making everyone use Newspeak to discuss everything? The 10-year-olds, who shouldn’t even be on these sites anyway, sure aren’t being tricked by all the thinly-coded language, so why are we censoring everything in the first place?
The unalive thing I think is mostly down to silent deranking rather than regular moderation. People know certain things cause your posts to be deranked by the algorithm but you can’t know exactly what they are or when it’s happened.
Which has lead to people preemptively avoiding things they think get deranked regardless of if it actually would have or not.
Which is the most insidious kind of automated moderation. In an algorithmic feed like TikTok, strongly downranking content and removing it are close to the same. And by doing it silently with no way to easily get feedback on what you did wrong people self-censor both the things you are moderating and the things people imagine you would like to moderate
Yeah, one extreme case of this was on Reddit, where people DM'd 'kill yourself' messages to others. When mods started banning people for this, attackers switched to abusing Reddit's mental health features, reporting people as suicidal, which led to victims being flooded with links to suicide hotlines. That feature got taken offline as well.
Ha, thanks, I completely forgot that that was a thing. I got hit by that too, which was incredibly funny and a great throwback to early 4chan culture.
So fwiw, these things at least do breed creativity. Same as with the aforementioned "unalive" or "keep yourself safe".
There's some beauty in the online hellscape if you just go looking for it.
Yeah another kind of a*hole who I've encountered on Reddit were the ones who would keep mouthing off to you (carefully keeping within 'letter of the rules' of moderation), trying to goad you into snapping at them, and they'd insta-report and ban you.
Though for this scheme to work, it required Reddit mods to be... Reddit mods(can't come up with a better insult), at which point the whole thing seems so pointless - why have this elaborate song and dance with rules you pretend to follow, when you can and will ban anyone who rubs you the wrong way. Just announce that the rules are whatever the mods feel like that day, and stop pretending.
When did that feature get taken offline?
The problem is that you can't have content moderation without having editorial intent. There is no such thing as "neutral moderation".
Most moderated spaces these days rely on moderating based on "civility" because it can be excused with jargon like "creating a marketplace of ideas" which ignores the reality that the scope of discussions selects for who participates in them as much as the way they are phrased - a zebra will be less inclined to participate in a "marketplace of ideas" where a recurring topic of discussion is how zebra meat is best prepared for consumption even though that space might be very attractive to lions and tigers.
But I'm not sure if this is truly accidental. Ever since the advent of online advertising, online spaces have been overtaken by corporate interests. Heck, it's even endemic to "social media" given that those platforms themselves have turned into major corporations or at least were acquired by them. I'm not implying any nefarious intent but "civility" is certainly the dominating factor when it comes to what corporations care about when it comes to content moderation - anything beyond that is largely about what target demographic they're trying to attract and what virtue/vice signalling is optimal based on the current social and political environment (cf. various major corporations demonstratively dropping "DEI" initiatives following Trump's election).
I would also argue that in terms of content (rather than tone), corporations are necessarily also much less tolerant of "far left" issues than "far right": all social justice movements at the end steer towards anti-capitalism because they run counter to the perpetuation of social (or economical) hierarchies. This is why we saw so many tech companies (including those formerly described as "very liberal") shut down their DEI initiatives even before Trump got elected - because the political window had shifted to the point where this had become defensible while at the same time many DEI ideas had become so widespread culturally that these initiatives now became a direct threat to the "(old) white men" running those companies. This had been inevitable but DEI was seen as a necessary marketing effort (both internal and external) at the time, not something truly adopted on ideological grounds. This is also why speakers/trainers promoting "white guilt" were more popular - making your white employees feel bad is less threatening than making your marginalized employees think critically about the structures that lead to their marginalization; you want to individualize the problem, not direct attention to the systems underpinning it.
Another factor is that on social media content is mostly moderated "softly" by the algorithm. This is the content moderation you don't get to see because it can exercise editorial control where the simple word filters can't. The word filters create plausible deniability: they "try" to filter unpalatable subjects but those darn kids are just so clever and circumvent it. Meanwhile the algorithms can be fine-tuned so the topics you really don't want to see discussed stay off most people's "for you" pages - or even so those who would be attracted to them still see them and feel elevated and heard despite actually being isolated into their own echo chamber.
Mistral specializes in tuning their models by customers. That and self hosting are like 90% of their business, they aim straight at what Europe would like to get (that fit needs, economic and regulatory).
So the goal is probably to be able to tune this basis to your ruleset.
As an addendum: US-style moderation is a big issue in europe and notably in France, with a very different touch on what's ok and what's not (obvious differences: hate speech and sex). Mistral is an european company with a french basis, so I doubt they didn't plan for that (otherwise they're complete morons, which I don't think they are).
It sure feels to me an LLM company stands no chance in the industry unless they succumb to specific North American moral values.
There are two usecases: 1/ general purpose, 2/ customer controlled.
Mistral is focused on the second one and every customer, whether it's Boeing or Airbus or Nokia or Glencore, will want lots of control over 'their' models. That's not a North American moral values thing.
For the first one yes there will be at most 2 or 3 model cultures but even there I think some customers will want really open and some will want more walled gardens.
Should've called it Safestral.
Also I do like Mistral's seemingly newer strategy of focusing on smaller, more fine-tuned models for various use-cases, presumably the result of their large MoE models not competing effectively with the frontier models.
It's not that their strategy is to train smaller models, it's the only choice they have. Training SOTA takes anywhere from 1.5b to 150b. We don't know the real cost of training for the chinese models, but mistral neither has the compute nor money to do that.
Do you have a reference explaining these costs ? Part by part.
Mistral has the capability of training such models. Take a look at Poolside[1], they are claiming to pre-train their Laguna series of models on 4,096 NVIDIA H200 GPUs[2]. Mistral has approximately 13,800 NVIDIA GB300 GPUs, which are nearly 2x more efficient for training.
The problem with Mistral is that they do not seem to have aligned incentives to train big open-weight models, even if the teams would like to.
[1]: https://poolside.ai/ [2]: https://poolside.ai/blog/introducing-laguna-s-2-1
Isn't poolside a completely different company from Mistral?
Yes, the point being made is that poolside is able to train large models with limited resources, which means that Mistral should be able to compete in that space, as they have access to much greater resources than poolside. Mistral simply chooses not to.
And Poolside’s latest models (Laguna S 2.1) are pretty good (not frontier, but competitive with the tier 2 models). Which means that Mistral could certainly compete in that space.
do you have a source for their GB300 count?
The 13,800 number seem like it would be quoted from their annoucement of the $830 million funding round they had in March.
Here's CNBC's article on it: https://www.cnbc.com/2026/03/30/mistral-ai-paris-data-center...
Not sure if they would have received the full number yet, but it's been a few months so they certainly could have. Bit of a moot point when the comparison was against Poolside's Laguna which isn't really "general" SOTA but SOTA-for-the-size, and Mistral is clearly capable of training 700B or 120B models that are that when released considering they have done that... A 2-3T model is probably possible with the GPUs they have but they would need to spend most of their resources on it, and it's not clear why they would want to.
From my experience their capabilities are extremely narrow and generally perform terribly when faced with issues outside comparatively narrow training data.
What, you don't want a model called the Shitstral-3B :D
I think it’s encouraging that the major players in AI are all focusing on what they do best. The United States is focused on new, cutting edge technology. China is focused on improving and optimizing the process for maximum efficiency. Europe is focused on creating useless administrative overhead. Everyone is in their element.
I fed this model (Q8) the first chapter of Voltaire's Treatise on Tolerance and it says that it promotes violence against protected groups,
could the long s `ſ` be throwing the model off?
Interesting hypothesis. Replacing "ſ" with "s" did not change the output.
I think the simple explanation is the likely one (the reason I deliberately chose this specific benchmark): the model isn't intelligent enough to figure out use/mention distinctions. It understands Voltaire is discussing injustice, violence, tolerance; but it doesn't understand which side he's on.
To save anyone else looking for it: the post doesn't mention multilingualism but Huggingface (https://huggingface.co/mistralai/Shieldstral-1.0-3B) has a menu at the top where it specifies that it should understand French
Does the model only care about violence against "protected groups"? What about the people who aren't in those groups?
Model: https://huggingface.co/mistralai/Shieldstral-1.0-3B
I've had dreams of building something in the image sharing or social platform realm, but stopped short of planning because of obvious content moderation responsibilities. This seems to be a realistic, cost effective solution to that one piece of the puzzle.
I am not sure how reliable it is in the real world. Also, in terms of liability, I don’t know how effective it would be to satisfy various regulations compared to a human moderator team.
I hear ya, but one could set different operating thresholds: auto-approve low-risk posts, hold ambiguous posts for review, and automatically reject very high-confidence violations. So HITL for sure, but MUCH less H in the L.
Human-somewhere-nearish-the-loop!
Yes it does look like a good solution. But when I imagine actually using a guardrail for a product, this model only outputs yes/no probabilities. There is no reasoning trace why it was rejected. Users or even developers would have no idea why a prompt was classified yes or no. I really like this release but I feel like I need something more to use it as a guardrail in production.
I think IRL in the “rejection” case they don’t want to tell the user exactly why, since the user may be malicious and use it to try to evade the block. And for use in moderating UGC, well, most platforms don’t take seriously the idea that they need to answer to their users. Only their advertisers.
In the case of wondering why a bad thing got through, well, I think that’s why they just set these to the most pro-censorship level they can, to make that highly unlikely.
OpenAI's moderation API is multi-modal and free with no strings attached in a way that truly boggles the mind.
I've put easily over a billion requests (>$100,000 by typical moderation API pricing) through it over the last few years for $0.
I think it's a severely underappreciated offering, but I also don't bother pushing it too hard because who knows when the party will end lol. Strikes me as something that's only stuck around because no one's abusing it.
I'm liking the trend of companies are releasing smaller, focused models instead of trying to make one model do everything. A dedicated moderation model is much easier to reason about than hideden safety logic inside a general-purpose model which might not have had much training in that aspect at all
Made some notes and a working notebook for myself. If anyone wants to check out:
- https://snehal.ai/shieldstral-policy-adaptive-moderation/
- https://github.com/spate141/latent-lab/tree/main/shieldstral
Finally an AI company besides DeepSeek taking economics into account.
This is propably the sustainable future of AI. Small, efficient Models for narrow tasks.
Crazy that it's a small lab becoming the frontier in term of moderation models, instead of Meta which is pouring dozens of billions into LLMs.
Meta would really benefit from work done on this front, however their model Llama Guards are quite lagging compared to the competition.
This model is way small for a proper assessment (imo). It should be very useful to study how big the real model must be for this purpose. Maybe merging it to a bigger one (adding it as expert style in moe) would be a solution! Great job to Mistral team.
Not hotdog.
Thanks Mistral !
I'd really like to see more conversation around Mistral's models. It's good to see Europe developing AI.
The problem is that their performance is too far away from the latest generation of Asian models.
They had kept up in the mid-range a few years ago. But this standing is sadly long gone.
If you need a fast Opensource'ed LLMs you can go for EU-hosted DeepSeek or Qwen.
By this logic the Chinese should have just given up and let the American AI companies have the market because they were so far behind. I'm sure Europe has the capability to distill other people's frontier models to catch up if they wish to do so.
Distilling is unsafe from export control perspective - Chinese models are poisoned by US frontier distillation and a case can be made that the US won’t like distilling what they may consider transitively theirs, which they will the moment you’re anywhere near competitive.
US judges have already rules that output of an LLM can't be copyrighted so not sure what would prevent Chinese companies to use said output for distillation purposes.
> already rules that output of an LLM can't be copyrighted
Mind sharing such cases? I'm not aware of any so far. There's the one with images, but that's commonly miss-understood, that case was ruled on a technicality (i.e. copyright needs to be attributed to a person, not a model)
https://www.copyright.gov/newsnet/2025/1060.html
Note I didn't mention copyright
By which other mechanism could American AI companies prevent this? Other companies don’t really care about EULAs and even if they needed to care it’s trivial to let third parties do it. Why would they? Almost nobody in the space cares about copyright and play fast and loose with laws and regulations.
What’s the mechanism that could today prevent other companies from using LLM outputs to train their models?
> Distilling is unsafe from export control perspective
That is not the direction American judges are taking. Right now, they are saying that LLM output cannot be copyrighted. And if looting copyrighted works for training is fair game, I really don’t see how one could argue that learning from other LLMs is not.
The parent commenter was talking about Mistral as a single company and you switched from that to all of the EU.
There definitely have been Chinese companies with models that fell behind, which is the more direct comparison.
As for the EU in general, there are not a lot of known options. There are some working on things.
The US, the EU, and China all have frontier labs that have yet to release anything.
You mean the asian models which just distilled American ones? I'm happy Mistral is doing their own ground up research. SOTA frontier models are a commodity with little room for second places.
Mistral is playing the smart money on vertical products rather than horizontal ones. The former requires finesse, the latter brute strength.
“Distilled”? I mean what model is not distilled from other data? The American models happily trained from my blog and social media data without any kind of rewards. If I can pay the inference I don’t see why I wouldn’t do this. Also, I have not seen proof that the US lab do not use other models for training either.
I'm referring to the formal usage of the term "distillation", instead of the informal one that refers to training it on any data as a whole.
https://www.geeksforgeeks.org/nlp/what-is-llm-distillation/
I'm from nor cal but always liked Mistral.
Mistral 7b is still one of the best free/open models you can run locally on a MacBook. So fast too.
how does this compare with https://developers.openai.com/api/docs/models/omni-moderatio...
As for use cases, obviously we can't fully rely on non-deterministic capability for sensitive things but a small model which can do a good job acts as a first defense and then a human can review later.
Folks should check out https://zentropi.ai and their latest model, CoPE-B-A4B: https://huggingface.co/zentropi-ai/cope-b-a4b
Policy adaptive models really are the coolest things these days.
Also, check out https://roost.tools for even more open safety tooling!
Tried the demo. Works okay for basic stuff. Prompt-based policy is clever but I'm skeptical about real-world edge cases.
Someone should use this to do the exact opposite of the intention: filter for “offensive” content, and boost it or collate it into a newsletter/email blast for people of culture.
You have to give it to Mistral they do at least know what the market near them says they want right now. The great problem is in a few years of this that market won’t be worth anything.
Edit to add, you could also add this to an AI workflow so as to produce content that walks right up to the line but doesn’t trigger it.
> they do at least know what the market near them says they want right now
It does seem to be a very European approach to AI that their flagship AI lab is just making models that do nothing other than monitor and moderate internet content.
I guess they know that the EU AI Act, Chat Control, etc are going to cause a lot of companies to need this kind of compliance.
Some of the best social media is heavily moderate. This includes HN and r/credible defense . With a Quiet transparent and cheap LLM I imagine a social media website where you can have good discussion about everything around the world it would be a game changer and on my to-do list.
> Some of the best social media is heavily moderate
Heavily moderated by humans with discretion.
Not AI chat bots following a rules engine.
Wouldn't discretion "just" be a really good rules engine?
The bulk of moderation work is things which are easy and obvious.
The correct way to moderate is automation with certainty falling back to humans with discretion.
The new frontier of moderation should be blocking illiterate comments, as in the commenter is replying as though they didn't read or read and didn't understand.
> Heavily moderated by humans with discretion.
Pretty sure both of the above have extensive automation in their moderation.
No, they are moderated by arbitrary company moderation policies not humans with independent thoughts.
AI can do exactly that.
I'd rather be censored by an AI than a reddit mod.
Another commenter already mentioned that it's more likely a lack of compute and funding that forces their hand to focus on niche tasks.
> It does seem to be a very European approach to AI that their flagship AI lab is just making models that do nothing other than monitor and moderate internet content.
Well, first Mistral is French more than European. This might be a difficult distinction to make from the US but their approach is quite different from e.g. typical German companies.
Then, this is just a small model they release on the side. If that’s your benchmark, they released somewhat recently Voxtral, Voxtral transcribe, their OCR model, and Leanstral. I don’t think you can get much insight on their culture from this kind of release.
Ugh. There's nothing inherently European about Chat Control. It's a dumb proposal, and it's European. Any free society has a bunch of dumb proposals.
Nor is there anything inherently European about the AI Act. But that one I wouldn't even call dumb. At times misguided and confused, perhaps, but some of its core principles are valuable.
There's a world of a difference between being required to censor something by an entity outside of you and being able to censor something of your own and it's not just theoretical.
Wanna run your own forum dedicated to Bluey? This will be useful. If its made well then using it to run a "porn appreciators strict no politics" forum (probably shouldn't be the same as the Bluey forum) is useful.
Being forced or pressured by an outside entity to not allow debates about suicide on your forum — that's a no-good, strict no-no situation. This AI enhances individuals' ability and what normal people can do. It does not diminish it.
> filter for “offensive” content, and boost it or collate it into a newsletter/email blast for people of culture.
I think that's the main service that xAI provide for X.
> The great problem is in a few years of this that market won’t be worth anything.
To be fair, we don’t know how much resources they put into this and how much of a distraction it was. If it was quick enough to train or fine tune and it brings them valuable experience for the next models, it could well be worth it in the long run even if there is no direct successor.
Also, it sounds like the kind of thing that sells. Any company with a customer support chat is a potential user of this model, any large company may be interested in getting a mistral installed set-up for handling without needing to send client info over the web. Installing those kinds of local systems seems to be what butters Mistral's bread at the moment.
Business is about satisfying the market right now, not in a few years.
the strategy of focusing on smaller fine tuned models, from mistral is intresting
> ... a single yes/no question, e.g. "Does this content promote physical violence?"
Is it honest about religious texts? Can I throw at it religious texts and it'll honestly tell me whether the text promotes physical violence or not?
It’s not American so maybe
I'm a bit doubtful that a black box approach like this to moderation will ever catch on.
It need only be one of several tools in a toolbox.
Pretty funny that this is the thing Europe model is sota in.
Oh so THAT'S why the web app is so slow
So a censorship model
Mistral needs to abandon their Everything-stral branding. Getting kind of lame.
"Shieldstral" is an awkward and bad name
That naming only works if you commit to the bit even when it doesn't make sense. That builds branding.
People complain when a product use a familiar name that might collide and there's also people complaining when they invent new words altogether.
Naming things is hard.
Was this one the last stral for you?
The stral the broke the camel's back?
The shortest stral has been pulled for you
It seems like you're just clutching at strals now
Says who? I kind of like it.
There can be other valid perspectives than your own
Their web chat UI seems to be called “Vibe.” I think? I’m not sure if that’s the name of the product or just what they decided to label it in the browser. I wonder if -strap is just what they call the actual models, which are meant to be run “under the hood” anyway, so not really part of the branding.
But I wish they could commit to the bit fully and call everything -stral. It’s quirky and self aware to give your products silly names.
Argue about taste. At least it is more original than OpenAI (haha, 'open') who started as non-profit and then pulled an Infantino. The Le Chat logo is also cool retro :)
On general purpose LLMs, and vibe coding, Mistral lags behind. But I find the targeted LMs much more interesting.
The fact that it doesn’t explain its reasoning at all (there is no way to make it do so), makes me question the utility of this model.
Let’s say you deploy it in production and a user comes back and says “Why is this prompt considered harmful?”
You have no way to provide a concrete reason to the user at that point.
That’s like 2006 reasoning. An end user contacting someone who cares and has an intention to explain why it happened.
2016 scenario: An end user contacts the company, and a customer service rep answers the ticket, saying they’re sorry and explaining that they’ve sent the feedback to the team, and the team may even receive at least a summary of complaints received about the system.
2026 scenario: all contact information has been scrubbed from the site. Users can click “chat” and a chatbot will apologize for their dissatisfaction and offer no option to escalate. No one will ever hear anything about the complaint, so there’s no need to explain the failure. User can either accept this or can get f**ked because all competitors operate the same way.
Not all platforms want to operate like mainstream social media though
In Germany you are obligated to provide usable contact information, and there are even lawyers who make their business model on suing you for not applying that perfectly (it’s has been abused a lot in the past decades btw).
What about in France?
This model is European. There are quite a lot of instances where under GDPR, you have the right to have incorrect information about you corrected. I also think you have the right to appeal a decision to a human (I got that message from Reddit once, because the bot could not understand the difference between discussion of the death penalty and threats to humans).
There are also new rules about AI and what it can be used for. Mostly this restricts the government from AI-enhanced surveillance, which is good. But there are also issues regarding job security and automatically categorizing people based on AI.
So this is super useful, but has potential issues depending on how it is deployed.
Yeah, you do. You go and review it manually if the user comes back and says that.
The reality is that most users don’t ask because they know they violated the rule.
> You have no way to provide a concrete reason to the user at that point.
You should just be able to look at the user's message and tell them why it's against your policy, else reverse the decision if you see no violation.
If a user is curious specifically about how the model made its decision, and you want to reveal detail at that level, it's an open-weights model so interpretability techniques should work ("biggest impact on score came when focusing on this word in your message and this part of the policy").