My 7 Best 7-8B Models For OobaBooga – AI Roleplay & Chatting

Here are the absolute best uncensored models I’ve found and personally tested both for AI RP/ERP, chatting, coding and other LLM related tasks that can be done locally on your own PC. Of course it goes without saying that these will work with any other large language model software, however I’ve ran all of them using the Oobabooga WebUI and KoboldCpp/SillyTavern. Let’s roll!

Looking for the list of 24B models instead? – You can find it here: 6 Best 24B GGUF Models for 24GB VRAM Local AI RP

Updated – August, 2026 – all of the original models and my notes on them are still here, as they remain surprisingly relevant even today, but some new worthwhile ones I’ve tested have also been added (Anubis Mini 8B v1, Stheno v3.2, Lunaris v1 and Kunoichi DPO v2).

7B, 8B Models? – What Does The “B” Stand For?

The “B” you probably keep on seeing in various large language model names is short for “billion” and it refers to the amount of parameters the model has. In this case it means that a model has 7 or 8 billion parameters.

In general, the more parameters an LLM has, the better they can perform in various more complex tasks (because of their greater capacity for learning), however they are also larger in size, and depending on their precision level, optimization methods at play and a few other factors, do require more VRAM to load up and use.

A Q4_K_M (medium quant) version of an 8B model is usually around 4.9GB, so it can still fit on many 8GB graphics cards with room left for context, provided the context window is capped around 8k tokens size to prevent operating system overhead and KV cache from triggering out-of-memory errors.

This is a part of the larger LLM naming convention in which the model names are created from a few connected parts most commonly including: the model family, model’s current version/iteration, parameter size and quantization format.

Although not all model names follow this convention, it is basically agreed upon and widely used when publishing trained large language models online.

Need a GPU upgrade to run larger, higher quality models with OobaBooga or KoboldCpp/SillyTavern? – Here is my curated GPU list for LLMsBest GPUs For Local LLMs This Year (My Top Picks!)

What Were My Criteria? – Why These Models

Oobabooga WebUI running a customized AI character chat.
An example conversation with a locally hosted 7B language model in Oobabooga’s chat interface.

This list goes with 7B and 8B models, as these are among the most used by people without access to larger amounts of video memory, and in many cases can even be loaded on older 8GB VRAM GPUs. Moreover, on graphics cards with more video memory on board, they can be used to make use of the larger context windows and achieve longer conversations.

All of the original models listed were used by me and tested both during casual conversations, roleplay scenarios and simple logic/language processing/text formatting tasks, and their performance turned out to be reasonably good.

Remember that ideally you want to test out a few different models to see which model’s response style fits your RP/chatting style the best. The outputs for the same input prompt may largely vary between different models.

If you want direct benchmark tests, some of these models do feature benchmark scores on their main Hugging Face repository pages, so feel free to check that out. The UGI Leaderboard is another great reference point. Set its #P filter to between 7 and 8, open the Intelligence view and compare the Pop Culture and World Model scores.

Pop Culture tests details from games, movies, music and internet culture, which helps with characters based on existing fiction, including of course various anime franchises. The Writing view adds style and repetition data, and W/10 shows how far a model can be pushed before it refuses or drifts from the prompt.

Is The Output Quality of 7B Models Any Good?

7B model vs. popular large language models in terms of roleplay, writing, STEM and other testing categories.
This, now very much dated 2023 Zephyr 7B chart still shows how one good 7B fine-tune could beat much larger models in selected tests.| Source: Hugging Face, zephyr-7b-beta repository.

As you might probably know, larger doesn’t always necessarily mean better. But still, it often does. One could argue that a smaller model trained on a better dataset could perform better than a large model trained on half-garbage data, and probably would be right.

The reality is that for less complex tasks like roleplaying, casual conversations, simple text comprehension tasks, writing simple algorithms and solving general knowledge tests, the smaller 7B models can be surprisingly efficient and give you more than satisfying outputs with the right configuration.

So let’s get to the list, there are quite a few models to choose from here!

The Four Small Models I Would Try First Today

Here are my recent additions to the updated list, which consist of four later releases that are in my eyes one of the best performing small models out there for RP/creative writing purposes.

1. Anubis Mini 8B v1

Anubis Mini 8B v1 is the one genuinely new small model added with my recent update to this lineup. This Llama 3.3-based fine-tune handles detailed instructions, sticks closely to character cards and can adapt its writing style to your prompt.

The early feedback is unusually good. Better instruction following, higher general intelligence and character adherence, and a few other advantages it has over the older options present here. Stheno still has a distinctive way with words, so Anubis is not a straight upgrade in every chat, but it’s very much worth trying if you want the absolute newest pick you can get.

2. Stheno v3.2 8B

Stheno v3.2 one of the easiest small models to recommend for one-on-one RP and character cards. It can give you detailed scenes, realistic dialogue and good instruction following for its size.

The praise it got in various online communities did not stop after its release, and it’s still very popular. You can read one of its rather detailed user reviews here.

3. Lunaris v1 8B

Lunaris v1 is a Stheno-based merge made to keep its creativity and improve its logic. The difference is easy to notice. Lunaris tends to write longer narration, move the scene forward on its own and handle a small group better, and I’ve confirmed that to be the case in my personal testing.

It does have a smaller user trail than Stheno, but it’s a great alternative if you don’t quite like the style or mannerisms of the original model.

4. Kunoichi DPO v2 7B

Kunoichi DPO v2 is the one I would try if you still want to stay at exactly 7B for one reason or another. It is expressive, easy to steer and still has a surprisingly loyal group of users.

Another flavor, another low-VRAM option for you to try. And at the same time a great addition to our list.

What About My Original 7B Picks?

These are technically “older” picks now, but keep in mind that the creative writing/RP models do age in a completely different way than models meant to keep up-to-date knowledge and state-of-the-art reasoning capabilities.

All of the models listed below are still, in my eyes, great picks for local RP sessions, of course bearing all the usual limitations that smaller 7-8B models have.

1. Dolphin-2.8-Mistral-7B

Dolphin-2.8-Mistral-7B is an uncensored Dolphin model based on Mistral with alignment/bias data removed from the training set, which is designed for coding tasks, but really excels at simulating natural conversations and can easily go beyond that, for instance acting as a writing assistant.

I was really surprised how good this model does with various roleplaying tasks and character-based chat sessions considering it was originally meant for solving simple programming tasks. Overall, still one of the best hidden gems I came across. You should definitely try it out!

2. Wizard-Vicuna-7B-Uncensored

Wizard-Vicuna-7B-Uncensored is a LLaMA-7B model trained using a subset of the Wizard-Vicuna dataset with any responses which could contain moralizing or alignment of any kind were removed to avoid any unwanted “guidance” from the model when generating responses, which would otherwise be subject to the classic on-the-go judgement from the AI during your chat.

This unfiltered model is commonly recommended for roleplay use, and being honest it’s up there when it comes to the output quality. Together with the OpenHermes-2.5-Mistral-7B and Dolphin it was among my 3 most used 7B models for a pretty long time. All in all, another one really worth testing out.

3. Yarn-Mistral-7b-128k

Yarn-Mistral-7b-128k. There isn’t much to say about this one, other than it’s a really solid 7B model which can be used both for casual AI character roleplay and some more serious purposes. It’s a direct extension of the base Mistral-7B-v0.1 model, and it supports a rather large 128k token context window.

Good for RP, great for most of the tasks you throw at it. Works well with guided character data and is in general a great choice when it comes to higher quality 7B models out there. Not censored, but also not further mixed with more explicit datasets.

4. OpenHermes-2.5-Mistral-7B

The OpenHermes-2.5-Mistral-7B model is one of the most popular fine-tunes of the original Mistral model. Being another flavor of Mistral, it also inherits its main features. This model is trained on additional coding datasets, however it also passes various non-code benchmarks with great scores. While being one of the older options, it still performs pretty well in RP contexts!

If you didn’t really like the base Mistral model or the YaRN version, be sure to test it out – the outputs from this one are significantly different. Now let us continue. Beware: more “heavily uncensored” models ahead!

5. Pygmalion-2-7b

Pygmalion-2-7b is based on the Llama-2-7B model from Meta AI. It is meant to be a model designed and fine-tuned specifically for roleplaying, casual conversations, storytelling and writing assistant use. And it does its job really well.

Trained on a large amount of RP convos, short stories and conversation data it is specifically prepared for using with character data. Its first version was one of the most popular open-source large language models made for private roleplay sessions. It is also fully uncensored. Not much more to say here!

6. Starling-LM-7B-alpha

Starling-LM-7B-alpha, fine-tuned from the base Openchat 3.5 model which was originally based on the first version of the Mistral-7B is also one of the most popular picks for local LLM use. In many benchmarking categories its scores are surprisingly close to GPT-4, and it is generally a very solid model.

While it isn’t my first pick for AI roleplay, it’s a solid model and can be rather easily steered towards a natural simulated conversation. Expect quality outputs in the style of other Mistral-based models with the right guidance. Yet another great pick.

7. Toppy-M-7B

The Toppy-M-7B is a model with a rather telling name often recommended when it comes to ERP, and for a good reason. Based on the uncensored merge of a few different models and LoRAs, this one is definitely worth trying out, especially if you’re into generating some more spicy roleplay scenarios.

That’s pretty much it. Definitely not my first pick, but I can see how many of you may be interested in it. As a fun challenge you can always try and force it to do some light coding for you. Test away!

8. Loyal-Toppy-Bruins-Maid-7B-DARE

Loyal-Toppy-Bruins-Maid-7B-DARE is the last model on this list, which is a DARE TIES merge of four Mistral-based models. Yes, this name is also real. And yes, it is a model trained using the previously mentioned Starling-LM-7B-alpha, the Toppy-M-7B and a few different models as a base. It’s optimized for roleplaying but can also perform many other tasks with satisfying output quality given the right prompt.

This one is quite an interesting merge, and it can give surprisingly high quality results considering its funky name and merge source models training material. And with that, I’m out!

Check out also: Character.ai Offline & Without Filter? – Free And Local Alternatives

Tom Smigla
Tom Smiglahttps://tomsmigla.com/
Tom is the founder of TechTactician.com with years of experience as a professional tech journalist and hardware & software reviewer. Armed with a master's degree in Cultural Studies / Cyberculture & Media, he created the "Consumer Usability Benchmark Methodology" to ensure all the content he produces is practical and real-world focused.

Check out also:

Latest Articles