Here is the "memo". It's not a secret.
I wrote this guide myself a time ago on another platform, describing my tools and models i use to make the game. Maybe it is useful for you.
------------
I myself started with some techniques that are completely outdated by now. You should skip these.
But to really start a VN (with Videos and animations or without), you should get familiar with a few tools first and also think about the basics.
Once you got the hang of it, you can work out a strategy and toolstack that works towards your stylistic idea of the game.
The Seven Seas has been developed with a highly dynamic model landscape beginning with very poor first-gen models, and by now development has slowed down a bit at least for gamechanging open source stuff, and the existing tools are much better and mature, allowign us to cherrypick them for stylistic consistency.
Hardware
First of all your strategy should adapt to the hardware you have, unless you have the cash to adapt the other way. I do all my generations and also some training locally on a 3090. Not the newest model, but generous 24 GB VRAM are still more than most mid range graphic cards until today. Do you happen to have at least 16 GB of VRAM, ideally on a NVIDIA GPU? That makes things a lot easier. If not, there are still options.
Secondly RAM on your local AI generation machine is important: RAM is damn expensive, but since i upgraded from 32 GB to 96, i got rid of a lot of trouble. 3090s are still not cheap, and for a reason, but maybe you can get lucky and get one, or you are a rich kid and get a 5090 or an RTX6000 ideally.
Then also, i would advice to invest in a SSD not smaller than 2GB. Once you start downloading models and Loras, you immediately stack hundreds of GB of model data. Most HDDs are way too slow. Don't be cheap on that or it will be a bottleneck.
General AI Toolset
ComfyUI is king and the most important tool for local AI. It has a learning curve, but it is the most versatile tool, able to generate images, video and even sound, and there are lots of workflows being shared around. It is the "industry" standard. There are tons of good ComfyUI tuts on Youtube. Unfortunately, to even install it, you need a little toolstack (Python, Git etc.). Don't know what your previous experience with these is, but you don't need to be an expert there. Most of the necessary stuff is automatically installed, but the installation and keeping it up to date can be challenging for a complete beginner.
A few remarks:
- Important: get the ComfyUI portable version. The fix installed version will ruin your day.
- install the ComfyUIManager. It is a fine tool to look up missing nodes, addins and models and automatically download them
- To get models and especially Loras for Comfy, you should use Civitai or Civitai.red. For some older models, there also is a variety of Loras shared around here, or you even dive into training you own models, but that is another story
Models for Images
Right now i see the following options. For each of these, you will find really good tutorials on youtube. Of course not on how to do smols with them. You should decide which style you are aiming at. Pixar style? Anime style? Something semirealistic/artistic? The hardest part is to maintain consistency between different models. I would not advice to mix them too much. Only the very best models can adapt to styles you feed them with, the older ones will dictate their style on your image.
-
Pony XL / IllustriousXL models: B-Tier Still often used. Excels at hardcore scenes. Very little intelligence about how the worlds outside looks like. Very bad prompt adherence. You need to know the keywords to prompt it. Still useful for Anime styles and for hardcore gangbang scenes for which the models below are not trained out of the box.
- Flux 1S/Flux 1D: B-Tier: The first intellgent model. Will do nice backgrounds and complex scenes. Not great for characters or NSFW. I use it rarely, but sometimes it is a good addition for some very specific scenes like an archeologic site or something like that.
- Chroma: A-Tier and secret tip: A mix of Flux and Pony. Combines strengths and some weaknesses of both. Does a fair amount of NSFW stuff even without Loras. Maybe the best model to start with. High prompt adherence, a very intellgient straight forward model
- Qwen and especially Qwen Image Edit 2511: Highly recommended. Highly intellgent model with highest prompt adherence. Will create the most complex scenes you can imagine with ease. With Qwen image edit you can use Character images or poses as input and just tell it how to combine it or use it to create something different from it. Sadly the number of Loras is quite low, but there are good ones out there. I could not do without Qwen Image Edit, it is critical for me. I trained some Loras on my own for it. It can adapt to various styles and is the secret solution for all consistency issues! Also it can create keyframes for video generation just great!
-
Krea 2: S-Tier. Just came out recently, Loras are already exploding for this model. Easy to uncensor. Capable to depict very lewd scenes in fantastic stylization. It requires proficiency with Loras, but it is about to surpass Chroma for me, especially because the Lora makers love it and support it a lot, which allows to create hardcore scenes and fantastic styles.
Video models:
IF you also want to create videos/animations: On the open source/local market, there are just two worth your time. I would advice you to focus on Image2Vid or FFLV (first Frame Last Frame) approaches. FFLV is better for consistency and control, but sometimes you only have a start image. You can then generate a video and just pick the final frame to generate the next section. This is not 100% ideal, but with video editing one can make it work.
Consistency also here is a big issue, therefore FFLV. But sometimes the video glides into absurd variants, or into realism that often comes with hardcore Loras. It is especially hard to keep the consistency between your still images and your videos.
- LTX 2.3: Blazing fast. Generates sound. Can generate vids up to 1 minute on my 3090 with sound. Can either create sound or adapt (dance, lipsync) to a sound you put in. The more modern, fresher of the two. It also is slightly better at keeping the consistency than WAN. But strangely it lacks on some sorts of hardcore scenes.
- WAN 2.2: Still the Open Source King. May be the last WAN model that is open source for quite some time. Tons of Loras, good ones and bad ones. It is highly intelligent and a huge model, it can make a ship move around, it can make a whale swim around, it can create a thunderstorm. And it is, with Lora Support, better at NSFW/hardcore scenes. But it takes about 2-3x as long as LTX to generate, video length and resolutions are more limited, and it can not really handle sound well.
- Minimax H3: Very intelligent model. The newest of the bunch. Extremely uncensored from the start, capable of creating both video and audio in one go just as LTX. Right now just a few Loras are available, but for many scenarios, the base model is insanely good. With new Loras and workflows coming out, this model might be the upcoming king very soon, and it is worth a try if you have a decent amount of VRAM (16 GB or higher). It just is a bit on the slower side right now.
So i still combine both TX and WAN often, but i also experiment with Minimax H3. If you want to start with one, i would suggest to pick LTX 2.3. There are also great tuts on youtube. For LTX, it is important to keep your PyTorch up to date or else it will be slower and kind of instable.
For both LTX and WAN, you can also work with lower VRAM, WAN does not make much sense under 16 GB though, while LTX can also run in normal RAM mostly as long as you have enough. So for lower VRAM like 4-8GB, LTX 2.3 also is the better option.
VN Game Engine
Renpy is the classic and well supported and accepted. Unity can be an alternative. I sticked with Renpy after trying it. If you have a preference on some special programming language, it might be worth to look for a VN techstack that uses it. But for me, Renpy works fine.
It is important to also learn a bit on how to use and initialize it, or how to play sounds and videos, how to use sprites or even animated sprites. Renpy is based on Python, and python is fairly easy. AI can also help you do some programming for Renpy. I used that for example on the ship game.
Sound generation
Suno or Udio are way better than the open source alternatives i tried for making music. Unfortunately both platforms started or ar about to start fingerprinting.
For spoken Language, Qwen3 TTS is a great tool to generate or replicate voices. You can put in a sound file and create new speech that sounds the same, or you prompt the voice and mood you want and then feed your text in. Works quite well.
You can also use open Soundpacks like
You must be registered to see links
for some hardcore moaning and such.
Tools for Video editing:
- Any tool you like fits. I use Davinci Resolve, the free version is very powerful and easy to use. You might also need a converter, i use Handbrake to convert MP4 files into WebM files that Renpy prefers