Hello! I'm still Vladimir Sadovsky, I'm still working as a toolset programmer at Nau Engine, I still love games, and artificial intelligence has not yet replaced me. Last time, I talked about procedural generation in the game industry. It's time to continue the topic and take a look at the other side of the generative medal - artificial intelligence.
The question of creating AI has been fascinating the greatest minds for a long time. Asimov's positronic robots, Ellison's omnipotent supercomputer, Dick's replicants, and hundreds, thousands of other artificial intelligences invented by science fiction writers have so deeply penetrated popular culture that the birth of real AI was only a matter of time.

It's hard to imagine today's game industry without AI, although it was still a novelty just a few years ago. Neural networks help game designers find ideas and automate routine work for artists.
According to Unity, in 2023, 62% of game studios used AI in their work for rapid prototyping, asset creation, coding, and other tasks. Moreover, 71% of developers using neural networks claim that AI speeds up and simplifies their work.

It would be impossible to ignore such a massive phenomenon. So I suggest recalling the history of generative AI, discussing its role in modern game development, and presenting what it can offer the game industry in the future.
The Underbelly of GAI
Generative Artificial Intelligence (GAI) is capable of creating various forms of media content, such as text or images, using so-called generative models.

Generative models work roughly the same way as other neural networks. They study patterns in the input data and then generate new data that resembles what they were directly trained on.
The beginning of artificial intelligence's history can be marked by Alan Turing's 1950 article "Computing Machinery and Intelligence". It posed the first questions about what constitutes machine intelligence, how similar it could be to human intelligence, and overall introduced some fundamental understanding of what AI could represent.

The real active development of the AI theme began with the Dartmouth Summer Research Project on Artificial Intelligence in 1956. It was at this scientific seminar that the first calls were made for more active study of artificial intelligence issues.
By the early 1970s, the first works related to AI appeared. For example, the computer program AARON, intended to create paintings.
At the beginning of the 21st century, progress in the field of AI was significantly stimulated by the emergence of deep learning methods. This approach allowed further research into image classification, speech recognition, and natural language processing. However, due to large data volumes at that time, neural networks were trained as discriminative (conditional) models.

In 2017, a new architecture of neural networks using deep learning - the "transformer" (or "converter") was presented, which allowed for improving existing models. This led to the creation of what is called the "first generative pre-trained converter" or GPT-1 in 2018. And in 2021, DALL-E was released - a model generating raster images based on the transformer architecture.

The latest headlines in this field relate to Microsoft Research's announcement made in March 2023 that GPT-4 can be considered a system of general artificial intelligence (AGI). In other words, almost a complete analog of human consciousness. However, this achievement is still disputed by some researchers who believe that GPT-4 does not meet the necessary level.

However, for this material, it's not particularly important how similar a specific AI is to humans. The main thing is that modern neural networks create media content in almost all possible formats.
GAI at Human Service
Among the relevant generative models are GPT, Copilot, as well as several artistic systems of artificial intelligence: Stable Diffusion, Midjourney, DALL-E and Kandinsky - a domestic neural network.

Recently, Yandex's Shchedyur (Masterpiece) joined the top three most popular applications based on GAI. According to the corporation itself, since its launch in April 2023, the neural network has been installed by almost 8.5 million users.

And how not to mention the release of a new neural network OpenAI Sora on February 15, 2024! Potentially it can become a full-fledged supplier of visual content for video hosting platforms. More information about it can be read on the developer's website, and there are already several detailed articles on Habr.
Although it takes colossal computational resources and about an hour of time to generate just one minute of FullHD video.
Unfortunately, there hasn't been a big boom in solutions from small commercial companies in this area. The reason is that such neural networks require fantastically large amounts of data for training. Giants like Microsoft or Google have been collecting information for AI training for years. And the company responsible for creating China's largest internet search engine Baidu has recently released its own GPT-like model called ERNIE, which was trained on two trillion tokens. Small studios don't have access to such a large amount of data and therefore are in no hurry to develop AI-related solutions.
The situation is similar in the game engines market. Only the largest and most popular solutions like Unity or Unreal Engine can create their own GAI. However, they have alternatives. I will tell you below about what's available on the market now and what you can already use.
AI as a game development tool
The widespread use of neural networks has only just begun, so there are few solutions specifically tailored to the gaming industry. Usually, in game development, general-purpose generative models are used. For example, you can use ChatGPT to develop an already conceived story or ask Midjourney to draw art that you can then use to create a doc or a quick prototype. However, there are several projects focused on the gaming industry.
The most interesting tool, in my opinion, is Unity Muse. Already in early access, it allows you to create textures and sprites directly in the engine interface, and for working with the generative model offers an equivalent of ChatGPT. The neural network readily generates, for example, new ideas or pieces of frequently encountered code. In the future, Unity plans to add new capabilities for creating humanoid animations, game AI behavior trees, 3D models for prototyping games.
Epic Games has not yet announced official tools for working with native artificial intelligence. However, there are already community-made solutions.
Promethean AI has been automating developer routine since 2019. This tool is designed to manage game assets and create a virtual world.

Thanks to the fact that Promethean AI combines both a generative neural network and an asset manager, users can quickly design game levels. For creating large volumes of content, an integrated assistant is provided. You can ask it to create a section for generating a large array of vegetation or a small relaxation zone for the bedroom. Plus, you can always call up the built-in asset manager and quickly collect necessary visual elements from existing solutions.
This approach differs slightly from Muse's concept. In Unity, they probably want AI-generated game assets not to need editing. While Promethean AI looks more like an analog of a visual assistant based on artificial intelligence, which allows simplifying and accelerating work, for example, of a technical artist.
Another useful tool for game developers is integrated into MANU Video Game Maker. The creators position their project as a 3D game engine that allows creating the game of your dreams without programming - but with the use of an integrated AI assistant. In perspective, it can generate even the sky almost all game content, including levels and cutscenes. But for now, everything is mostly experimental features inaccessible to the general public. However, MANU already helps with story development and game mechanics now.

It's enough to write a request and press the "Imagine storytelling" button...

..Voila! MANU presents a ready-made story in the form of a diagram with detailed comments. For example, here is what may result if you ask AI to come up with a plot for a platformer about zombies that navigate by scent.


And here's how the logic of an enemy NPC looks like, generated by artificial intelligence.


I spent several hours in MANU and can say that this AI assistant has room for growth, but it can already provide a good idea and set the direction of work. It's left to connect it with asset generation, and you can leave just one button - "Create game" in the interface. Fantasy, of course, but not too far from reality.
Let's leave our dreams aside and take a look at a project from our compatriots, which we'd like to discuss in more detail for two reasons. Firstly, it already brings real benefits to game development today. Secondly, its principle of operation clearly demonstrates the features of neural networks.
In 2019, Moscow-based Banzai Games began working on a new tool for creating animations called Cascadeur. The authors had accumulated a lot of material after working on Shadow Fight 3, and they reasonably decided that over 1100 animations shouldn't be gathering dust on a shelf. The developers supplemented the existing dataset, broke down the animations into individual poses, and trained several neural networks to predict the behavior of all key points in an animation based on a few control point values.

Banzai Games' approach is based on using neural network composition. Initially, they tried to use a standard neural network architecture with key points as inputs. The accuracy was evaluated depending on how much the predicted value deviated from the actual one.
Initially, the deviations were high - around 3.5 centimeters. However, the developers noticed that when increasing the input number of points to 16, the accuracy improved by two times. This led them to think about using neural network composition.
So, the first neural network can take in those same 6 key points, while the next one will take in 15 points predicted by the previous neural network and so on. Thus, it is possible to achieve a reduction in error without increasing the workload of the animator, who still uses only 6 points.


It would be interesting to see how Cascadeur handles the prediction of animation behavior for non-anthropomorphic characters. After all, judging by open information, it was trained on human-like animations. However, Banzai Games seems to have solved this problem. At least, the authors were able to solve the task of positioning a cat in space when falling.

Recently, on March 13, 2024, the original code was released under the name Animatable Gaussians: Learning Pose-dependent Gaussian Maps for High-fidelity Human Avatar Modeling. This method of creating animated humanoid models is trained on RGB-video and thanks to technology 3D Gaussian splatting, as well as an innovative approach to training a neural network, it outperforms existing solutions based on neural radiance fields (NeRF).

I mentioned the creation of 3D models by their text description when I talked about Unity Muse. But you can't help but notice that today there is already a similar solution called Luma AI.

The solution allows for generating full-fledged 3D models with texture and textures in convenient formats for the game industry, such as .fbx or .gltf. Luma AI also has integration with Unreal Engine, specifically UEFN.
But talking about full-fledged generators of 3D models by requests on natural language is still premature. Any more or less complex task makes such GAI generate unrelated images - generative oatmeal. This is similar to the first versions of neural networks when generating 2D images.
Although generative models for working with 2D images have already developed and now produce fairly adequate results even on complex requests. There are also specialized solutions for game development, such as PixelLab, which simplify routine work with sprites.

It's hard to pinpoint a single tool, because each one uses what is convenient for them. However, the field of application of such models is more interesting to study.
Neural networks can be used mainly for two purposes. The first is fast prototyping of 2D assets, whether it's icons for the user interface or simple diagrams. The second is quick image quality improvement. Especially since now some of these neural networks are being integrated into hardware by microprocessor developers. An example of this is Deep Learning Super Sampling (DLSS). This can be useful when working on a remaster of an existing game.

GAI continues to progress
One of the promising areas for using generative artificial intelligence is space scanning.
Some time ago, high-quality photogrammetry caused a stir in the gaming community. When Epic Games bought Quixel, the largest library of available assets based on photogrammetry technology, it greatly simplified entry into game development. You can start creating a game not with a simple set of basic assets provided with Unreal Engine, but use high-quality assets from the Quixel Megascan library and save time and money.
However, the process of high-quality photogrammetry remains complex and long. In field conditions, you can use a laser scanner, but the quality of photogrammetry will be very different from professional.
But Luma AI provides an experimental tool for 3D space scanning that improves the quality of photogrammetry by using neural networks. This method is called NeRF or Neural Radiance Field and is based on deep learning. It allows you to restore a three-dimensional space from its two-dimensional images. Unlike other methods of restoring scenes, NeRF can take into account reflections.

Interactive version: lumalabs.ai
Another promising direction in the field of generative artificial intelligence is AI for non-player characters (NPCs).
NVIDIA offers using language models as a basis for NPC communication with each other and with players. Such an approach allows you to make every conversation with NPCs unique. This functionality is presented in NVIDIA ACE demo.

While I was working on this material, a BuildBox 4 announcement appeared in the network - a game engine based on AI. The authors promise integrated generation of assets, scenes, logic, and almost the entire game by text queries. You can already register for a waiting list and participate in the alpha version.
It's pleasing to see how broad the range of generative artificial intelligence applications is becoming in the gaming industry. However, technical progress cannot exist without problems - first and foremost, legal ones.
Problems with the law
There is still no clear understanding of what to do with artificial intelligence that was trained on texts or images protected by copyright. Especially since it could have happened accidentally. For example, during web resource tagging necessary to form a context for generative AI over a certain period of time. It's not a secret that ChatGPT knows only about events that occurred up to the moment of its training.
Of course, this issue is not only relevant to game development. Japanese author Rié Kudan, winner of the Ryūnosuke Akutagawa Prize, admitted that about 5% of her novel "Tokyo Tower of Empathy" was written by AI. The book itself deals with a moral question - the use of artificial intelligence.

Rie Qudan decided to test for herself how AI affects our lives. As a result, she received a wave of negative feedback, as not everyone believes that without using neural networks she could have achieved such success. The organizers remain neutral.
Steam has been warning developers for a long time that using content generated with generative models in games may lead to problems - up to blocking. The fact is that inside supposedly unique images generated by a neural network, there may be fragments of someone else's work protected by copyright. And you need to respect it.

Not so long ago, Steam introduced rules regulating this issue. Now, in case of using GAI when creating a game, it must be indicated. When publishing such content, it must be marked and proved that it does not infringe on anyone's rights.

No restrictions can prevent enthusiasts from exploring the boundaries of AI applicability in the gaming industry. And now there are already worthy games with AI in Steam from independent developers. Here are a few examples:
Playing with GAI
In the stylish adventure in the spirit of "Twin Peaks" Who's Lila by Russian solo developer Garage Heathen, we don't choose dialogue options - we change the protagonist's expression using a mouse. The neural network recognizes the emotion and it affects the story.

Another solo project "AIadventure" offers to go on a text adventure that is created by artificial intelligence in real-time. You can write anything you want in a living language - almost any, since the game supports auto-translation using the same AI.

DREAMIO: AI-Powered Adventures goes even further and generates not only text but also illustrations and voiceover. It's like an endless visual novel where you can do anything you want.

Interactive detective Vaudeville allows you to communicate with witnesses and suspects as if they were living people - with voices. You can talk to characters on almost any topic, except obscene or unethical ones, and they will maintain the conversation without deviating from their role. It seems like the future is already here.

Not much time has passed since generative models became a tool for developers on a massive scale. Of course, this field will eventually be formalized better. But it's worth being cautious when using GAI now, especially since AI is gradually covering more areas of game development.
Where does AI show the most benefit?
Neural networks are most useful in visual assistance, where they take over routine work such as placing objects. Such proposals do well on the market (think Promethean AI). According to promises made by Unity Muse creators, this direction will continue to develop.
Additionally, neural networks can become helpful assistants for programmers in studying game libraries.

It's hard to find examples of using GAI for creating, say, game levels during the program's work. This is due to the fact that neural networks are prone to the problem of generative mush and are not deterministic in general. A developer can influence the training process and thus improve the results obtained later on. But after training, it is no longer possible to affect what results the AI will produce.
For games, a guaranteed standard gaming experience is very important. However, achieving this with a neural network that cannot even understand that it created an impassable level, for example, is difficult. Moreover, we cannot guarantee a given game pace or that the complexity of the generated game level (say, in a first-person shooter) will be the same from generation to generation. In contrast, when using PCG, these metrics can be checked even at the stage of generation.
After creating a level with a neural network, you can launch the process of testing its compliance with gaming metrics using agent systems. However, this is still a rather long process, and agent systems are not much different from PCG, as they are also algorithmic.
In combination, these problems lead to the fact that there are no commercial products on the market that can create game levels. But I hope that in the future, the situation will change.
Moreover, one of the largest players in the industry openly states: "AI will save the gaming industry." Tencent Holdings CEO Pony Ma stated that despite high competition, they were able to reduce the technological lag from leading companies in AI technologies by 2023. However, it's not yet clear what we can expect - PUBG Mobile and Genshin Impact with the use of these technologies or massive layoffs. Time will show.

***
So now neural networks can be used to generate simple game content such as sprites, textures, and (with significant restrictions) 3D models. There are initial attempts to use text models as a basis for some part of the game's artificial intelligence. However, all this must be done with caution, because the legal and ethical aspects of the issue have not been fully worked out yet.
If during development it is necessary to guarantee any deterministic permanent result that relies on understandable generation parameters, then, of course, PCG algorithms should be used.
However, if a fast tool for prototyping games is needed to collect first rough drafts of game assets already at early stages of work or simply an artistic assistant tool, then GAI can be safely used with neural networks, which show themselves much better in the creative key.
For developers of both games and game engines, it's very important to distinguish between PCG and GAI, because these are absolutely different tools that can give incredible results when used together skillfully.
P.S. Is the future so close? And how many artificially created faces do you think you can guess?
