Run a Local LLM in Spring Boot with Spring AI, Ollama and Docker
Add AI features to a Spring Boot app without paying per token or sending data to a third party. We will run a model locally with Ollama and call it through Spring AI.
When I built DishGenie, an API that turns a list of ingredients into a recipe, I wanted the whole thing to run on my own machine: no API keys, no usage bills, and no user data leaving the server. Spring AI and Ollama make that surprisingly simple. This guide walks through the same setup with current versions.
Why run the model locally
- Privacy. Prompts and responses never leave your infrastructure, which matters for health, finance and internal company data.
- Predictable cost. You pay for hardware, not per token, so heavy usage does not surprise you at the end of the month.
- Offline development. Your team can build and test AI features without shared API keys or rate limits.
- Easy to swap. Spring AI uses the same
ChatClientAPI for Ollama, OpenAI, Anthropic and others, so moving to a hosted model later is mostly a config change.
Start Ollama in Docker
Ollama serves open models behind a simple HTTP API on port 11434. Run it in a container with a named volume so downloaded models survive restarts:
docker run -d --name ollama \
-p 11434:11434 \
-v ollama:/root/.ollama \
ollama/ollama
# Download a model (a few GB, only needed once)
docker exec -it ollama ollama pull llama3.2
# Quick check that it answers
docker exec -it ollama ollama run llama3.2 "Say hello in one sentence"
DishGenie originally used Llama 2. Newer models like llama3.2 are faster and follow instructions better, but any model from the Ollama library works with the code below.
Add Spring AI to your project
Import the Spring AI BOM so all Spring AI modules share one version, then add the Ollama starter:
<dependencyManagement>
<dependencies>
<dependency>
<groupId>org.springframework.ai</groupId>
<artifactId>spring-ai-bom</artifactId>
<version>1.0.0</version>
<type>pom</type>
<scope>import</scope>
</dependency>
</dependencies>
</dependencyManagement>
<dependencies>
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-web</artifactId>
</dependency>
<dependency>
<groupId>org.springframework.ai</groupId>
<artifactId>spring-ai-starter-model-ollama</artifactId>
</dependency>
</dependencies>
Use the latest Spring AI release that matches your Spring Boot version. The code in this post works with 1.0 and later.
Configure the model
spring:
ai:
ollama:
base-url: http://localhost:11434
init:
pull-model-strategy: when_missing # download the model on startup if needed
chat:
options:
model: llama3.2
temperature: 0.7
A lower temperature gives more consistent answers. For structured output like the recipe below, values between 0.2 and 0.7 work well.
Call the model with ChatClient
Spring Boot auto-configures a ChatClient.Builder. Build a client once with a default system prompt, then reuse it:
@Service
public class RecipeService {
private final ChatClient chatClient;
public RecipeService(ChatClient.Builder builder) {
this.chatClient = builder
.defaultSystem("""
You are a helpful chef. Suggest recipes that only use the
ingredients given plus common pantry staples like salt, oil and water.
""")
.build();
}
public String suggest(List<String> ingredients) {
return chatClient.prompt()
.user(u -> u.text("Suggest one recipe using: {ingredients}")
.param("ingredients", String.join(", ", ingredients)))
.call()
.content();
}
}
Using a template parameter instead of string concatenation keeps prompts readable and makes it clear which part of the prompt came from the user.
Get structured output
Free text is hard to use in an API. Spring AI can map the model's answer straight into a Java record. It adds format instructions to the prompt and parses the JSON response for you:
public record Recipe(
String title,
List<String> ingredients,
List<String> steps,
int totalMinutes) {}
public Recipe generate(List<String> ingredients) {
return chatClient.prompt()
.user(u -> u.text("Create one recipe using: {ingredients}")
.param("ingredients", String.join(", ", ingredients)))
.call()
.entity(Recipe.class);
}
Smaller local models occasionally return invalid JSON. Catch the exception and retry once, or fall back to a friendly error, rather than letting a 500 reach your users.
Expose a REST endpoint
public record RecipeRequest(@NotEmpty @Size(max = 20) List<@NotBlank String> ingredients) {}
@RestController
@RequestMapping("/api/recipes")
public class RecipeController {
private final RecipeService recipeService;
public RecipeController(RecipeService recipeService) {
this.recipeService = recipeService;
}
@PostMapping
public Recipe create(@Valid @RequestBody RecipeRequest request) {
return recipeService.generate(request.ingredients());
}
}
Try it:
curl -s -X POST localhost:8080/api/recipes \
-H "Content-Type: application/json" \
-d '{"ingredients":["rice","eggs","spring onion","soy sauce"]}'
{
"title": "Quick Egg Fried Rice",
"ingredients": ["2 cups cooked rice", "2 eggs", "2 spring onions", "1 tbsp soy sauce", "1 tbsp oil"],
"steps": ["Heat the oil in a wok...", "Scramble the eggs...", "Add the rice and soy sauce..."],
"totalMinutes": 15
}
Run everything with Docker Compose
For a team or a server, put the app and Ollama in one Compose file. Inside the Compose network the app reaches Ollama by its service name:
services:
ollama:
image: ollama/ollama
ports:
- "11434:11434"
volumes:
- ollama:/root/.ollama
app:
build: .
ports:
- "8080:8080"
environment:
SPRING_AI_OLLAMA_BASE_URL: http://ollama:11434
depends_on:
- ollama
volumes:
ollama:
Run docker compose up -d. With pull-model-strategy: when_missing, the app downloads the model on its first start.
Tips for production
- Size the hardware for the model. Small models (1B to 3B parameters) run on a CPU. For 7B and larger, a GPU makes responses many times faster. Ollama supports NVIDIA GPUs in Docker with
--gpus=all. - Expect a slow first request. The model loads into memory on first use, so increase your HTTP client timeouts and consider a warm-up call at startup.
- Stream long answers. Replace
.call()with.stream().content()to send tokens to the client as they are generated. - Treat user input as untrusted. Validate lengths, keep instructions in the system prompt, and never let model output run code or queries directly.
- Keep Ollama private. Do not expose port 11434 to the internet. Only your app should be able to reach it.
That is all it takes to ship a private, self-hosted AI feature with the Spring tools you already know.