PolarSPARC

Decision Model(s) on Ollama & Llama.cpp


Bhaskar S 10/03/2026


Overview

Both of the popular open source LLM serving platforms - Ollama and llama.cpp now provide support for Jev like Decision models.

Jev is *NOT* a frontier LLM model, but a fast and performant Decision model that outputs structured decisions (along with a confidence scores) for a given input state and a set of question(s). It does *NOT* generate any text like the frontier LLM models.

A decision model supports the following three types of question types:

In this article, we will demonstrate how one can setup and use decision models on a local desktop using Ollama OR llama.cpp from a Python script.


Installation and Setup

The installation and setup will can on a Ubuntu 24.04 LTS based Linux desktop. Ensure that Docker is installed and setup on the desktop (see instructions).

Also, ensure that the Python 3.1x programming language is installed and setup.

We will setup two required directories by executing the following commands in a terminal window:


$ mkdir -p $HOME/.ollama

$ mkdir -p $HOME/.llama_cpp/models


Next, we pull and download the required docker images for Ollama and llama.cpp by executing the following commands in a terminal window:


$ docker pull ollama/ollama:0.35.1

$ docker pull ghcr.io/ggml-org/llama.cpp:full-cuda-b11371


Finally, we install the necessary Python packages by executing the following command:


$ pip install dotenv typesafe-sdk


This completes all the system installation and setup for the hands-on demonstration.


Hands-on with Ollama


Assuming the linux desktop has Nvidia GPU with decent amount of VRAM (at least 16 GB) and has been enabled for use with docker (see instructions) and has the ip address of 192.168.1.25, execute the following command to start Ollama:


$ docker run --rm --name ollama --gpus=all -e OLLAMA_KEEP_ALIVE=-1 -p 192.168.1.25:11434:11434 -v $HOME/.ollama:/root/.ollama ollama/ollama:0.35.1


For the hands-on demonstration, we will download and use the Tev1 0.8B model.

Open a new terminal window and execute the following docker command to download the mentioned decision model:

$ docker exec -it ollama ollama run tev1:0.8b


Now, time to test the decision model tev1 served by Ollama using TypeSafe SDK in Python.

Create a file called ollama.env with the following environment variables defined:


#
### Date: 10/03/2026
### Ollama: tev1:0.8b
#
TYPESAFE_BASE_URL='http://192.168.1.25:11434'
TYPESAFE_API_KEY='ollama'
TYPESAFE_DEFAULT_MODEL='tev1:0.8b'

To load the environment variables and assign them to corresponding Python variables, execute the following code snippet:


from dotenv import dotenv_values

config = dotenv_values('ollama.env')

api_key = config.get('TYPESAFE_API_KEY')
base_url = config.get('TYPESAFE_BASE_URL')
model = config.get('TYPESAFE_DEFAULT_MODEL')

To initialize an instance of the TypeSafe client to connect to use the decision model running on the host URL, execute the following code snippet:


from typesafe_sdk import TypeSafeClient

client = TypeSafeClient(
    api_key=api_key,
    base_url=base_url,
    model=model
)

The first use-case demonstrates the use of the Noul decision type. To test this case, execute the following code snippet:


from typesafe_sdk import Noul

response = client.system_one(
  state='It is very cold outside today',
  questions={
    'weather': Noul(instructions='Is the user talking of the weather?')
  }
)
print(response.answers)

The following should be the typical output:


Output.1

{'weather': NoulAnswer(type='noul', noul=0.8419231565703403)}

The second use-case demonstrates the use of the Choice decision type. To test this case, execute the following code snippet:


from typesafe_sdk import Choice

response2 = client.system_one(
  state='Want to build a slick web app. Trying to decide which language to write it in',
  questions={
    'languages': Choice(
      instructions='Which language the web app should be written in?',
      criteria={
        'golang': 'Preferred when building cloud platform tools',
        'java': 'Preferred for platform portability and server side',
        'typescript': 'Preferred when building web applications'
      }
    )
  }
)
print(response2.choices)

The following should be the typical output:


Output.2

{'languages': ChoiceAnswer(type='choice', choice='typescript', confidence=0.9515392346479077, probabilities={'golang': 0.0020730589728125515, 'java': 0.00633350755977683, 'typescript': 0.9915934334674106})}

The final use-case demonstrates the use of the Score decision type. To test this case, execute the following code snippet:


from typesafe_sdk import Score

response3 = client.system_one(
  state='Like to drink tea; it is my favorite beverage',
  questions={
    'beverage': Score(
      instructions='Help me decide a beverage?',
      criteria=['Water', 'Coffee', 'Masala tea']
    )
  }
)
print(response3.scores)

The following should be the typical output:


Output.3

{'beverage': ScoreAnswer(type='score', score=1.8814333228138191, confidence=0.756360714277704, legend={0: 'Water', 1: 'Coffee', 2: 'Masala tea'}, probabilities={0: 0.054075586504366564, 1: 0.01041550417744781, 2: 0.9355089093181856})}

With this, we conclude the hands-on demonstration on using the Ollama platform for running and working with decision model(s) locally !!!


Hands-on with Llama.cpp


For this hands-on demostration, we will download the Kev 0.8B decision model from Huggingface.

After the model download, move the model GGUF file to the $HOME/.llama_cpp/models folder. To serve the just downloaded decision model, execute the following command in the terminal window:


$ docker run --rm --name llama_cpp_jev --gpus all --network host -v $HOME/.llama_cpp/models:/models ghcr.io/ggml-org/llama.cpp:full-cuda-b11371 --server --model /models/Kev-0.8B-Q8_0.gguf --alias kev:0.8b --host 192.168.1.25 --port 8000 -ngl auto --threads 4 --parallel 1


Now, time to test the decision model kev served by llama.cpp using TypeSafe SDK in Python.

Create a file called llama_cpp.env with the following environment variables defined:


#
### Date: 10/03/2026
### HuggingFace: https://huggingface.co/ggml-org/Kev-0.8B-GGUF
#
TYPESAFE_BASE_URL='http://192.168.1.25:8000'
TYPESAFE_API_KEY='llama_cpp'
TYPESAFE_DEFAULT_MODEL='kev:0.8b'

To load the environment variables and assign them to corresponding Python variables, execute the following code snippet:


from dotenv import dotenv_values

config = dotenv_values('llama_cpp.env')

api_key = config.get('TYPESAFE_API_KEY')
base_url = config.get('TYPESAFE_BASE_URL')
model = config.get('TYPESAFE_DEFAULT_MODEL')

To initialize an instance of the TypeSafe client to connect to use the decision model running on the host URL, execute the following code snippet:


from typesafe_sdk import TypeSafeClient

client = TypeSafeClient(
    api_key=api_key,
    base_url=base_url,
    model=model
)

The first use-case demonstrates the use of the Noul decision type. To test this case, execute the following code snippet:


from typesafe_sdk import Noul

response = client.system_one(
  state='It is very cold outside today',
  questions={
    'weather': Noul(instructions='Is the user talking of the weather?')
  }
)
print(response.nouls)

The following should be the typical output:


Output.4

{'weather': NoulAnswer(type='noul', noul=0.9398799506544198)}

The second use-case demonstrates the use of the Choice decision type. To test this case, execute the following code snippet:


from typesafe_sdk import Choice

response2 = client.system_one(
  state='Want to build a slick web app. Trying to decide which language to write it in',
  questions={
    'languages': Choice(
      instructions='Which language the web app should be written in?',
      criteria={
        'golang': 'Preferred when building cloud platform tools',
        'java': 'Preferred for platform portability and server side',
        'typescript': 'Preferred when building web applications'
      }
    )
  }
)
print(response2.answers['languages'])

The following should be the typical output:


Output.5

type='choice' choice='typescript' confidence=0.37731794988497735 probabilities={'golang': 0.24518399008599562, 'java': 0.1699373766573529, 'typescript': 0.5848786332566516}

The final use-case demonstrates the use of the Score decision type. To test this case, execute the following code snippet:


from typesafe_sdk import Score

response3 = client.system_one(
  state='Like to drink tea; it is my favorite beverage',
  questions={
    'beverage': Score(
      instructions='Help me decide a beverage?',
      criteria=['Water', 'Coffee', 'Masala tea']
    )
  }
)
print(response3.answers['beverage'])

The following should be the typical output:


Output.6

type='score' score=1.6817728488969708 confidence=0.5226592733454563 legend={0: 'Water', 1: 'Coffee', 2: 'Masala tea'} probabilities={0: 0.04304915858852094, 1: 0.2321288339259873, 2: 0.7248220074854917}

With this, we conclude the hands-on demonstration on using the llama.cpp platform for running and working with decision model(s) locally !!!


References

TypeSafe AI Python SDK



© PolarSPARC