---
title: Functionary
slug: functionary.meetkai.com
icon: 🐑
docTags: 
createdAt: 2024-01-15T07:58:22.530Z
---

::Image[]{src="https://github.com/musabgultekin/functionary/assets/3749407/c7a1972d-6ad7-40dc-8000-dceabe6baabd" size="30" width="512" height="512" position="center" showCaption="false"}

Functionary is a language model that can interpret and execute functions/plugins.

The model determines when to execute functions, whether in parallel or serially, and can understand their outputs. It only triggers functions as needed. Function definitions are given as JSON Schema Objects, similar to OpenAI GPT function calls.

# Key Features

## Model Features

- Intelligent **parallel function-calls/tool-uses**
- Able to analyze functions/tools outputs and provide relevant responses grounded in the function/tools outputs
- Able to decide **when to not use tools/call functions&#x20;**&#x61;nd provide normal chat response

## Server/Inference Features

- OpenAI-compatible server based on the blazing fast [vLLM](https://github.com/vllm-project/vllm)
- **Grammar Sampling** resulting in 100% accuracy in generating function and argument names
- Automatically **execute python functions&#x20;**&#x63;alled by Functionary

## Us vs the Industry

| **Feature/Project**                                          | [Functionary](https://github.com/MeetKai/functionary) | [NexusRaven](https://github.com/nexusflowai/NexusRaven) | [Gorilla](https://github.com/ShishirPatil/gorilla) | [Glaive](https://huggingface.co/glaiveai/glaive-function-calling-v1) | [GPT-4-1106-preview](https://github.com/openai/openai-python) |
| ------------------------------------------------------------ | ----------------------------------------------------- | ------------------------------------------------------- | -------------------------------------------------- | -------------------------------------------------------------------- | ------------------------------------------------------------- |
| Single Function Call                                         | -                                                     | *                                                       | -                                                  | *                                                                    | -                                                             |
| Parallel Function Calls                                      | *                                                     | -                                                       | *                                                  | -                                                                    | *                                                             |
| Nested Function Calls                                        | -                                                     | *                                                       | -                                                  | *                                                                    | -                                                             |
| Following up on Missing Function Arguments                   | *                                                     | -                                                       | *                                                  | -                                                                    | *                                                             |
| Multi-turn                                                   | -                                                     | *                                                       | -                                                  | *                                                                    | -                                                             |
| Generate Model Responses Grounded in Tools Execution Results | *                                                     | -                                                       | *                                                  | -                                                                    | *                                                             |
| Chit-chat                                                    | -                                                     | *                                                       | -                                                  | *                                                                    | -                                                             |

# Performance

## Function Prediction Evaluation

We evaluated Functionary's function call prediction capabilities on MeetKai's in-house benchmark dataset. The accuracy metric measures the overall correctness of predicted function calls, including function name prediction and arguments extraction.

![](https://github.com/MeetKai/functionary/raw/main/assets/Functioncall_acc_chart.jpeg)

| Dataset       | Model Name                      | Function Calling  Accuracy (Name & Arguments) |
| ------------- | ------------------------------- | --------------------------------------------- |
| In-house data | MeetKai-functionary-small-v2.2  | 0.546                                         |
| In-house data | MeetKai-functionary-medium-v2.2 | **0.664**                                     |
| In-house data | OpenAI-gpt-3.5-turbo-1106       | 0.531                                         |
| In-house data | OpenAI-gpt-4-1106-preview       | **0.737**                                     |

# Get Started

## Installation

Make sure you have [PyTorch](https://pytorch.org/get-started/locally/) installed. Then to install the required dependencies, run:

```shell
pip install -r requirements.txt
```

## Running the server

Now you can start a blazing fast [vLLM](https://vllm.readthedocs.io/en/latest/getting_started/installation.html) server:

```shell
python3 server_vllm.py --model "meetkai/functionary-small-v2.2" --host 0.0.0.0
```

### Server Options

Since the server is running on vLLM, you can pass in various arguments for the vLLM Engine. They are listed [here.](https://github.com/vllm-project/vllm/blob/main/vllm/engine/arg_utils.py#L11-L37) The more important arguments are:

- **--model&#x20;**: the name of model (e.g.: "meetkai/functionary-medium-v2.2")
- **--tensor-parallel-size (default = 1)** : the number of GPUs to use for the server. This should be set to greater than 1 for larger models like "meetkai/functionary-medium-v2.2" depending on the type of GPU that you are using
- **--max-model-len** : the context window to set for your server.

There are also some features implemented independent of vLLM by us:

- **--host (default="0.0.0.0")** : the host IP address for the server
- **--port (default="8000")** : the port to be exposed for the server
- **--grammar\_sampling (default=True)** : when enabled, the server uses our [implementation of grammar sampling](https://github.com/MeetKai/functionary/pull/74) that ensures 100% accuracy for function and argument names. It works by constraining function and argument names to only the list of functions provided in the API call or no function call at all.

## Running the client

### OpenAI Compatible Usage

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="functionary")

client.chat.completions.create(
    model="meetkai/functionary-small-v2.2",
    messages=[{"role": "user",
            "content": "What is the weather for Istanbul?"}
    ],
    tools=[{
            "type": "function",
            "function": {
                "name": "get_current_weather",
                "description": "Get the current weather",
                "parameters": {
                    "type": "object",
                    "properties": {
                        "location": {
                            "type": "string",
                            "description": "The city and state, e.g. San Francisco, CA"
                        }
                    },
                    "required": ["location"]
                }
            }
        }],
    tool_choice="auto"
)
```

### Raw Usage

```python
import requests

data = {
    'model': 'meetkai/functionary-small-v2.2', # model name here is the value of argument "--model" in deploying: server_vllm.py or server.py
    'messages': [
        {
            "role": "user",
            "content": "What is the weather for Istanbul?"
        }
    ],
    'tools':[ # For functionary-7b-v2 we use "tools"; for functionary-7b-v1.4 we use "functions" = [{"name": "get_current_weather", "description":..., "parameters": ....}]
        {
            "type": "function",
            "function": {
                "name": "get_current_weather",
                "description": "Get the current weather",
                "parameters": {
                    "type": "object",
                    "properties": {
                        "location": {
                            "type": "string",
                            "description": "The city and state, e.g. San Francisco, CA"
                        }
                    },
                    "required": ["location"]
                }
            }
        }
    ]
}

response = requests.post("http://127.0.0.1:8000/v1/chat/completions", json=data, headers={
    "Content-Type": "application/json",
    "Authorization": "Bearer xxxx"
})

# Print the response text
print(response.text)
```

If you're having trouble with dependencies, and you have [nvidia-container-toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html#setting-up-nvidia-container-toolkit),
you can start your environment like this:

```shell
sudo docker run --gpus all -it --shm-size=8g --name functionary -v ${PWD}/functionary_workspace:/workspace -p 8000:8000 nvcr.io/nvidia/pytorch:22.12-py3
```

## Models Available

| Model                                                                                                                                                   | Description                                              | Compute Requirements (for FP16 HF model weights) |
| ------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------- | ------------------------------------------------ |
| [functionary-small-v2.2](https://huggingface.co/meetkai/functionary-small-v2.2) / [GGUF](https://huggingface.co/meetkai/functionary-small-v2.2-GGUF)    | 8k context                                               | Any GPU with 24GB VRAM                           |
| [functionary-medium-v2.2](https://huggingface.co/meetkai/functionary-medium-v2.2) / [GGUF](https://huggingface.co/meetkai/functionary-medium-v2.2-GGUF) | 8k context, better accuracy                              | 2 x A100-80GB or equivalent                      |
| [functionary-7b-v2.1](https://huggingface.co/meetkai/functionary-7b-v2.1) / [GGUF](https://huggingface.co/meetkai/functionary-7b-v2.1-GGUF)             | 8k context                                               | Any GPU with 24GB VRAM                           |
| [functionary-7b-v2](https://huggingface.co/meetkai/functionary-7b-v2) / [GGUF](https://huggingface.co/meetkai/functionary-7b-v2-GGUF)                   | Parallel function call support.                          | Any GPU with 24GB VRAM                           |
| [functionary-7b-v1.4](https://huggingface.co/meetkai/functionary-7b-v1.4) / [GGUF](https://huggingface.co/meetkai/functionary-7b-v1.4-GGUF)             | 4k context, better accuracy (deprecated)                 | Any GPU with 24GB VRAM                           |
| [functionary-7b-v1.1](https://huggingface.co/meetkai/functionary-7b-v1.1)                                                                               | 4k context (deprecated)                                  | Any GPU with 24GB VRAM                           |
| functionary-7b-v0.1                                                                                                                                     | 2k context (deprecated) Not recommended, use 2.1 onwards | Any GPU with 24GB VRAM                           |

### Compatibility information

- v1 models are compatible with both OpenAI-python v0 and v1.
- v2 models are designed for compatibility with OpenAI-python v1.

You may refer to the official OpenAI documentation [here](https://platform.openai.com/docs/api-reference/chat/create#chat-create-tools) for the difference between OpenAI-python v0 and v1&#x20;

# How it works?

We convert function definitions to a similar text to TypeScript definitions. Then we inject these definitions as system prompts. After that, we inject the default system prompt.
Then we start the conversation messages.

We use specially designed prompt templates which we call "v2PromptTemplate" and "v1PromptTemplate". v2PromptTemplate breaks down each turns into from, recipient and content portions. "v1PromptTemplate" uses a variety of special tokens in each turn. These includes a start-of-function-call token to indicate an assistant turn with function call and role-specific stop token for each turn.

The prompt example can be found here: [V1](https://github.com/MeetKai/functionary/blob/main/tests/prompt_test_v1.txt) and [V2](https://github.com/MeetKai/functionary/blob/main/tests/prompt_test_v2.txt)

We don't change the logit probabilities to conform to a certain schema, but the model itself knows how to conform. This allows us to use existing tools and caching systems with ease.

# License

This project is licensed under the terms of the MIT license.

