Mistral's newest tokenizer has two major improvements:
System prompt
Similar to other tokenization schemes the system prompt is now treated as a "normal" message encapsulated by [SYSTEM_PROMPT] ...[\SYSTEM_PROMPT]
E.g.
from mistral_common.protocol.instruct.messages import (
UserMessage,
SystemMessage,
AssistantMessage,
)
from mistral_common.protocol.instruct.request import ChatCompletionRequest
from mistral_common.tokens.tokenizers.mistral import MistralTokenizer
# Load Mistral tokenizer
tokenizer = MistralTokenizer.v7()
# Tokenize a list of messages
tokenized = tokenizer.encode_chat_completion(
ChatCompletionRequest(
messages=[
SystemMessage(content="You are a funny AI assistant. Always make jokes."),
UserMessage(content="What's the weather like today in Paris"),
],
model="joker",
)
)
tokens, text = tokenized.tokens, tokenized.text
print(text)
# <s>[SYSTEM_PROMPT]▁You▁are▁a▁funny▁AI▁assistant.▁Always▁make▁jokes.[/SYSTEM_PROMPT][INST]▁What's▁the▁weather▁like▁today▁in▁Paris[/INST]
Improve function calling
A new [TOOL_CONTENT] is added if trained with correctly should improve the accuracy of function calling.
from mistral_common.protocol.instruct.messages import (
UserMessage,
SystemMessage,
AssistantMessage,
ToolMessage
)
from mistral_common.protocol.instruct.request import ChatCompletionRequest …