유럽 최고의 AI 스타트업 Mistral AI의 오픈소스 모델 시리즈. Mistral 7B의 뛰어난 효율성부터 Mixtral MoE(전문가 혼합) 아키텍처, Function Calling, 다국어 지원까지 상세히 다룹니다.
Mistral AI는 2023년 프랑스에서 설립된 AI 스타트업으로, 효율성과 성능 모두를 추구하는 오픈소스 LLM을 개발합니다. Mistral 7B는 출시 당시 같은 크기의 어떤 모델보다 뛰어난 성능을 보여주며 주목받았습니다.
| 모델 | 특징 | 라이선스 |
|---|---|---|
| Mistral 7B | 소형 고성능, Llama 13B 능가 | Apache 2.0 |
| Mixtral 8x7B | MoE, 실제 활성 파라미터 12B | Apache 2.0 |
| Mixtral 8x22B | 대형 MoE, GPT-3.5 수준 | Apache 2.0 |
| Mistral Large | API 전용 프리미엄 모델 | 상용 |
ollama run mistral # Mistral 7B
ollama run mixtral # Mixtral 8x7B
ollama run mixtral:8x22b # Mixtral 8x22Bpip install mistralaifrom mistralai import Mistral
client = Mistral(api_key="YOUR_API_KEY")
response = client.chat.complete(
model="mistral-small-latest",
messages=[{"role": "user", "content": "Mixtral MoE가 뭔가요?"}],
)
print(response.choices[0].message.content)로컬 Ollama를 사용하면 API 키 없이 무료로 사용 가능합니다. OpenAI 호환 엔드포인트: http://localhost:11434/v1
Mistral은 강력한 Function Calling(도구 사용)을 지원합니다.
from mistralai import Mistral
import json
client = Mistral(api_key="YOUR_API_KEY")
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "특정 도시의 현재 날씨를 가져옵니다.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "도시 이름"},
},
"required": ["city"],
},
},
}]
response = client.chat.complete(
model="mistral-small-latest",
messages=[{"role": "user", "content": "서울 날씨 알려줘"}],
tools=tools,
tool_choice="auto",
)
# 도구 호출 처리
if response.choices[0].finish_reason == "tool_calls":
tool_call = response.choices[0].message.tool_calls[0]
args = json.loads(tool_call.function.arguments)
print(f"함수 호출: {tool_call.function.name}({args})")messages = [
{"role": "system", "content": "당신은 친절한 한국어 코딩 어시스턴트입니다."},
{"role": "user", "content": "Python 데코레이터를 설명해주세요."},
]history = []
def chat(user_input: str) -> str:
history.append({"role": "user", "content": user_input})
response = client.chat.complete(model="mistral-small-latest", messages=history)
assistant_msg = response.choices[0].message.content
history.append({"role": "assistant", "content": assistant_msg})
return assistant_msgMixtral은 MoE(Mixture of Experts) 아키텍처를 사용합니다. 총 8개의 전문가(Expert) 레이어가 있지만, 각 토큰 처리 시 2개만 활성화됩니다. 이 덕분에 56B 파라미터를 가지면서도 실제 계산량은 12B 수준입니다.
| 항목 | Mistral 7B | Mixtral 8x7B |
|---|---|---|
| 총 파라미터 | 7B | 56B |
| 활성 파라미터 | 7B | ~12B |
| 컨텍스트 길이 | 32K | 64K |
| 추론 속도 | 빠름 | 7B와 유사 |