create_chat_completion | Generate a model reply for a conversation (non-streaming, no tools). Arguments: model (string, a model id from list_models, e.g. “llama-3.3-70b-versatile”); messages (array of at least one object {role: “system”|“user”|“assistant”, content: string (plain text), name?: string} — the conversation so far, oldest first, usually one optional “system” message then alternating “user”/“assistant”, ending with a “user” message); temperature (number 0–2); top_p (number 0–1); max_completion_tokens (integer ≥ 1); stop (a string, or an array of up to 4 strings); seed (integer); response_format (object — {type: “text”}, {type: “json_object”} (the messages must also ask for JSON), or {type: “json_schema”, json_schema: {name: string, schema?: object (a JSON Schema), description?: string, strict?: boolean}} for Structured Outputs on models that support it). |
create_response | Generate a model response with the OpenAI-compatible Responses API (non-streaming, no tools, nothing stored). Arguments: model (string, a model id from list_models); input (either a plain string, treated as one user message, or an array of at least one object {role: “system”|“developer”|“user”|“assistant”, content: string (plain text)}); instructions (string — a system message inserted first); max_output_tokens (integer ≥ 1, includes reasoning tokens); temperature (number 0–2); top_p (number 0–1); text (object {format: {type: “text”} | {type: “json_object”} | {type: “json_schema”, name: string, schema: object (a JSON Schema), description?: string, strict?: boolean}}). |
rerank_documents | Rank documents by relevance to a query (scores sorted in descending order). Arguments: model (string, a reranking model id from list_models); query (string — the search query); docs (array of 1 to 100 non-empty strings — the documents’ text); instruction (string, optional — guides the ranking; a default is used otherwise). |