yera.models.interfaces.llms.base
Base interface for LLM implementations.
This module defines the abstract BaseLLMInterface that all llm provider implementations must inherit from. It establishes the contract for:
- Streaming chat completions (chat method)
- Generating structured outputs conforming to a schema (make_struct method)
- Managing llm client lifecycle (start/stop methods)
Concrete implementations (e.g., AnthropicLLM, OpenAILLM, AwsBedrockLLM) provide provider-specific implementations of these abstract methods whilst handling their respective API clients and configuration.
Symbols
BaseLLMInterface
ABCMistralLLM, AwsBedrockLLM, NoLLM, AnthropicLLM, OllamaLLM, OpenAILLM, OpenRouterLLM, LlamaCppLLM, GeminiLLMAbstract base interface for llm implementations.
Defines the contract that all concrete llm provider implementations must satisfy. Subclasses handle provider-specific client initialisation, authentication, and API interaction whilst conforming to the streaming chat and structured output methods defined here.
Methods
BaseLLMInterface.chat
chat(
messages: list[Message],
reasoning_level: ReasoningLevel | None = None,
**kwargs,
) → Iterator[LLMToken]Stream a chat completion response.
Abstract method that must be implemented by concrete llm providers. Sends a conversation to the llm and streams the response as text tokens.
Parameters
List of Message objects representing the conversation history.
Set the reasoning effort level overriding the default (medium)
Provider-specific keyword arguments to customise llm behaviour (e.g., temperature, top_p, max_tokens).
BaseLLMInterface.make_struct
make_struct(
messages: list[Message],
reasoning_level: ReasoningLevel | None = None,
**kwargs,
) → Iterator[LLMToken]Stream a structured output response conforming to a provided schema.
Abstract method that must be implemented by concrete llm providers. Generates a response that strictly conforms to the structure defined by the provided schema class.
Parameters
List of Message objects representing the conversation history.
A pydantic model class defining the output structure. The provider will transform this into the format required by its respective API.
Set the reasoning effort level overriding the default (off)
Provider-specific keyword arguments to customise llm behaviour.
BaseLLMInterface.make_request_struct
make_request_struct(
messages: list[Message],
reasoning_level: ReasoningLevel | None = None,
**kwargs,
) → Iterator[LLMToken]Generate a tool-like request structure with an initial call_id token.
Prepares a unique identifier for the tool call and streams structured output
tokens. This method is intended for tools that require a tool_call → tool_result
interaction pattern, where the call_id must be sent first to associate results.
Parameters
Conversation history.
Struct subclass defining the tool's input/output schema.
Set the reasoning effort level overriding the default (off)
Per-call LLM overrides (e.g., temperature, max_tokens).
BaseLLMInterface.start
start() → NoneInitialise the llm client.
Lifecycle method called before making API requests. Concrete implementations override this to instantiate and configure their provider-specific client. Default implementation does nothing.
BaseLLMInterface.stop
stop() → NoneShut down and clear the llm client.
Lifecycle method called when finished with the llm. Concrete implementations override this to release resources and clean up the provider-specific client. Default implementation does nothing.
BaseLLMInterface.with_instruction
with_instruction(
messages: list[Message],
instruction: str | None,
) → list[Message]Prepend or insert an instruction into the conversation history.
Parameters
The existing conversation history.
The extra instruction to inject.
Returns
A new list of messages with the instruction incorporated.
LLMToken
Class representing a token in an LLM response.
Can be either "thinking" from a reasoning (CoT) trace, or "response" as in a normal response to the user.
Attributes
a string that defines this as a thinking or response token.
the content of the token.
RateLimitError
RuntimeErrorA provider rate-limit response that may succeed after waiting.