Skip to content

Repository files navigation

SmallestAI Python SDK

pypi

pip install smallestai gives you one package with two surfaces, plus a CLI:

  • Atoms — build, configure, deploy, and phone-call voice AI agents (client.atoms).
  • Waves — low-latency text-to-speech and speech-to-text, sync/async and streaming (client.waves).
  • CLIsmallestai for managing agents and deploying agent-crew code.
pip install smallestai

Table of Contents

Quickstart: create an agent and call it

from smallestai import SmallestAI

client = SmallestAI(api_key="<your-api-key>")

# create an agent (the response .data is the new agent id)
agent_id = client.atoms.agents.create_agent(name="my-first-agent").data

# place an outbound call (from_product_id is a rented number's product id)
client.atoms.calls.start_outbound_call(
    agent_id=agent_id,
    phone_number="+1XXXXXXXXXX",
    from_product_id="<rented-number-product-id>",
)

Text-to-speech and speech-to-text (Waves)

synthesize_tts streams audio bytes. List available voices with client.waves.get_voices().

from smallestai import SmallestAI

client = SmallestAI(api_key="<your-api-key>")

with open("out.wav", "wb") as f:
    for chunk in client.waves.synthesize_tts(text="Hello from Smallest.", voice_id="<voice-id>"):
        f.write(chunk)

Streaming speech-to-text helper:

from smallestai.waves.helpers import stream_speech_to_text

for event in stream_speech_to_text(client, language="en"):
    print(event)

Agent crew: your own LLM in the middle

An agent crew runs the LLM turn on a model you choose while Smallest handles STT and TTS. Point the crew node's OpenAIClient at any OpenAI-compatible endpoint (a hosted API, or a local model via Ollama):

from smallestai.atoms.crew.nodes import OutputCrewNode
from smallestai.atoms.crew.clients.openai import OpenAIClient

class Assistant(OutputCrewNode):
    def __init__(self):
        super().__init__(name="assistant")
        self.llm = OpenAIClient(
            model="claude-haiku-4-5",
            api_key="<your-llm-key>",
            base_url="https://api.anthropic.com/v1/",   # or http://localhost:11434/v1 for Ollama
        )

    async def generate_response(self):
        async for chunk in await self.llm.chat(self.context.messages, stream=True):
            if chunk.content:
                yield chunk.content

Conversation history is handled for you: every turn is appended to self.context, and you send self.context.messages to the model each turn.

Deploy it with the CLI:

smallestai auth login
smallestai agent-crew init --agent-id <agent-id>
smallestai agent-crew deploy --entry-point server.py
smallestai agent-crew builds        # pick the build -> Make Live

A flat directory (server.py + requirements.txt at the root) is the simplest layout; a src/ layout with a pyproject.toml also works (declare all runtime deps in the pyproject).

CLI

smallestai auth login               # store your API key
smallestai agents list              # list, get, call, and manage agents
smallestai agent-crew deploy ...     # package and deploy crew code
smallestai agent-crew chat           # talk to a running crew locally

Async client

The SDK exports an async client with the same surface:

import asyncio
from smallestai import AsyncSmallestAI

async def main():
    client = AsyncSmallestAI(api_key="<your-api-key>")
    agents = await client.atoms.agents.list_agents()
    print(agents.data)

asyncio.run(main())

Environments

from smallestai import SmallestAI
from smallestai.environment import SmallestAIEnvironment

client = SmallestAI(environment=SmallestAIEnvironment.PRODUCTION)

Exception handling

from smallestai.core.api_error import ApiError

try:
    client.atoms.agents.get_agent(id="does-not-exist")
except ApiError as e:
    print(e.status_code, e.body)

Streaming and websockets

Waves supports real-time, low-latency streaming over websockets. stream() returns a context manager; iterate it to process messages as they arrive.

from smallestai import SmallestAI

client = SmallestAI(api_key="<your-api-key>")

# real-time speech-to-text
with client.waves.speech_to_text.stream() as socket:
    for message in socket:
        print(message)

The async client mirrors this with async with / async for. Text-to-speech streams too: synthesize_tts(...) yields audio bytes as they are generated (see above).

Advanced

Access raw response data

Use .with_raw_response to get the response headers and status alongside the parsed data.

response = client.atoms.agents.with_raw_response.get_agent(id="<agent-id>")
print(response.headers)       # response headers
print(response.status_code)   # status code
print(response.data)          # parsed object

Retries

The SDK retries retryable requests with exponential backoff (default 2). Configure it at the client or per request.

client = SmallestAI(api_key="<your-api-key>", max_retries=3)

# or per request
client.atoms.agents.get_agent(id="<agent-id>", request_options={"max_retries": 1})

Timeouts

Defaults to 60 seconds. Configure at the client or per request.

client = SmallestAI(api_key="<your-api-key>", timeout=20.0)

# or per request
client.atoms.agents.get_agent(id="<agent-id>", request_options={"timeout_in_seconds": 5})

Custom client

Override the httpx client for proxies, custom transports, and similar.

import httpx
from smallestai import SmallestAI

client = SmallestAI(
    api_key="<your-api-key>",
    httpx_client=httpx.Client(
        proxy="http://my.test.proxy.example.com",
        transport=httpx.HTTPTransport(local_address="0.0.0.0"),
    ),
)

Reference and docs

Contributing

Most of src/ is generated from an API spec and gets overwritten on regeneration, so hand edits there will not stick. If you spot a bug or a gap, open an issue.

About

Official Python client for the Smallest AI API

Resources

Contributing

Security policy

Stars

Watchers

Forks

Releases

Packages

Used by

Contributors

Languages