Skip to content

Configuration

The library uses a tiered configuration system combining TOML project files, user configurations, environment variables, and programmatic overrides. Settings are validated strictly on load; unknown keys or invalid data types raise a ConfigurationError.

Configuration Precedence

Settings.load() evaluates configuration sources in the following order, from lowest to highest precedence:

  • Library defaults
  • User configuration at ~/.config/indic-language-utils/config.toml (or $XDG_CONFIG_HOME/indic-language-utils/config.toml)
  • The nearest .indic-language-utils.toml found by walking upward from the current working directory
  • An explicit configuration file specified via Settings.load(path=...) or the ILU_CONFIG_FILE environment variable
  • Environment variables (ILU_*, BHASHINI_*, and TRANSLATION_SERVICE_PROVIDER)
  • Programmatic overrides passed via Settings.load(overrides=...)

Only existing discovered files are merged. If an explicit configuration path is provided, that file must exist.

Project Configuration File

A project configuration file (.indic-language-utils.toml) can be placed at the root of a project repository. The following example illustrates all supported configuration sections:

[cache]
enabled = true
backend = "sqlite"
path = ".cache/translations.sqlite3"
namespace = "my-application"
max_entries = 50000
tts_max_bytes = 67108864
ttl_seconds = 86400

[retry]
max_attempts = 3
base_delay_seconds = 0.25
max_delay_seconds = 5.0

[telemetry]
logging_enabled = true
metrics_enabled = true
traces_enabled = true
include_content = false

[providers.bhashini]
endpoint = "https://dhruva-api.bhashini.gov.in/services/inference/pipeline"
translation_service_id = "default-translation-model-id"
detection_service_id = "default-tld-model-id"
transliteration_service_id = "default-transliteration-model-id"
stt_model_id = "default-asr-service-id"
tts_model_id = "default-tts-service-id"
timeout_seconds = 20.0
max_concurrency = 8

[providers.bhashini.translation_service_ids]
"hi-IN" = "service-for-any-source-to-hindi"
"en-IN>ta-IN" = "service-for-english-to-tamil"

[providers.bhashini.stt_model_ids]
"hi-IN" = "asr-service-for-hindi"
"ta-IN" = "asr-service-for-tamil"

[providers.bhashini.tts_model_ids]
"hi-IN" = "tts-service-for-hindi"

[providers.googletrans]
timeout_seconds = 20.0
max_concurrency = 4

[providers.sarvam]
endpoint = "https://api.sarvam.ai"
model = "sarvam-translate:v1"
stt_model_id = "saaras:v4"
tts_model_id = "bulbul:v3"
timeout_seconds = 20.0
max_concurrency = 8

[providers.navana]
endpoint = "https://tts.navana.ai"
timeout_seconds = 120.0
max_concurrency = 8

[providers.gnani]
endpoint = "https://api.vachana.ai"
timeout_seconds = 120.0
max_concurrency = 8

[providers.my_adapter.options]
region = "south"
batch_limit = 12

[routes]
translation = ["sarvam", "bhashini", "googletrans"]
text_language_detection = ["sarvam", "bhashini", "fasttext"]
transliteration = ["bhashini", "aksharamukha"]
speech_to_text = ["bhashini", "sarvam"]
text_to_speech = ["bhashini", "sarvam", "navana"]

Relative cache paths resolve relative to the current working directory of the process. In production containers or multi-directory environments, specify an absolute path.

Provider-specific options

An adapter can read settings under [providers.<id>.options]. Options support TOML scalar values, arrays, and nested tables. The loader validates their shape and rejects secret-like keys, including in nested options. It merges these values across configuration files and programmatic overrides using the normal precedence rules. Each adapter must validate the option names and value types it uses:

from indic_language_utils import Settings
from indic_language_utils.errors import ConfigurationError

settings = Settings.load()
options = settings.providers["my_adapter"].options
batch_limit = options.get("batch_limit", 12)
if not isinstance(batch_limit, int) or isinstance(batch_limit, bool) or batch_limit < 1:
    raise ConfigurationError("my_adapter.options.batch_limit must be a positive integer")

Keep credentials in environment variables or pass a Secret directly to the adapter. An options section cannot contain keys such as api_key, token, or password.

Secrets and Environment Variables

TOML files are strictly prohibited from storing credentials, passwords, tokens, or API keys. Attempting to define fields containing sensitive terms (such as api_key, token, secret, or password) raises a ConfigurationError.

All sensitive values must be supplied via environment variables.

Provider Credentials and Overrides

Bhashini credentials and service endpoints:

# Required for Bhashini live API calls
export BHASHINI_API_KEY="your-bhashini-api-key"

# Optional overrides for endpoints and service IDs
export BHASHINI_ENDPOINT_URL="https://dhruva-api.bhashini.gov.in/services/inference/pipeline"
export BHASHINI_TRANSLATION_SERVICE_ID="your-translation-service-id"
export BHASHINI_DETECTION_SERVICE_ID="your-tld-service-id"
export BHASHINI_TRANSLITERATION_SERVICE_ID="your-transliteration-service-id"
export BHASHINI_STT_MODEL_ID="your-default-asr-service-id"
export BHASHINI_TTS_MODEL_ID="your-default-tts-service-id"
export BHASHINI_TIMEOUT_SECONDS="20"
export BHASHINI_MAX_CONCURRENCY="8"

Sarvam AI credentials and overrides:

# Required for Sarvam live API calls
export SARVAM_API_KEY="your-sarvam-api-key"

# Optional overrides
export SARVAM_ENDPOINT_URL="https://api.sarvam.ai"
export SARVAM_MODEL="sarvam-translate:v1"
export SARVAM_STT_MODEL_ID="saaras:v4"
export SARVAM_TTS_MODEL_ID="bulbul:v3"
export SARVAM_TIMEOUT_SECONDS="20"
export SARVAM_MAX_CONCURRENCY="8"

Navana AI TTS credentials and endpoint overrides:

export NAVANA_API_KEY="your-navana-api-key"
export NAVANA_ENDPOINT_URL="https://tts.navana.ai"
export NAVANA_TIMEOUT_SECONDS="120"
export NAVANA_MAX_CONCURRENCY="8"

Gnani STT and TTS credentials and runtime overrides:

export GNANI_API_KEY="your-gnani-api-key"
export GNANI_ENDPOINT_URL="https://api.vachana.ai"
export GNANI_TIMEOUT_SECONDS="120"
export GNANI_MAX_CONCURRENCY="8"

Google Free STT overrides

export GOOGLE_FREE_STT_TIMEOUT_SECONDS="20" export GOOGLE_FREE_STT_MAX_CONCURRENCY="4" export GOOGLE_FREE_STT_DEFAULT_LANGUAGE="en-IN"

Faster-Whisper STT overrides

export FASTER_WHISPER_MODEL="base" export FASTER_WHISPER_DEVICE="auto" export FASTER_WHISPER_COMPUTE_TYPE="default" export FASTER_WHISPER_TIMEOUT_SECONDS="60" export FASTER_WHISPER_MAX_CONCURRENCY="2"

To quickly select an active provider during development without editing configuration files:

export TRANSLATION_SERVICE_PROVIDER="sarvam"
export STT_SERVICE_PROVIDER="google_free"  # or "faster_whisper"

Shared System Settings

Global cache, retry, telemetry, and routing settings use the ILU_ prefix:

# Caching settings
export ILU_CACHE_ENABLED="true"
export ILU_CACHE_BACKEND="sqlite"
export ILU_CACHE_PATH="/var/lib/indic-language-utils/cache.sqlite3"
export ILU_CACHE_NAMESPACE="my-app"
export ILU_CACHE_MAX_ENTRIES="50000"
export ILU_CACHE_TTS_MAX_BYTES="67108864"
export ILU_CACHE_TTL_SECONDS="86400"

# Retry settings
export ILU_RETRY_MAX_ATTEMPTS="3"
export ILU_RETRY_BASE_DELAY_SECONDS="0.25"
export ILU_RETRY_MAX_DELAY_SECONDS="5.0"

# Telemetry settings
export ILU_LOGGING_ENABLED="true"
export ILU_METRICS_ENABLED="true"
export ILU_TRACES_ENABLED="true"
export ILU_TELEMETRY_INCLUDE_CONTENT="false"

# Route overrides (comma-separated provider names in priority order)
export ILU_ROUTE_TRANSLATION="bhashini,googletrans"
export ILU_ROUTE_TEXT_LANGUAGE_DETECTION="bhashini,fasttext"
export ILU_ROUTE_SPEECH_TO_TEXT="bhashini,sarvam"
export ILU_ROUTE_TEXT_TO_SPEECH="bhashini,sarvam"

Programmatic Loading and Overrides

Automatic Discovery

The easiest way to initialize a client is via the built-in factories, which call Settings.load() automatically:

from indic_language_utils import get_translation_client

# Automatically reads .indic-language-utils.toml and environment variables
client = get_translation_client()

Explicit Settings Loading

For granular control over settings loading:

from pathlib import Path
from indic_language_utils import Settings

# Load settings from a specific TOML file
settings = Settings.load(path=Path("/etc/indic-language-utils/production.toml"))

# Load settings with test overrides
test_settings = Settings.load(
    overrides={
        "cache": {"enabled": False},
        "retry": {"max_attempts": 1},
    }
)

Building Custom Clients from Settings

Assemble a client pipeline explicitly using loaded settings:

from indic_language_utils import (
    BhashiniConfig,
    BhashiniTranslationProvider,
    CacheKeyBuilder,
    CapabilityId,
    GoogleTranslateProvider,
    OrderedRouter,
    ProviderRegistry,
    Settings,
    TranslationClient,
    create_translation_cache,
    create_translation_segment_cache,
)

settings = Settings.load()

registry = ProviderRegistry()
if "bhashini" in settings.providers:
    registry.register(BhashiniTranslationProvider(BhashiniConfig.from_settings(settings)))
registry.register(GoogleTranslateProvider())

route = settings.routes.get("translation", ("bhashini", "googletrans"))
router = OrderedRouter(registry, {CapabilityId.TRANSLATION: route})

client = TranslationClient(
    router=router,
    cache=create_translation_cache(settings.cache),
    segment_cache=create_translation_segment_cache(settings.cache),
    cache_keys=CacheKeyBuilder(settings.cache.namespace),
)