Skip to content

BitBullet

See BitBullet: bitbullet.co.uk

BitBullet is a fully managed, no-code machine learning platform. Describe your objective to the built-in AI agent, or use guided workstations, to create classification, regression, forecasting, and clustering experiments. You stay in control while BitBullet handles training infrastructure, tuning, diagnostics, and portable bundles.

The BitBullet SDK is the open-source Python toolkit for data scientists who want to write and orchestrate their own workflows with composable building blocks for transformations, training, forecasting, clustering, evaluation, and reproducible artefact metadata.

Full documentation and practical SDK tutorials are available here.

SDK Or Platform?

Use the BitBullet SDK when you want the raw Python building blocks to write and orchestrate your own data-science workflows.

Use BitBullet Platform when you want to configure and manage a supported modelling lifecycle through a guided, no-code environment. It centralises projects, datasets, managed compute, storage, and repeatable workflows while keeping you in control of every decision. Configure and compare experiments with clicks, or ask the AI assistant to prepare a draft for review; inspect the evidence, then export fitted artefacts, preprocessing, metadata, and generated inference code.

Explore BitBullet Platform | Start free

pip install bitbullet

Optional extras keep installations lean:

pip install "bitbullet[inference-models]"   # LightGBM and XGBoost wrappers
pip install "bitbullet[inference-cluster]"  # clustering extras (kmodes, umap, hdbscan)
pip install "bitbullet[forecast-statistical]"  # optional StatsForecast adapter
pip install "bitbullet[train,viz]"          # training, SHAP, and plotting tools
pip install "bitbullet[all]"                # complete SDK

Modules

Module Purpose
bitbullet.transform Fitted transformation pipelines for numerical, categorical, and datetime features.
bitbullet.train Supervised training with Optuna-backed search, feature selection, sample weights, threshold optimisation, and reports.
bitbullet.forecast Ordered and panel forecasting, reduction strategies, rolling-origin evaluation, intervals, and reconciliation.
bitbullet.model_selection Ordered holdouts and temporal windows with explicit gaps and audit metadata.
bitbullet.cluster K-Means, K-Modes, K-Prototypes, DBSCAN, GMM, gamma estimation, categorical weighting, and clustering metrics.
bitbullet.evaluate Structured classification and regression evaluation metrics ready for reports and metadata.
bitbullet.model Model wrappers, metadata, dataset metadata, and serialization helpers.

Quick Examples

Transform Data

from bitbullet.transform import TransformPipeline

pipeline = TransformPipeline(name="credit_features")
pipeline.add("numerical", "standard_scale", columns=["income", "balance"])
pipeline.add("categorical", "onehot_encode", columns=["region"])

X_transformed = pipeline.fit_transform(X_train)
X_new = pipeline.transform(X_new_raw)
pipeline.save("artifacts/transform_pipeline.joblib")

Target-aware encoders receive y directly. target_encode is leakage-aware: fit_transform(..., y=...) returns out-of-fold training encodings, while later transform(...) calls use the stored full-training smoothed mapping.

pipeline = TransformPipeline()
pipeline.add(
    "categorical",
    "target_encode",
    columns=["merchant_category"],
    params={"target_type": "classification", "cv_folds": 5, "cv_strategy": "stratified"},
)
X_encoded = pipeline.fit_transform(X_train, y=y_train)

Train a Classifier

from bitbullet.train import TrainConfig, OptunaTrainer

config = TrainConfig(
    name="default_risk_lgbm",
    model_type="lgbm",
    task="binary_classification",
    n_trials=30,
    optimization_metric="roc_auc",
    optuna_sampler="tpe",  # tpe, random, grid, cmaes
)

trainer = OptunaTrainer(config)
model = trainer.fit(X_train, y_train, X_val=X_val, y_val=y_val)

print(trainer.best_params)
print(trainer.state.optimal_threshold)

Evaluate a Model

from bitbullet.evaluate import evaluate_classification

report = evaluate_classification(
    y_true=y_test,
    y_pred_proba=model.predict_proba(X_test),
    threshold=trainer.state.optimal_threshold or 0.5,
)
print(report.to_dict())

Forecast a Panel

from sklearn.linear_model import Ridge

from bitbullet.forecast import (
    ForecastConfig,
    ForecastFeatureBuilder,
    ForecastFrame,
    ForecastSchema,
    TabularForecaster,
)

schema = ForecastSchema(
    time="date",
    targets="sales",
    entities="store",
    future_covariates="promotion",
    cadence="D",
)
history = ForecastFrame(history_df, schema)

forecaster = TabularForecaster(
    Ridge(alpha=1.0),
    config=ForecastConfig(horizons=(1, 2, 3), strategy="direct"),
    feature_builder=ForecastFeatureBuilder(target_lags=(1, 7)),
).fit(history)

predictions = forecaster.forecast(future_df)

Cluster Data

from bitbullet.cluster.core import ClusterConfig
from bitbullet.cluster.algorithms.partitional import KPrototypesClusterer

config = ClusterConfig(
    name="customer_segments",
    algorithm_type="partitional",
    method="kprototypes",
    n_clusters=5,
    numerical_columns=["income", "spend"],
    categorical_columns=["region", "channel"],
)

clusterer = KPrototypesClusterer(config)
labels = clusterer.fit_predict(df)

Save a Model with Metadata

from bitbullet.model import ModelMetadata, ModelSerializer

metadata = ModelMetadata(
    name="default_risk_lgbm",
    model_type=model.model_type,
    framework=model.framework,
    task="binary_classification",
    metrics=report.metrics,
)
metadata.add_feature_schema(X_train)

ModelSerializer.save(
    model=model,
    path="artifacts/default_risk_lgbm.pkl",
    metadata=metadata,
    train_data=(X_train, y_train),
    test_data=(X_test, y_test),
    include_datasets=False,
)

License

MIT

Support

Questions and problem reports can be sent to contact@bitbullet.ai.