BitBullet
See BitBullet: bitbullet.co.uk
BitBullet is a fully managed, no-code machine learning platform. Describe your objective to the built-in AI agent, or use guided workstations, to create classification, regression, forecasting, and clustering experiments. You stay in control while BitBullet handles training infrastructure, tuning, diagnostics, and portable bundles.
The BitBullet SDK is the open-source Python toolkit for data scientists who want to write and orchestrate their own workflows with composable building blocks for transformations, training, forecasting, clustering, evaluation, and reproducible artefact metadata.
Full documentation and practical SDK tutorials are available here.
SDK Or Platform?
Use the BitBullet SDK when you want the raw Python building blocks to write and orchestrate your own data-science workflows.
Use BitBullet Platform when you want to configure and manage a supported modelling lifecycle through a guided, no-code environment. It centralises projects, datasets, managed compute, storage, and repeatable workflows while keeping you in control of every decision. Configure and compare experiments with clicks, or ask the AI assistant to prepare a draft for review; inspect the evidence, then export fitted artefacts, preprocessing, metadata, and generated inference code.
Explore BitBullet Platform | Start free
pip install bitbullet
Optional extras keep installations lean:
pip install "bitbullet[inference-models]" # LightGBM and XGBoost wrappers
pip install "bitbullet[inference-cluster]" # clustering extras (kmodes, umap, hdbscan)
pip install "bitbullet[forecast-statistical]" # optional StatsForecast adapter
pip install "bitbullet[train,viz]" # training, SHAP, and plotting tools
pip install "bitbullet[all]" # complete SDK
Modules
| Module | Purpose |
|---|---|
bitbullet.transform |
Fitted transformation pipelines for numerical, categorical, and datetime features. |
bitbullet.train |
Supervised training with Optuna-backed search, feature selection, sample weights, threshold optimisation, and reports. |
bitbullet.forecast |
Ordered and panel forecasting, reduction strategies, rolling-origin evaluation, intervals, and reconciliation. |
bitbullet.model_selection |
Ordered holdouts and temporal windows with explicit gaps and audit metadata. |
bitbullet.cluster |
K-Means, K-Modes, K-Prototypes, DBSCAN, GMM, gamma estimation, categorical weighting, and clustering metrics. |
bitbullet.evaluate |
Structured classification and regression evaluation metrics ready for reports and metadata. |
bitbullet.model |
Model wrappers, metadata, dataset metadata, and serialization helpers. |
Quick Examples
Transform Data
from bitbullet.transform import TransformPipeline
pipeline = TransformPipeline(name="credit_features")
pipeline.add("numerical", "standard_scale", columns=["income", "balance"])
pipeline.add("categorical", "onehot_encode", columns=["region"])
X_transformed = pipeline.fit_transform(X_train)
X_new = pipeline.transform(X_new_raw)
pipeline.save("artifacts/transform_pipeline.joblib")
Target-aware encoders receive y directly. target_encode is leakage-aware:
fit_transform(..., y=...) returns out-of-fold training encodings, while later
transform(...) calls use the stored full-training smoothed mapping.
pipeline = TransformPipeline()
pipeline.add(
"categorical",
"target_encode",
columns=["merchant_category"],
params={"target_type": "classification", "cv_folds": 5, "cv_strategy": "stratified"},
)
X_encoded = pipeline.fit_transform(X_train, y=y_train)
Train a Classifier
from bitbullet.train import TrainConfig, OptunaTrainer
config = TrainConfig(
name="default_risk_lgbm",
model_type="lgbm",
task="binary_classification",
n_trials=30,
optimization_metric="roc_auc",
optuna_sampler="tpe", # tpe, random, grid, cmaes
)
trainer = OptunaTrainer(config)
model = trainer.fit(X_train, y_train, X_val=X_val, y_val=y_val)
print(trainer.best_params)
print(trainer.state.optimal_threshold)
Evaluate a Model
from bitbullet.evaluate import evaluate_classification
report = evaluate_classification(
y_true=y_test,
y_pred_proba=model.predict_proba(X_test),
threshold=trainer.state.optimal_threshold or 0.5,
)
print(report.to_dict())
Forecast a Panel
from sklearn.linear_model import Ridge
from bitbullet.forecast import (
ForecastConfig,
ForecastFeatureBuilder,
ForecastFrame,
ForecastSchema,
TabularForecaster,
)
schema = ForecastSchema(
time="date",
targets="sales",
entities="store",
future_covariates="promotion",
cadence="D",
)
history = ForecastFrame(history_df, schema)
forecaster = TabularForecaster(
Ridge(alpha=1.0),
config=ForecastConfig(horizons=(1, 2, 3), strategy="direct"),
feature_builder=ForecastFeatureBuilder(target_lags=(1, 7)),
).fit(history)
predictions = forecaster.forecast(future_df)
Cluster Data
from bitbullet.cluster.core import ClusterConfig
from bitbullet.cluster.algorithms.partitional import KPrototypesClusterer
config = ClusterConfig(
name="customer_segments",
algorithm_type="partitional",
method="kprototypes",
n_clusters=5,
numerical_columns=["income", "spend"],
categorical_columns=["region", "channel"],
)
clusterer = KPrototypesClusterer(config)
labels = clusterer.fit_predict(df)
Save a Model with Metadata
from bitbullet.model import ModelMetadata, ModelSerializer
metadata = ModelMetadata(
name="default_risk_lgbm",
model_type=model.model_type,
framework=model.framework,
task="binary_classification",
metrics=report.metrics,
)
metadata.add_feature_schema(X_train)
ModelSerializer.save(
model=model,
path="artifacts/default_risk_lgbm.pkl",
metadata=metadata,
train_data=(X_train, y_train),
test_data=(X_test, y_test),
include_datasets=False,
)
License
MIT
Support
Questions and problem reports can be sent to contact@bitbullet.ai.