Alternatives
Self-hosted alternatives to scikit-learn
Projects in the registry that can run on your own hardware and serve at least one of the same use cases as scikit-learn.
- Source
- Registry relations over fetched repository records — the evidence sits on each repository page
- Verified
- Evidence not verified
- Confidence
- High
01How this list was built
- Compared against
- scikit-learn
- Delivery
- Libraries
- Shared use cases
- Prediction from operational data, Spreadsheet analysis, Document classification
- Selection rule
- open source, or self-hostable, and shares a use case
- Candidates found
- 6
02Alternatives
MLflow
Cloud platforms
Experiment tracking, a model registry and deployment packaging. The component that lets a prediction be traced back to the model version and data that produced it.
- Repository
- mlflow/mlflow
- Stars
- —
- Licence
- unknown
- In stacks
- 1
Shares: Predictive analytics
n8n
Hybrid — hosted or self-hosted
Workflow automation tool with several hundred integrations, branching logic, code steps and AI nodes. Can be self-hosted, which is why it appears in on-premise automation stacks.
- Repository
- n8n-io/n8n
- Stars
- 202,375
- Licence
- Other
- In stacks
- 4
Shares: Document classification
PaddleOCR
Libraries
OCR toolkit with detection, recognition and layout models, including Chinese and other East Asian scripts, plus table and formula recognition. Runs offline on CPU or GPU.
- Repository
- PaddlePaddle/PaddleOCR
- Stars
- —
- Licence
- unknown
- In stacks
- 0
Shares: Document classification
paperless-ngx
Self-hosted projects
Self-hosted document management system that OCRs incoming scans and email, applies tags and correspondents through trainable rules, and keeps a searchable archive.
- Repository
- paperless-ngx/paperless-ngx
- Stars
- 44,585
- Licence
- GPL-3.0
- In stacks
- 2
Shares: Document classification
Tesseract OCR
Libraries
Long-established open-source OCR engine with trained data for many languages and scripts, usable offline and embedded in most self-hosted document pipelines.
- Repository
- tesseract-ocr/tesseract
- Stars
- —
- Licence
- unknown
- In stacks
- 0
Shares: Document classification
XGBoost
Libraries
Gradient-boosted decision trees with distributed training and GPU support. The usual first model to try on a tabular prediction problem before anything more elaborate.
- Repository
- dmlc/xgboost
- Stars
- —
- Licence
- unknown
- In stacks
- 1
Shares: Predictive analytics · Spreadsheet analysis
03Repository facts
| Repository | Stars | Health |
|---|---|---|
| mlflow/mlflow | — | Health not measuredunknown |
| n8n-io/n8nFair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations. | 202k | Maintenance 100%, Releases 100%, Contributors 50%, Issues 95%, Community 100%92 |
| PaddlePaddle/PaddleOCR | — | Health not measuredunknown |
| paperless-ngx/paperless-ngxA community-supported supercharged document management system: scan, index and archive all your documents | 45k | Maintenance 100%, Releases 80%, Contributors 50%, Issues 100%, Community 93%87 |
| tesseract-ocr/tesseract | — | Health not measuredunknown |
| dmlc/xgboost | — | Health not measuredunknown |
04Ask about your own case
Which alternative is right depends on how many people will use it, what the data is and what you can operate. Ask and the answer is researched for your organisation.