This blog records some of my experiences to build a Python Package from zero experience for a UQ software.
[Demo] How to quantify the uncertainty when using LLM?
Product outline:

Target User Flow:
user data → Adapter → UnifiedSample
↓
Collate / DataLoader
↓
Batch
↓
Predictor.predict(batch) → Prediction
↓
UQMethod.quantify(...) → UncertainPrediction
↓
Evaluator.score(...) → Report

Package Layout:
MM_UQ/
├── pyproject.toml
├── README.md
├── LICENSE
├── src/MM_UQ/
│ ├── __init__.py # small public API only
│ ├── types.py # UnifiedSample, Batch, Prediction, ...
│ ├── config.py # RunConfig (one dataclass, see below)
│ ├── io/
│ ├── models/
│ ├── uq/
│ ├── eval/
│ └── demo/ # examples only
├── tests/
└── web/ # v1: stub only
External Tools:
Pytorch, matplotlib and other needed libs