feat(data): add Iris dataset provider (#1044) - #1101
michalharakal merged 2 commits into
Conversation
…orials (SKaiNET-developers#1044) Add a first-class Iris dataset provider to skainet-data-simple, enabling: val (train, test) = Iris.load().split(0.8, seed = 42L, stratified = true) Three new commonMain files: - IrisData.kt: IrisSample data class with named fields (sepalLength, sepalWidth, petalLength, petalWidth, label), embedded 150-row CSV (UCI canonical ordering, public domain), and parseIrisCsv() with strict error messages naming line number and column. - IrisDataset.kt: Dataset<FloatArray, Int> implementation. X is FP32 [batch, 4] raw centimetres; Y is FP32 [batch, 3] one-hot. Both createDataBatch and createIndexedDataBatch are overridden (trap 1 from the issue) so that split(), shuffle() and filter() views batch correctly. getY() returns Int class index, not one-hot (trap 2). - Iris.kt: entry object with featureNames, classNames, and suspend fun load(). The suspend modifier is for call-site symmetry with MNIST.loadTrain() and CIFAR10.load() even though loading is in-memory. Embedded CSV (no network, no cache directory) keeps the provider platform-agnostic including JS and Wasm browser targets where the KMP resources plugin is not wired. IrisDatasetTest lives in jvmTest (same convention as MNIST/FashionMNIST/CIFAR tests) and covers: - 150 samples, 50 per class - feature values in documented ranges (sepal length 4.3-7.9, petal width 0.1-2.5) - class indices match classNames ordering (setosa=0, versicolor=1, virginica=2) - dataBatch produces [n,4] and [n,3] tensors with correct one-hot rows - stratified split(0.8, seed=42L) yields 120/30 with 40/10 per class - split-then-batch regression test for trap 1 (non-contiguous index views) - seed determinism across load calls - parser error messages name line number and column/field README.md and data-sources-getting-started.adoc updated to list Iris in the built-in loaders section.
There was a problem hiding this comment.
First of all: welcome, and thank you, the quality bar you set here is genuinely impressive.
The createIndexedDataBatch override for non-contiguous views, deferring one-hot to batch construction so stratified splitting works, and the test proving the split-then-batch path — that's exactly the subtle stuff that's usually missed. 🎉
I pushed one tiny fix directly to your branch (0dd4620) so we can merge without another round-trip: in the unknown-species error message, $Iris.classNames interpolated Iris.toString() plus the literal text .classNames — it needed braces, ${Iris.classNames}. I also added an assertion to the parser test so it would catch this in future.
Two purely optional, non-blocking notes — fine as follow-ups or to just leave as-is:
- The no-arg
shuffle()override usesRandom.Default, so it's non-deterministic. Probably intentional, but a short KDoc note pointing reproducibility-minded users at the seededshuffle(seed)overload would help. - The plain
split(splitRatio)cuts at a raw index, and since the embedded data is ordered by species (50 per class), an unstratified 0.8 split puts every virginica into the test set. A KDoc warning recommendingsplit(ratio, seed, stratified = true)would save someone a confusing afternoon.
Thanks again — looking forward to more! I am preparing another similar first good issues so we can reach funstion/Ops parity with KotlinDL :-) #290
What
Adds a first-class Iris dataset provider to
skainet-data-simple, enabling one-line access to the 1936 Fisher Iris dataset for classification tutorials.Files changed
Design notes
Checklist