Skip to content

feat(data): add Iris dataset provider (#1044) - #1101

Merged
michalharakal merged 2 commits into
SKaiNET-developers:developfrom
AjithGoveas:feature/1044-iris-dataset-provider
Aug 24, 2026
Merged

michalharakal merged 2 commits into
SKaiNET-developers:developfrom
AjithGoveas:feature/1044-iris-dataset-provider

Conversation

@AjithGoveas

Copy link
Copy Markdown
Contributor

What

Adds a first-class Iris dataset provider to skainet-data-simple, enabling one-line access to the 1936 Fisher Iris dataset for classification tutorials.

val (train, test) = Iris.load().split(0.8, seed = 42L, stratified = true)

Files changed

  • Iris.kt — entry object: featureNames, classNames, suspend fun load()
  • IrisData.kt — IrisSample data class, embedded 150-row CSV (UCI canonical ordering, public domain), strict parser with error messages naming line + column
  • IrisDataset.kt — Dataset<FloatArray, Int>: FP32 [n,4] features, FP32 [n,3] one-hot labels. Both createDataBatch and createIndexedDataBatch overridden so split()/shuffle() views batch correctly
  • IrisDatasetTest.kt — 8 tests in jvmTest (same convention as MNIST/CIFAR tests)
  • README.md — Iris listed in built-in loaders
  • data-sources-getting-started.adoc — Iris usage snippet added

Design notes

  • Data is source-embedded (no file resource) to stay platform-agnostic including JS and Wasm browser targets, matching the repo's existing approach
  • load() is suspend for call-site symmetry with MNIST.loadTrain() / CIFAR10.load()
  • Raw centimetres returned; no standardisation (scaling should be fitted on training split only)
  • getY() returns Int, not one-hot — one-hot conversion is deferred to batch construction so stratified splitting works correctly

Checklist

AjithGoveas and others added 2 commits August 24, 2026 22:32
…orials (SKaiNET-developers#1044)

Add a first-class Iris dataset provider to skainet-data-simple, enabling:

    val (train, test) = Iris.load().split(0.8, seed = 42L, stratified = true)

Three new commonMain files:

- IrisData.kt: IrisSample data class with named fields (sepalLength,
  sepalWidth, petalLength, petalWidth, label), embedded 150-row CSV
  (UCI canonical ordering, public domain), and parseIrisCsv() with strict
  error messages naming line number and column.

- IrisDataset.kt: Dataset<FloatArray, Int> implementation. X is FP32
  [batch, 4] raw centimetres; Y is FP32 [batch, 3] one-hot. Both
  createDataBatch and createIndexedDataBatch are overridden (trap 1 from
  the issue) so that split(), shuffle() and filter() views batch correctly.
  getY() returns Int class index, not one-hot (trap 2).

- Iris.kt: entry object with featureNames, classNames, and suspend fun
  load(). The suspend modifier is for call-site symmetry with
  MNIST.loadTrain() and CIFAR10.load() even though loading is in-memory.

Embedded CSV (no network, no cache directory) keeps the provider
platform-agnostic including JS and Wasm browser targets where the KMP
resources plugin is not wired. IrisDatasetTest lives in jvmTest (same
convention as MNIST/FashionMNIST/CIFAR tests) and covers:

  - 150 samples, 50 per class
  - feature values in documented ranges (sepal length 4.3-7.9, petal width 0.1-2.5)
  - class indices match classNames ordering (setosa=0, versicolor=1, virginica=2)
  - dataBatch produces [n,4] and [n,3] tensors with correct one-hot rows
  - stratified split(0.8, seed=42L) yields 120/30 with 40/10 per class
  - split-then-batch regression test for trap 1 (non-contiguous index views)
  - seed determinism across load calls
  - parser error messages name line number and column/field

README.md and data-sources-getting-started.adoc updated to list Iris in
the built-in loaders section.

@michalharakal michalharakal left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

First of all: welcome, and thank you, the quality bar you set here is genuinely impressive.
The createIndexedDataBatch override for non-contiguous views, deferring one-hot to batch construction so stratified splitting works, and the test proving the split-then-batch path — that's exactly the subtle stuff that's usually missed. 🎉

I pushed one tiny fix directly to your branch (0dd4620) so we can merge without another round-trip: in the unknown-species error message, $Iris.classNames interpolated Iris.toString() plus the literal text .classNames — it needed braces, ${Iris.classNames}. I also added an assertion to the parser test so it would catch this in future.

Two purely optional, non-blocking notes — fine as follow-ups or to just leave as-is:

  1. The no-arg shuffle() override uses Random.Default, so it's non-deterministic. Probably intentional, but a short KDoc note pointing reproducibility-minded users at the seeded shuffle(seed) overload would help.
  2. The plain split(splitRatio) cuts at a raw index, and since the embedded data is ordered by species (50 per class), an unstratified 0.8 split puts every virginica into the test set. A KDoc warning recommending split(ratio, seed, stratified = true) would save someone a confusing afternoon.

Thanks again — looking forward to more! I am preparing another similar first good issues so we can reach funstion/Ops parity with KotlinDL :-) #290

@michalharakal
michalharakal merged commit 25f081d into SKaiNET-developers:develop Aug 24, 2026
16 checks passed
@AjithGoveas
AjithGoveas deleted the feature/1044-iris-dataset-provider branch August 25, 2026 07:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature]: Add an Iris dataset provider (and a tabular RawDataset -> Dataset bridge)

2 participants