Skip to content

Releases: guizmaii-opensource/csvzen

v0.6.0

Choose a tag to compare

@github-actions github-actions released this 10 Sep 00:21
c2c1874

What's Changed

v0.5.1

Choose a tag to compare

@github-actions github-actions released this 13 Jun 06:20
85316e4

What's Changed

v0.5.0

Choose a tag to compare

@github-actions github-actions released this 28 Apr 02:34
4ab7c91

What's Changed

  • Add FieldEmitter::emit(Option[A]) for ergonomic optional fields (#11) @guizmaii

v0.4.0

Choose a tag to compare

@github-actions github-actions released this 28 Apr 00:03
4725804

What's Changed

  • Add csvSinkDiscard: a ZSink variant that returns Unit (#10) @guizmaii

v0.3.0

Choose a tag to compare

@github-actions github-actions released this 25 Apr 12:19
4653010

csvzen 0.3.0 introduces csvzen-zio — a small, focused ZIO 2 integration on top of csvzen-core.

If you've been writing a CSV file from a ZStream[A] and gluing the resource handling together by hand, this is for you. Two helpers, no surprises:

  • openCsvWriter(path, config, …): ZIO[Scope, Throwable, CsvWriter] — opens a CsvWriter inside a Scope, closing it automatically at scope exit (success, failure, or interruption).
  • csvSink[A: CsvRowEncoder](path, config, …): ZSink[Any, Throwable, A, Nothing, Long] — writes a header row followed by one row per consumed A, returning the number of rows written. The underlying file handle is owned by the sink for its lifetime.

✨ Usage

libraryDependencies += "com.guizmaii" %% "csvzen-zio" % "0.3.0"
import com.guizmaii.csvzen.core.*
import com.guizmaii.csvzen.zio.*
import zio.*
import zio.stream.ZStream
import java.nio.file.Paths

final case class Person(name: String, age: Int) derives CsvRowEncoder

val people: ZStream[Any, Throwable, Person] = ZStream.fromIterable(
  Vector(Person("Ada", 36), Person("Linus", 55), Person("Grace", 85))
)

val rowsWritten: ZIO[Any, Throwable, Long] =
  people.run(csvSink[Person](Paths.get("people.csv"), CsvConfig.default))

The sink wraps CsvWriter lifecycle in a Scope, so abnormal terminations (failure, interruption) close the file handle deterministically. Setup runs inside a single ZIO.blocking shift; per-chunk writes use attemptBlocking. For tightly-locked execution on the blocking pool — no executor ping-pong between chunks — wrap the run call:

ZIO.blocking(stream.run(csvSink[Person](path, CsvConfig.default)))

🔧 Implementation note (for the curious)

csvSink is hand-rolled as a ZChannel rather than built on ZSink.foldLeftChunksZIO. The recursion uses @threadUnsafe lazy val loop + a captured var count, so the channel value is built once and reused per chunk (no per-chunk ZChannel.ReadWithCause allocation), and the row count is a stack-local mutation rather than a parameter threaded through. Each stream.run(sink) allocates a fresh closure, so concurrent or repeated runs of the same sink instance are isolated.

🛠 Other changes

Repository reorganisation

All published sub-projects now live under modules/:

modules/
├── core/        (csvzen-core)
├── test-kit/    (csvzen-test-kit)
└── zio/         (csvzen-zio)

The root keeps configuration, docs and build files separate from the published artefacts. No source changes — pure layout. If you're contributing or browsing source, paths are now modules/<name>/src/….

📦 Installation

libraryDependencies += "com.guizmaii" %% "csvzen-core"     % "0.3.0"
libraryDependencies += "com.guizmaii" %% "csvzen-test-kit" % "0.3.0" % Test
libraryDependencies += "com.guizmaii" %% "csvzen-zio"      % "0.3.0"  // optional

Targets Scala 3.3.7. JVM-only. csvzen-core has no runtime dependencies beyond the standard library; csvzen-zio pulls in dev.zio:zio + dev.zio:zio-streams.

Full diff: v0.2.0...v0.3.0

v0.2.0

Choose a tag to compare

@github-actions github-actions released this 25 Apr 09:25
6840a92

csvzen 0.2.0 introduces csvzen-test-kit — a zio-test integration providing golden-file (snapshot) testing for CSV encoder output.

To the author's knowledge, no other Scala CSV library ships golden testing. scala-csv, kantan-csv and zio-blocks/schema-csv all leave it to you to roll your own. csvzen 0.2.0 closes that gap, and the API mirrors zio-json-golden so anyone coming from the JSON side gets a familiar workflow with no relearning.

🤔 Why golden tests for CSV?

CSV output is a wire format. Once consumers exist — a downstream pipeline, a partner who imports the file daily, an Excel sheet someone wired up two years ago — your encoder's output is a contract. Any change in the bytes the encoder produces is a change in that contract: a column reordered, a date format tweaked, a CRLF turned into LF, a None rendered as "" instead of "null". Each is the kind of "harmless cleanup" that quietly breaks a consumer in production three weeks later.

A golden test is a small, opinionated answer:

  1. You commit a reference file — the golden — that captures what the encoder produces today for a representative set of inputs.
  2. On every test run, the encoder is re-executed against the same inputs and the output is compared byte-for-byte to the golden.
  3. If anything in the encoder's output changes, the test fails and shows you the diff.

The value comes from what it forces:

  • No silent format changes. Refactoring a CsvFieldEncoder, swapping a date library, "fixing" a quoting rule — all of it surfaces as a failing test with a visible diff, not as a wire-format regression that ships.
  • Cheap to write, dense in coverage. One csvGoldenTest(gen) call exercises 50 rows of randomised but stable input through the entire encoder stack. You don't write per-field assertions; you commit one file.
  • The diff is the spec. When the change is intentional, you open the _changed.csv next to the original, eyeball the diff to confirm the delta is what you wanted, and rename it over the original. The PR review then has a one-file diff that says exactly how the wire format moved.
  • Catches the boring stuff for free. Line-terminator drift, accidental quoting of a previously-unquoted column, an extra trailing newline, an Instant formatter that started emitting +00:00 instead of Z — all surface immediately, not at 3 AM in production.

The cost is one checked-in file per encoder shape and one rename when the contract intentionally changes. Worth it.

✨ Usage

libraryDependencies += "com.guizmaii" %% "csvzen-test-kit" % "0.2.0" % Test
import com.guizmaii.csvzen.core.*
import com.guizmaii.csvzen.testkit.*
import zio.test.*

object UserSpec extends ZIOSpecDefault {
  final case class User(id: Int, name: String, active: Boolean) derives CsvRowEncoder

  val gen: Gen[Sized, User] =
    for {
      id     <- Gen.int
      name   <- Gen.alphaNumericString
      active <- Gen.boolean
    } yield User(id, name, active)

  override def spec = suite("UserSpec")(
    csvGoldenTest(gen)
  )
}

Workflow

  • First run → writes src/test/resources/golden/User_new.csv and fails the test with "Remove _new from the suffix and re-run." That promotes the snapshot.
  • Subsequent runs → encoder output is compared byte-for-byte to the on-disk file. On mismatch a <Name>_changed.csv is written next to the original so you can diff. If the change is intentional, overwrite the original; if not, the test caught a regression.
  • No env-var auto-update mode. Promotion is always an explicit file rename.
  • On a passing run, leftover _changed.csv / _new.csv files are best-effort deleted so the workspace converges to clean.

Configuration

csvGoldenTest(
  gen,
  GoldenConfiguration(
    relativePath = "users",   // golden lives at src/test/resources/golden/users/User.csv
    sampleSize   = 50,        // default — bump for wider coverage
    csvConfig    = CsvConfig(delimiter = '\t', lineTerminator = "\n"),
  ),
)

GoldenConfiguration is a regular default-valued parameter — no implicit-config juggling at the call site.

📚 Docs

🛠 Other changes

  • CsvRowEncoder.derived scaladoc now flags the -Xmax-inlines requirement for case classes with ~25+ fields. Bump it in your build if you hit the "Maximal number of successive inlines (32) exceeded" compile error:
    scalacOptions ++= Seq("-Xmax-inlines:128")
  • CI plumbing fix for the release-drafter workflow (#6).

📦 Installation

libraryDependencies += "com.guizmaii" %% "csvzen-core"     % "0.2.0"
libraryDependencies += "com.guizmaii" %% "csvzen-test-kit" % "0.2.0" % Test

Targets Scala 3.3.7. JVM-only. csvzen-core has no runtime dependencies beyond the standard library; csvzen-test-kit pulls in zio-test, zio-test-magnolia, and zio for its compile-scope contract.

🙏 Acknowledgements

csvzen-test-kit is modelled directly on zio-json-golden — same workflow, same _new.csv / _changed.csv suffix dance, same GoldenConfiguration shape. Credit to the zio-json contributors for the design; csvzen-test-kit is the CSV-shaped translation.

Full diff: v0.1.0...v0.2.0

v0.1.0

Choose a tag to compare

@github-actions github-actions released this 25 Apr 07:13
06da197

csvzen is a zero-allocation streaming CSV writer for Scala 3 (LTS). RFC 4180-compliant, java.io.Writer-based, with compile-time-derived row encoders and hand-rolled Int / Long digit conversion for primitive cells.

This is the first public release.

✨ Features

Streaming writer

  • CsvWriter.open(path, config, charset = UTF_8, options*) — buffered, file-backed. Configurable Charset and any OpenOption (APPEND, CREATE_NEW, …).
  • Implements AutoCloseable and Flushable. Single-threaded by design.
  • Methods: writeHeader[A](), writeHeader(IndexedSeq[String]), writeRow[A], writeRow(FieldEmitter => Unit) (escape hatch), writeAll[A](Iterable[A]), flush(), close().

Encoders

  • CsvRowEncoder[A] with derives CsvRowEncoder for any flat case class — header names taken from field labels in declaration order.
  • CsvRowEncoder.custom(headers)(encode) for hand-built encoders: project a subset of fields, reorder columns, rename headers.
  • CsvFieldEncoder[A] shipped for primitives, BigInt, BigDecimal, UUID, Currency, every java.time.* codec (Instant, LocalDate*, OffsetDateTime, ZonedDateTime, Duration, Period, Year, …) and Option[A] for any of those.

Dialect

  • CsvConfig(delimiter, quoteChar, lineTerminator) — defaults to ",", "\"", "\r\n". Validated in the constructor; invalid combinations throw IllegalArgumentException.

RFC 4180 escaping

  • Quote-on-demand: only fields containing the delimiter, quote char, \r or \n are wrapped in quotes; embedded quotes are doubled.
  • Plain fields take a fast path with zero allocations and one Writer.write(s) call.

Zero-allocation hot path

  • Hand-rolled Int / Long digit emission via a reusable per-writer scratch: Array[Char]. No per-row allocations for String, primitives, Boolean, or Option[None].
  • Documented one-String-per-cell carve-out for Float, Double, BigInt, BigDecimal, UUID, Currency, and java.time.* types — same trade-off as zio-blocks' schema-csv.

📦 Installation

libraryDependencies += "com.guizmaii" %% "csvzen-core" % "0.1.0"

Targets Scala 3.3.7. JVM-only. No runtime dependencies beyond the standard library.

🙏 Acknowledgements

The codec concept and primitive-set are inspired by zio-blocks' schema-csv; the FieldEmitter design (owns-the-Writer, inline emit, two-pass escape scan) comes from a prior internal implementation.