-
Notifications
You must be signed in to change notification settings - Fork 83
Lazy statistics for columns #1492
Copy link
Copy link
Closed
Labels
enhancementNew feature or requestNew feature or requestperformanceSomething related to how fast the library can handle dataSomething related to how fast the library can handle dataresearchThis requires a deeper dive to gather a better understandingThis requires a deeper dive to gather a better understanding
Description
Activity
Metadata
Metadata
Assignees
Labels
enhancementNew feature or requestNew feature or requestperformanceSomething related to how fast the library can handle dataSomething related to how fast the library can handle dataresearchThis requires a deeper dive to gather a better understandingThis requires a deeper dive to gather a better understanding
Let's say you write
df.filter { someValue > df().myColumn.max() }This is way faster:
Maybe we could solve this by having lazily calculated stats stored inside
ValueColumns. Columns are immutable after all, so it would be safe to do so and the performance gain should be significant!Of course, this wouldn't work when you write:
df.filter { someValue > (myColumn + 1).max() } // or df.filter { someValue > myColumn.maxOf { it + 1 } }but that's okay I think. We can't have it all :)