You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Selector to filter samples based on the frequency of a specified field.
This operator selects samples based on the frequency of values in a specified field. The field can be multi-level, with keys separated by dots. It supports filtering by either a top ratio or a fixed number (topk) of the most frequent values. If both top_ratio and topk are provided, the one resulting in fewer samples is used. The sorting order can be controlled with the reverse parameter. The operator processes the dataset and returns a new dataset containing only the selected samples.
Selector based on the specified value corresponding to the target key. The target key corresponding to multi-level field information need to be separated by '.'.
Ratio of selected top specified field value, samples will be selected if their specified field values are within this parameter. When both topk and top_ratio are set, the value corresponding to the smaller number of samples will be applied.
topk
typing.Optional[typing.Annotated[int, Gt(gt=0)]]
None
Number of selected top specified field value, samples will be selected if their specified field values are within this parameter. When both topk and top_ratio are set, the value corresponding to the smaller number of samples will be applied.
reverse
<class 'bool'>
True
Determine the sorting rule, if reverse=True, then sort in descending order.
The operator selects samples based on the frequency of 'meta.suffix' field, using a top ratio of 0.3 and a topk of 5, with reverse sorting. The target list contains the most frequent suffixes, '.pdf' and '.html', according to the specified criteria, while others are removed.
算子根据'meta.suffix'字段的频率选择样本,使用0.3的顶部比例和5的topk,并按降序排列。目标列表包含根据指定标准最频繁的后缀'.pdf'和'.html',而其他则被移除。
This example demonstrates the use of the operator with reverse set to False, selecting the least frequent values in the 'meta.key1.key2.count' field. Only two samples with the least frequent or null values in this field are kept, while all others are removed.
此示例展示了将reverse设置为False时算子的使用情况,选择'meta.key1.key2.count'字段中最不频繁的值。仅保留该字段中值最不频繁或为空的两个样本,其余全部被移除。