{"criteria": {"false": "The answer to the question is no, using the supplied passage as evidence.", "true": "The answer to the question is yes, using the supplied passage as evidence."}, "experiment": "boolq", "id": "boolq-dev-0715", "question": "is there another part of avengers infinity war?", "source": {"dataset": "BoolQ", "kind": "published_human_label", "label_mapping": "target_yes = int(answer)", "license": "CC BY-SA 3.0", "license_url": "https://creativecommons.org/licenses/by-sa/3.0/", "mirror": "https://huggingface.co/datasets/google/boolq", "overlap_policy": "New calls for every expanded case; pilot predictions are never reused.", "paper": "https://arxiv.org/abs/1905.10044", "pilot_overlap": false, "published_label": true, "raw_jsonl_sha256": "e8fb84fbf510b022e963cddf3a3aded04151afa0ea0ef1cc1bf22f260ddd2344", "repository": "https://github.com/google-research-datasets/boolean-questions", "row_index_zero_based": 715, "split": "development"}, "state": {"passage": "Avengers: Infinity War held its world premiere on April 23, 2018 in Los Angeles and was released in the United States on April 27, 2018, in IMAX and 3D. The film received praise for the performances of the cast (particularly Brolin's) and the emotional weight of the story, as well as the visual effects and action sequences. It was the fourth film and the first superhero film to gross over $2 billion worldwide, breaking numerous box office records and becoming the highest-grossing film of 2018, as well as the fourth-highest-grossing film of all time and in the United States and Canada. The currently untitled sequel is set to be released on May 3, 2019."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the question is no, using the supplied passage as evidence.", "true": "The answer to the question is yes, using the supplied passage as evidence."}, "experiment": "boolq", "id": "boolq-dev-0221", "question": "are ncaa and nba balls the same size?", "source": {"dataset": "BoolQ", "kind": "published_human_label", "label_mapping": "target_yes = int(answer)", "license": "CC BY-SA 3.0", "license_url": "https://creativecommons.org/licenses/by-sa/3.0/", "mirror": "https://huggingface.co/datasets/google/boolq", "overlap_policy": "New calls for every expanded case; pilot predictions are never reused.", "paper": "https://arxiv.org/abs/1905.10044", "pilot_overlap": false, "published_label": false, "raw_jsonl_sha256": "e8fb84fbf510b022e963cddf3a3aded04151afa0ea0ef1cc1bf22f260ddd2344", "repository": "https://github.com/google-research-datasets/boolean-questions", "row_index_zero_based": 221, "split": "development"}, "state": {"passage": "A basketball (basketball ball) is a spherical ball used in basketball games. Basketballs typically range in size from very small promotional items only a few inches in diameter to extra large balls nearly a foot in diameter used in training exercises. For example, a youth basketball could be 27 inches (69 cm) in circumference, while an National Collegiate Athletic Association (NCAA) men's ball would be a maximum of 30 inches (76 cm) and an NCAA women's ball would be a maximum of 29 inches (74 cm). The standard for a basketball in the National Basketball Association (NBA) is 29.5 inches (75 cm) in circumference and for the Women's National Basketball Association (WNBA), a maximum circumference of 29 inches (74 cm). High school and junior leagues normally use NCAA, NBA or WNBA sized balls."}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the question is no, using the supplied passage as evidence.", "true": "The answer to the question is yes, using the supplied passage as evidence."}, "experiment": "boolq", "id": "boolq-dev-3090", "question": "is a lion part of the dog family?", "source": {"dataset": "BoolQ", "kind": "published_human_label", "label_mapping": "target_yes = int(answer)", "license": "CC BY-SA 3.0", "license_url": "https://creativecommons.org/licenses/by-sa/3.0/", "mirror": "https://huggingface.co/datasets/google/boolq", "overlap_policy": "New calls for every expanded case; pilot predictions are never reused.", "paper": "https://arxiv.org/abs/1905.10044", "pilot_overlap": false, "published_label": false, "raw_jsonl_sha256": "e8fb84fbf510b022e963cddf3a3aded04151afa0ea0ef1cc1bf22f260ddd2344", "repository": "https://github.com/google-research-datasets/boolean-questions", "row_index_zero_based": 3090, "split": "development"}, "state": {"passage": "The lion (Panthera leo) is a species in the cat family (Felidae); it is a muscular, deep-chested cat with a short, rounded head, a reduced neck and round ears, and a hairy tuft at the end of its tail. The lion is sexually dimorphic; males are larger than females with a typical weight range of 150 to 250 kg (331 to 551 lb) for the former and 120 to 182 kg (265 to 401 lb) for the latter. Male lions have a prominent mane, which is the most recognisable feature of the species. A lion pride consists of a few adult males, related females and cubs. Groups of female lions typically hunt together, preying mostly on large ungulates. The species is an apex and keystone predator, although they scavenge when opportunities occur. Some lions have been known to hunt humans, although the species typically does not."}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the question is no, using the supplied passage as evidence.", "true": "The answer to the question is yes, using the supplied passage as evidence."}, "experiment": "boolq", "id": "boolq-dev-1240", "question": "is there a main group element in period 6?", "source": {"dataset": "BoolQ", "kind": "published_human_label", "label_mapping": "target_yes = int(answer)", "license": "CC BY-SA 3.0", "license_url": "https://creativecommons.org/licenses/by-sa/3.0/", "mirror": "https://huggingface.co/datasets/google/boolq", "overlap_policy": "New calls for every expanded case; pilot predictions are never reused.", "paper": "https://arxiv.org/abs/1905.10044", "pilot_overlap": false, "published_label": false, "raw_jsonl_sha256": "e8fb84fbf510b022e963cddf3a3aded04151afa0ea0ef1cc1bf22f260ddd2344", "repository": "https://github.com/google-research-datasets/boolean-questions", "row_index_zero_based": 1240, "split": "development"}, "state": {"passage": "In chemistry and atomic physics, the main group is the group of elements whose lightest members are represented by helium, lithium, beryllium, boron, carbon, nitrogen, oxygen, and fluorine as arranged in the periodic table of the elements. The main group includes the elements (except hydrogen, which is sometimes not included) in groups 1 and 2 (s-block), and groups 13 to 18 (p-block). The s-block elements are primarily characterised by one main oxidation state, and the p-block elements, when they have multiple oxidation states, often have common oxidation states separated by two units."}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the question is no, using the supplied passage as evidence.", "true": "The answer to the question is yes, using the supplied passage as evidence."}, "experiment": "boolq", "id": "boolq-dev-1609", "question": "is soft tissue damage the same as a sprain?", "source": {"dataset": "BoolQ", "kind": "published_human_label", "label_mapping": "target_yes = int(answer)", "license": "CC BY-SA 3.0", "license_url": "https://creativecommons.org/licenses/by-sa/3.0/", "mirror": "https://huggingface.co/datasets/google/boolq", "overlap_policy": "New calls for every expanded case; pilot predictions are never reused.", "paper": "https://arxiv.org/abs/1905.10044", "pilot_overlap": false, "published_label": false, "raw_jsonl_sha256": "e8fb84fbf510b022e963cddf3a3aded04151afa0ea0ef1cc1bf22f260ddd2344", "repository": "https://github.com/google-research-datasets/boolean-questions", "row_index_zero_based": 1609, "split": "development"}, "state": {"passage": "A soft tissue injury (STI) is the damage of muscles, ligaments and tendons throughout the body. Common soft tissue injuries usually occur from a sprain, strain, a one off blow resulting in a contusion or overuse of a particular part of the body. Soft tissue injuries can result in pain, swelling, bruising and loss of function."}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the question is no, using the supplied passage as evidence.", "true": "The answer to the question is yes, using the supplied passage as evidence."}, "experiment": "boolq", "id": "boolq-dev-1124", "question": "does need for speed have a story mode?", "source": {"dataset": "BoolQ", "kind": "published_human_label", "label_mapping": "target_yes = int(answer)", "license": "CC BY-SA 3.0", "license_url": "https://creativecommons.org/licenses/by-sa/3.0/", "mirror": "https://huggingface.co/datasets/google/boolq", "overlap_policy": "New calls for every expanded case; pilot predictions are never reused.", "paper": "https://arxiv.org/abs/1905.10044", "pilot_overlap": false, "published_label": true, "raw_jsonl_sha256": "e8fb84fbf510b022e963cddf3a3aded04151afa0ea0ef1cc1bf22f260ddd2344", "repository": "https://github.com/google-research-datasets/boolean-questions", "row_index_zero_based": 1124, "split": "development"}, "state": {"passage": "The first gameplay footage of the pre-alpha build for Need for Speed was revealed at EA's press conference at E3 on June 15, 2015. The E3 presentation shows a part of the story, followed by the customization of a Subaru BRZ which showed the new and improved customization system, and the 'action camera' which was later revealed to be one of the five different camera angles. There are five different gameplay types: Speed, Style, Crew, Build, and Outlaw where players can earn points for engaging in to progress in the game through five overlapping storylines. Need for Speed takes place in the fictional city of Ventura Bay and its surroundings which is based on Los Angeles."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the question is no, using the supplied passage as evidence.", "true": "The answer to the question is yes, using the supplied passage as evidence."}, "experiment": "boolq", "id": "boolq-dev-2142", "question": "did estonia used to be part of russia?", "source": {"dataset": "BoolQ", "kind": "published_human_label", "label_mapping": "target_yes = int(answer)", "license": "CC BY-SA 3.0", "license_url": "https://creativecommons.org/licenses/by-sa/3.0/", "mirror": "https://huggingface.co/datasets/google/boolq", "overlap_policy": "New calls for every expanded case; pilot predictions are never reused.", "paper": "https://arxiv.org/abs/1905.10044", "pilot_overlap": false, "published_label": true, "raw_jsonl_sha256": "e8fb84fbf510b022e963cddf3a3aded04151afa0ea0ef1cc1bf22f260ddd2344", "repository": "https://github.com/google-research-datasets/boolean-questions", "row_index_zero_based": 2142, "split": "development"}, "state": {"passage": "Sweden's defeat by Russia in the Great Northern War resulted in the capitulation of Estonia and Livonia in 1710, confirmed by the Treaty of Nystad in 1721, and Russian rule was then imposed on what later became modern Estonia. Nonetheless, the legal system, Lutheran church, local and town governments, and education remained mostly German until the late 19th century and partially until 1918."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the question is no, using the supplied passage as evidence.", "true": "The answer to the question is yes, using the supplied passage as evidence."}, "experiment": "boolq", "id": "boolq-dev-0781", "question": "do ac milan and inter milan share a stadium?", "source": {"dataset": "BoolQ", "kind": "published_human_label", "label_mapping": "target_yes = int(answer)", "license": "CC BY-SA 3.0", "license_url": "https://creativecommons.org/licenses/by-sa/3.0/", "mirror": "https://huggingface.co/datasets/google/boolq", "overlap_policy": "New calls for every expanded case; pilot predictions are never reused.", "paper": "https://arxiv.org/abs/1905.10044", "pilot_overlap": false, "published_label": true, "raw_jsonl_sha256": "e8fb84fbf510b022e963cddf3a3aded04151afa0ea0ef1cc1bf22f260ddd2344", "repository": "https://github.com/google-research-datasets/boolean-questions", "row_index_zero_based": 781, "split": "development"}, "state": {"passage": "The Giuseppe Meazza Stadium (Italian pronunciation: (dʒuˈzɛppe meˈattsa)), commonly known as San Siro, is a football stadium in the San Siro district of Milan, Italy, which is the home of A.C. Milan and Inter Milan. It has a seating capacity of 80,018, making it one of the largest stadiums in Europe, and the largest in Italy."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the question is no, using the supplied passage as evidence.", "true": "The answer to the question is yes, using the supplied passage as evidence."}, "experiment": "boolq", "id": "boolq-dev-3044", "question": "is it legal to drink in public edinburgh?", "source": {"dataset": "BoolQ", "kind": "published_human_label", "label_mapping": "target_yes = int(answer)", "license": "CC BY-SA 3.0", "license_url": "https://creativecommons.org/licenses/by-sa/3.0/", "mirror": "https://huggingface.co/datasets/google/boolq", "overlap_policy": "New calls for every expanded case; pilot predictions are never reused.", "paper": "https://arxiv.org/abs/1905.10044", "pilot_overlap": false, "published_label": true, "raw_jsonl_sha256": "e8fb84fbf510b022e963cddf3a3aded04151afa0ea0ef1cc1bf22f260ddd2344", "repository": "https://github.com/google-research-datasets/boolean-questions", "row_index_zero_based": 3044, "split": "development"}, "state": {"passage": "The City of Edinburgh allows the consumption of alcohol in public places but under the Edinburgh by-law, anyone drinking in public would have to stop if asked by police. In the Strathclyde region that includes Glasgow, the consumption of alcohol or possession of an open container of alcohol, in public places has been illegal since 1996. Breaking this law can mean a fine. This ban was enforced due to the increase in drink-related violent crime. In the Perth & Kinross local authority the consumption of alcohol in public places is illegal in the following places: Alyth, Crieff, Kinross, Scone, Aberfeldy, Blairgowrie, Dunkeld & Birnam, Milnathort, Coupar Angus, Errol, Perth City. Drinking publicly in these areas is chargeable offence. In St Andrews in Fife it is illegal to drink or even have an open drinks container on the street. On the spot fines can be handed out by the police. It is however legal to consume alcohol on any of the beaches in St Andrews."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the question is no, using the supplied passage as evidence.", "true": "The answer to the question is yes, using the supplied passage as evidence."}, "experiment": "boolq", "id": "boolq-dev-2890", "question": "is there goal line technology in the world cup?", "source": {"dataset": "BoolQ", "kind": "published_human_label", "label_mapping": "target_yes = int(answer)", "license": "CC BY-SA 3.0", "license_url": "https://creativecommons.org/licenses/by-sa/3.0/", "mirror": "https://huggingface.co/datasets/google/boolq", "overlap_policy": "New calls for every expanded case; pilot predictions are never reused.", "paper": "https://arxiv.org/abs/1905.10044", "pilot_overlap": false, "published_label": true, "raw_jsonl_sha256": "e8fb84fbf510b022e963cddf3a3aded04151afa0ea0ef1cc1bf22f260ddd2344", "repository": "https://github.com/google-research-datasets/boolean-questions", "row_index_zero_based": 2890, "split": "development"}, "state": {"passage": "Compared to similar technology in other sports, goal-line technology is a relatively recent addition to association football; its integration having been opposed by the sport's authorities. In July 2012, the International Football Association Board (IFAB) officially approved the use of goal line technology, amending the Laws of the Game to permit (but not require) its use. Due to its expense, goal-line technology is only used at the highest levels of the game. Goal-line technology is currently used in the top European domestic leagues, and at major international competitions such as the 2014 Men's, 2018 Men's and 2015 Women's FIFA World Cups."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the question is no, using the supplied passage as evidence.", "true": "The answer to the question is yes, using the supplied passage as evidence."}, "experiment": "boolq", "id": "boolq-dev-2750", "question": "was looking for mr goodbar based on a true story?", "source": {"dataset": "BoolQ", "kind": "published_human_label", "label_mapping": "target_yes = int(answer)", "license": "CC BY-SA 3.0", "license_url": "https://creativecommons.org/licenses/by-sa/3.0/", "mirror": "https://huggingface.co/datasets/google/boolq", "overlap_policy": "New calls for every expanded case; pilot predictions are never reused.", "paper": "https://arxiv.org/abs/1905.10044", "pilot_overlap": false, "published_label": true, "raw_jsonl_sha256": "e8fb84fbf510b022e963cddf3a3aded04151afa0ea0ef1cc1bf22f260ddd2344", "repository": "https://github.com/google-research-datasets/boolean-questions", "row_index_zero_based": 2750, "split": "development"}, "state": {"passage": "Roseann Quinn (November 17, 1944 -- January 2, 1973) was an American schoolteacher in New York City who was stabbed to death in 1973. Her murder inspired Judith Rossner's best-selling 1975 novel Looking for Mr. Goodbar, which was adapted as a 1977 film directed by Richard Brooks and starring Diane Keaton. Quinn's murder also inspired the 1977 account Closing Time: The True Story of the ``Goodbar'' Murder by New York Times journalist Lacey Fosburgh. The case was the subject of a Season 3 episode of Investigation Discovery's series A Crime to Remember in 2015 (``Last Night Stand'')."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the question is no, using the supplied passage as evidence.", "true": "The answer to the question is yes, using the supplied passage as evidence."}, "experiment": "boolq", "id": "boolq-dev-0519", "question": "does age of ultron come after winter soldier?", "source": {"dataset": "BoolQ", "kind": "published_human_label", "label_mapping": "target_yes = int(answer)", "license": "CC BY-SA 3.0", "license_url": "https://creativecommons.org/licenses/by-sa/3.0/", "mirror": "https://huggingface.co/datasets/google/boolq", "overlap_policy": "New calls for every expanded case; pilot predictions are never reused.", "paper": "https://arxiv.org/abs/1905.10044", "pilot_overlap": false, "published_label": true, "raw_jsonl_sha256": "e8fb84fbf510b022e963cddf3a3aded04151afa0ea0ef1cc1bf22f260ddd2344", "repository": "https://github.com/google-research-datasets/boolean-questions", "row_index_zero_based": 519, "split": "development"}, "state": {"passage": "The first film in the series was Iron Man (2008), which was distributed by Paramount Pictures. Paramount also distributed Iron Man 2 (2010), Thor (2011) and Captain America: The First Avenger (2011), while Universal Pictures distributed The Incredible Hulk (2008). Walt Disney Studios Motion Pictures began distributing the films with the 2012 crossover film The Avengers, which concluded Phase One of the franchise. Phase Two includes Iron Man 3 (2013), Thor: The Dark World (2013), Captain America: The Winter Soldier (2014), Guardians of the Galaxy (2014), Avengers: Age of Ultron (2015), and Ant-Man (2015)."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the stated question is no.", "true": "The answer to the stated question is yes."}, "experiment": "policy", "id": "expanded-policy-0244", "question": "Is this item eligible for a refund under the supplied policy?", "source": {"anchor_pair_id": null, "generator": "expanded/datasets.py", "kind": "synthetic_programmatic", "pilot_overlap": false, "rule_family": "refund", "sampling": "100 per family, 50 per label, including 6 yes/no contrast pairs; remaining records uniform without replacement within label", "seed": 20261012, "semantic_pilot_overlap_ids": [], "shared_pilot_templates": true, "truth_validation": "Boolean oracle plus independently specified condition-contrast anchors"}, "state": {"policy": "This is a fictional shop policy. A recalled item is eligible for a refund regardless of all other facts. For an item that is not recalled, a receipt is required and final-sale items are never eligible. With those requirements met, an item is eligible if it is defective and purchased at most 90 days ago, or if it is unused and purchased at most 30 days ago. Boundaries are inclusive. No other exceptions apply. The item price and customer membership do not affect eligibility.", "record": {"days_since_purchase": 32, "defective": false, "final_sale": true, "recalled": false, "receipt": false, "unused": true}}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the stated question is no.", "true": "The answer to the stated question is yes."}, "experiment": "policy", "id": "expanded-policy-0075", "question": "Is this person allowed laboratory access under the supplied policy?", "source": {"anchor_pair_id": "access-contrast-5", "generator": "expanded/datasets.py", "kind": "synthetic_programmatic", "pilot_overlap": false, "rule_family": "access", "sampling": "100 per family, 50 per label, including 6 yes/no contrast pairs; remaining records uniform without replacement within label", "seed": 20261012, "semantic_pilot_overlap_ids": [], "shared_pilot_templates": true, "truth_validation": "Boolean oracle plus independently specified condition-contrast anchors"}, "state": {"policy": "This is a fictional laboratory access policy. The account must be active and must not be suspended. These two requirements always apply. An emergency permit then grants access regardless of role, training, clearance, or sponsor. Without an emergency permit, training must be complete, and the person must either be staff with clearance at least 2 or be a contractor with clearance at least 3 and a sponsor. Visitors never qualify without an emergency permit. Boundaries are inclusive. Years of experience and the badge color do not affect access.", "record": {"account_active": true, "clearance": 2, "emergency_permit": false, "role": "visitor", "sponsor": false, "suspended": false, "training_complete": true}}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the stated question is no.", "true": "The answer to the stated question is yes."}, "experiment": "policy", "id": "expanded-policy-0202", "question": "Is this item eligible for a refund under the supplied policy?", "source": {"anchor_pair_id": null, "generator": "expanded/datasets.py", "kind": "synthetic_programmatic", "pilot_overlap": false, "rule_family": "refund", "sampling": "100 per family, 50 per label, including 6 yes/no contrast pairs; remaining records uniform without replacement within label", "seed": 20261012, "semantic_pilot_overlap_ids": [], "shared_pilot_templates": true, "truth_validation": "Boolean oracle plus independently specified condition-contrast anchors"}, "state": {"policy": "This is a fictional shop policy. A recalled item is eligible for a refund regardless of all other facts. For an item that is not recalled, a receipt is required and final-sale items are never eligible. With those requirements met, an item is eligible if it is defective and purchased at most 90 days ago, or if it is unused and purchased at most 30 days ago. Boundaries are inclusive. No other exceptions apply. The item price and customer membership do not affect eligibility.", "record": {"days_since_purchase": 0, "defective": false, "final_sale": false, "recalled": false, "receipt": false, "unused": false}}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the stated question is no.", "true": "The answer to the stated question is yes."}, "experiment": "policy", "id": "expanded-policy-0185", "question": "Is this item eligible for a refund under the supplied policy?", "source": {"anchor_pair_id": null, "generator": "expanded/datasets.py", "kind": "synthetic_programmatic", "pilot_overlap": false, "rule_family": "refund", "sampling": "100 per family, 50 per label, including 6 yes/no contrast pairs; remaining records uniform without replacement within label", "seed": 20261012, "semantic_pilot_overlap_ids": [], "shared_pilot_templates": true, "truth_validation": "Boolean oracle plus independently specified condition-contrast anchors"}, "state": {"policy": "This is a fictional shop policy. A recalled item is eligible for a refund regardless of all other facts. For an item that is not recalled, a receipt is required and final-sale items are never eligible. With those requirements met, an item is eligible if it is defective and purchased at most 90 days ago, or if it is unused and purchased at most 30 days ago. Boundaries are inclusive. No other exceptions apply. The item price and customer membership do not affect eligibility.", "record": {"days_since_purchase": 30, "defective": false, "final_sale": true, "recalled": true, "receipt": false, "unused": false}}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the stated question is no.", "true": "The answer to the stated question is yes."}, "experiment": "policy", "id": "expanded-policy-0107", "question": "Is this item eligible for a refund under the supplied policy?", "source": {"anchor_pair_id": null, "generator": "expanded/datasets.py", "kind": "synthetic_programmatic", "pilot_overlap": false, "rule_family": "refund", "sampling": "100 per family, 50 per label, including 6 yes/no contrast pairs; remaining records uniform without replacement within label", "seed": 20261012, "semantic_pilot_overlap_ids": [], "shared_pilot_templates": true, "truth_validation": "Boolean oracle plus independently specified condition-contrast anchors"}, "state": {"policy": "This is a fictional shop policy. A recalled item is eligible for a refund regardless of all other facts. For an item that is not recalled, a receipt is required and final-sale items are never eligible. With those requirements met, an item is eligible if it is defective and purchased at most 90 days ago, or if it is unused and purchased at most 30 days ago. Boundaries are inclusive. No other exceptions apply. The item price and customer membership do not affect eligibility.", "record": {"days_since_purchase": 30, "defective": true, "final_sale": true, "recalled": true, "receipt": false, "unused": true}}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the stated question is no.", "true": "The answer to the stated question is yes."}, "experiment": "policy", "id": "expanded-policy-0110", "question": "Is this person allowed laboratory access under the supplied policy?", "source": {"anchor_pair_id": null, "generator": "expanded/datasets.py", "kind": "synthetic_programmatic", "pilot_overlap": false, "rule_family": "access", "sampling": "100 per family, 50 per label, including 6 yes/no contrast pairs; remaining records uniform without replacement within label", "seed": 20261012, "semantic_pilot_overlap_ids": [], "shared_pilot_templates": true, "truth_validation": "Boolean oracle plus independently specified condition-contrast anchors"}, "state": {"policy": "This is a fictional laboratory access policy. The account must be active and must not be suspended. These two requirements always apply. An emergency permit then grants access regardless of role, training, clearance, or sponsor. Without an emergency permit, training must be complete, and the person must either be staff with clearance at least 2 or be a contractor with clearance at least 3 and a sponsor. Visitors never qualify without an emergency permit. Boundaries are inclusive. Years of experience and the badge color do not affect access.", "record": {"account_active": true, "clearance": 1, "emergency_permit": true, "role": "contractor", "sponsor": false, "suspended": true, "training_complete": true}}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the stated question is no.", "true": "The answer to the stated question is yes."}, "experiment": "policy", "id": "expanded-policy-0049", "question": "Is this item eligible for a refund under the supplied policy?", "source": {"anchor_pair_id": null, "generator": "expanded/datasets.py", "kind": "synthetic_programmatic", "pilot_overlap": false, "rule_family": "refund", "sampling": "100 per family, 50 per label, including 6 yes/no contrast pairs; remaining records uniform without replacement within label", "seed": 20261012, "semantic_pilot_overlap_ids": [], "shared_pilot_templates": true, "truth_validation": "Boolean oracle plus independently specified condition-contrast anchors"}, "state": {"policy": "This is a fictional shop policy. A recalled item is eligible for a refund regardless of all other facts. For an item that is not recalled, a receipt is required and final-sale items are never eligible. With those requirements met, an item is eligible if it is defective and purchased at most 90 days ago, or if it is unused and purchased at most 30 days ago. Boundaries are inclusive. No other exceptions apply. The item price and customer membership do not affect eligibility.", "record": {"days_since_purchase": 29, "defective": false, "final_sale": true, "recalled": false, "receipt": true, "unused": false}}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the stated question is no.", "true": "The answer to the stated question is yes."}, "experiment": "policy", "id": "expanded-policy-0128", "question": "Does this order qualify for free delivery under the supplied policy?", "source": {"anchor_pair_id": null, "generator": "expanded/datasets.py", "kind": "synthetic_programmatic", "pilot_overlap": false, "rule_family": "delivery", "sampling": "100 per family, 50 per label, including 6 yes/no contrast pairs; remaining records uniform without replacement within label", "seed": 20261012, "semantic_pilot_overlap_ids": [], "shared_pilot_templates": true, "truth_validation": "Boolean oracle plus independently specified condition-contrast anchors"}, "state": {"policy": "This is a fictional delivery policy. Free delivery is available only for mainland destinations, non-hazardous orders, and total weight at most 20 kilograms. These requirements cannot be waived. An order meeting them qualifies if its subtotal is at least 100 credits, or if the customer is a member and the subtotal is at least 50 credits, or if the order has an active free-delivery promotion. An active promotion waives only the subtotal and membership requirements. Boundaries are inclusive. Gift wrapping and customer account age do not affect the decision.", "record": {"active_promotion": true, "hazardous": true, "mainland": false, "member": false, "subtotal_credits": 99, "weight_kg": 0}}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the stated question is no.", "true": "The answer to the stated question is yes."}, "experiment": "policy", "id": "expanded-policy-0092", "question": "Does this order qualify for free delivery under the supplied policy?", "source": {"anchor_pair_id": null, "generator": "expanded/datasets.py", "kind": "synthetic_programmatic", "pilot_overlap": false, "rule_family": "delivery", "sampling": "100 per family, 50 per label, including 6 yes/no contrast pairs; remaining records uniform without replacement within label", "seed": 20261012, "semantic_pilot_overlap_ids": [], "shared_pilot_templates": true, "truth_validation": "Boolean oracle plus independently specified condition-contrast anchors"}, "state": {"policy": "This is a fictional delivery policy. Free delivery is available only for mainland destinations, non-hazardous orders, and total weight at most 20 kilograms. These requirements cannot be waived. An order meeting them qualifies if its subtotal is at least 100 credits, or if the customer is a member and the subtotal is at least 50 credits, or if the order has an active free-delivery promotion. An active promotion waives only the subtotal and membership requirements. Boundaries are inclusive. Gift wrapping and customer account age do not affect the decision.", "record": {"active_promotion": true, "hazardous": true, "mainland": true, "member": true, "subtotal_credits": 51, "weight_kg": 21}}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the stated question is no.", "true": "The answer to the stated question is yes."}, "experiment": "policy", "id": "expanded-policy-0137", "question": "Does this order qualify for free delivery under the supplied policy?", "source": {"anchor_pair_id": null, "generator": "expanded/datasets.py", "kind": "synthetic_programmatic", "pilot_overlap": false, "rule_family": "delivery", "sampling": "100 per family, 50 per label, including 6 yes/no contrast pairs; remaining records uniform without replacement within label", "seed": 20261012, "semantic_pilot_overlap_ids": [], "shared_pilot_templates": true, "truth_validation": "Boolean oracle plus independently specified condition-contrast anchors"}, "state": {"policy": "This is a fictional delivery policy. Free delivery is available only for mainland destinations, non-hazardous orders, and total weight at most 20 kilograms. These requirements cannot be waived. An order meeting them qualifies if its subtotal is at least 100 credits, or if the customer is a member and the subtotal is at least 50 credits, or if the order has an active free-delivery promotion. An active promotion waives only the subtotal and membership requirements. Boundaries are inclusive. Gift wrapping and customer account age do not affect the decision.", "record": {"active_promotion": true, "hazardous": false, "mainland": true, "member": false, "subtotal_credits": 101, "weight_kg": 20}}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the stated question is no.", "true": "The answer to the stated question is yes."}, "experiment": "policy", "id": "expanded-policy-0204", "question": "Is this item eligible for a refund under the supplied policy?", "source": {"anchor_pair_id": null, "generator": "expanded/datasets.py", "kind": "synthetic_programmatic", "pilot_overlap": false, "rule_family": "refund", "sampling": "100 per family, 50 per label, including 6 yes/no contrast pairs; remaining records uniform without replacement within label", "seed": 20261012, "semantic_pilot_overlap_ids": [], "shared_pilot_templates": true, "truth_validation": "Boolean oracle plus independently specified condition-contrast anchors"}, "state": {"policy": "This is a fictional shop policy. A recalled item is eligible for a refund regardless of all other facts. For an item that is not recalled, a receipt is required and final-sale items are never eligible. With those requirements met, an item is eligible if it is defective and purchased at most 90 days ago, or if it is unused and purchased at most 30 days ago. Boundaries are inclusive. No other exceptions apply. The item price and customer membership do not affect eligibility.", "record": {"days_since_purchase": 60, "defective": false, "final_sale": true, "recalled": false, "receipt": true, "unused": false}}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The answer to the stated question is no.", "true": "The answer to the stated question is yes."}, "experiment": "policy", "id": "expanded-policy-0134", "question": "Does this order qualify for free delivery under the supplied policy?", "source": {"anchor_pair_id": null, "generator": "expanded/datasets.py", "kind": "synthetic_programmatic", "pilot_overlap": false, "rule_family": "delivery", "sampling": "100 per family, 50 per label, including 6 yes/no contrast pairs; remaining records uniform without replacement within label", "seed": 20261012, "semantic_pilot_overlap_ids": [], "shared_pilot_templates": true, "truth_validation": "Boolean oracle plus independently specified condition-contrast anchors"}, "state": {"policy": "This is a fictional delivery policy. Free delivery is available only for mainland destinations, non-hazardous orders, and total weight at most 20 kilograms. These requirements cannot be waived. An order meeting them qualifies if its subtotal is at least 100 credits, or if the customer is a member and the subtotal is at least 50 credits, or if the order has an active free-delivery promotion. An active promotion waives only the subtotal and membership requirements. Boundaries are inclusive. Gift wrapping and customer account age do not affect the decision.", "record": {"active_promotion": false, "hazardous": false, "mainland": true, "member": true, "subtotal_credits": 101, "weight_kg": 19.9}}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The stated event does not occur in that random selection.", "true": "The stated event occurs in the random selection described in the state."}, "experiment": "probability", "id": "expanded-probability-0098", "question": "Is at least one of the two selected items marked?", "source": {"actual_rng_seed": 20261104, "batch_index": 1, "generator": "expanded/datasets.py using pinned pilot probability generator", "generator_version": 1, "kind": "synthetic_exact_probability", "pilot_overlap": false, "probability_family": "without_replacement", "seed": 20261103, "selection": "Keep first unique generated problems until 100 per family; exclude pilot-identical inputs.", "shared_pilot_templates": true, "truth_denominator": 20, "truth_numerator": 7, "validation": "Exact rational probability from the complete population and stated random selection mechanism"}, "state": {"counts": {"marked": 3, "unmarked": 13}, "event": "at_least_one_marked", "interpretation": "Estimate the probability of the event for these two random draws, using the supplied counts.", "selection": "Select two different items uniformly at random without replacement. Every unordered pair is equally likely.", "setting": "A fictional box contains the entire population described by these counts."}, "target": 0.35, "target_kind": "probability"}
{"criteria": {"false": "The stated event does not occur in that random selection.", "true": "The stated event occurs in the random selection described in the state."}, "experiment": "probability", "id": "expanded-probability-0080", "question": "Is the selected item marked?", "source": {"actual_rng_seed": 20261104, "batch_index": 1, "generator": "expanded/datasets.py using pinned pilot probability generator", "generator_version": 1, "kind": "synthetic_exact_probability", "pilot_overlap": false, "probability_family": "weighted_mixture", "seed": 20261103, "selection": "Keep first unique generated problems until 100 per family; exclude pilot-identical inputs.", "shared_pilot_templates": true, "truth_denominator": 3025, "truth_numerator": 1682, "validation": "Exact rational probability from the complete population and stated random selection mechanism"}, "state": {"box_selection_tickets": {"A": 8, "B": 3}, "boxes": {"A": {"marked": 28, "unmarked": 22}, "B": {"marked": 6, "unmarked": 5}}, "interpretation": "Estimate the probability of the event for this random item, accounting for both stages of selection.", "selection": "First choose one ticket uniformly at random from all tickets. The ticket names box A or box B. Then choose one item uniformly at random from that box. There is no further information about the chosen box.", "setting": "Two fictional boxes contain complete populations. Their contents are listed below."}, "target": 0.5560330578512397, "target_kind": "probability"}
{"criteria": {"false": "The stated event does not occur in that random selection.", "true": "The stated event occurs in the random selection described in the state."}, "experiment": "probability", "id": "expanded-probability-0127", "question": "Are both selected items marked?", "source": {"actual_rng_seed": 20261105, "batch_index": 2, "generator": "expanded/datasets.py using pinned pilot probability generator", "generator_version": 1, "kind": "synthetic_exact_probability", "pilot_overlap": false, "probability_family": "without_replacement", "seed": 20261104, "selection": "Keep first unique generated problems until 100 per family; exclude pilot-identical inputs.", "shared_pilot_templates": true, "truth_denominator": 561, "truth_numerator": 325, "validation": "Exact rational probability from the complete population and stated random selection mechanism"}, "state": {"counts": {"marked": 26, "unmarked": 8}, "event": "both_marked", "interpretation": "Estimate the probability of the event for these two random draws, using the supplied counts.", "selection": "Select two different items uniformly at random without replacement. Every unordered pair is equally likely.", "setting": "A fictional box contains the entire population described by these counts."}, "target": 0.5793226381461676, "target_kind": "probability"}
{"criteria": {"false": "The stated event does not occur in that random selection.", "true": "The stated event occurs in the random selection described in the state."}, "experiment": "probability", "id": "expanded-probability-0038", "question": "Is at least one of the two selected items marked?", "source": {"actual_rng_seed": 20261103, "batch_index": 0, "generator": "expanded/datasets.py using pinned pilot probability generator", "generator_version": 1, "kind": "synthetic_exact_probability", "pilot_overlap": false, "probability_family": "without_replacement", "seed": 20261102, "selection": "Keep first unique generated problems until 100 per family; exclude pilot-identical inputs.", "shared_pilot_templates": true, "truth_denominator": 260, "truth_numerator": 203, "validation": "Exact rational probability from the complete population and stated random selection mechanism"}, "state": {"counts": {"marked": 21, "unmarked": 19}, "event": "at_least_one_marked", "interpretation": "Estimate the probability of the event for these two random draws, using the supplied counts.", "selection": "Select two different items uniformly at random without replacement. Every unordered pair is equally likely.", "setting": "A fictional box contains the entire population described by these counts."}, "target": 0.7807692307692308, "target_kind": "probability"}
{"criteria": {"false": "The stated event does not occur in that random selection.", "true": "The stated event occurs in the random selection described in the state."}, "experiment": "probability", "id": "expanded-probability-0076", "question": "Are both selected items marked?", "source": {"actual_rng_seed": 20261104, "batch_index": 1, "generator": "expanded/datasets.py using pinned pilot probability generator", "generator_version": 1, "kind": "synthetic_exact_probability", "pilot_overlap": false, "probability_family": "without_replacement", "seed": 20261103, "selection": "Keep first unique generated problems until 100 per family; exclude pilot-identical inputs.", "shared_pilot_templates": true, "truth_denominator": 15, "truth_numerator": 14, "validation": "Exact rational probability from the complete population and stated random selection mechanism"}, "state": {"counts": {"marked": 29, "unmarked": 1}, "event": "both_marked", "interpretation": "Estimate the probability of the event for these two random draws, using the supplied counts.", "selection": "Select two different items uniformly at random without replacement. Every unordered pair is equally likely.", "setting": "A fictional box contains the entire population described by these counts."}, "target": 0.9333333333333333, "target_kind": "probability"}
{"criteria": {"false": "The stated event does not occur in that random selection.", "true": "The stated event occurs in the random selection described in the state."}, "experiment": "probability", "id": "expanded-probability-0216", "question": "Are both selected items marked?", "source": {"actual_rng_seed": 20261106, "batch_index": 3, "generator": "expanded/datasets.py using pinned pilot probability generator", "generator_version": 1, "kind": "synthetic_exact_probability", "pilot_overlap": false, "probability_family": "without_replacement", "seed": 20261105, "selection": "Keep first unique generated problems until 100 per family; exclude pilot-identical inputs.", "shared_pilot_templates": true, "truth_denominator": 52, "truth_numerator": 29, "validation": "Exact rational probability from the complete population and stated random selection mechanism"}, "state": {"counts": {"marked": 30, "unmarked": 10}, "event": "both_marked", "interpretation": "Estimate the probability of the event for these two random draws, using the supplied counts.", "selection": "Select two different items uniformly at random without replacement. Every unordered pair is equally likely.", "setting": "A fictional box contains the entire population described by these counts."}, "target": 0.5576923076923077, "target_kind": "probability"}
{"criteria": {"false": "The stated event does not occur in that random selection.", "true": "The stated event occurs in the random selection described in the state."}, "experiment": "probability", "id": "expanded-probability-0219", "question": "Are both selected items marked?", "source": {"actual_rng_seed": 20261106, "batch_index": 3, "generator": "expanded/datasets.py using pinned pilot probability generator", "generator_version": 1, "kind": "synthetic_exact_probability", "pilot_overlap": false, "probability_family": "without_replacement", "seed": 20261105, "selection": "Keep first unique generated problems until 100 per family; exclude pilot-identical inputs.", "shared_pilot_templates": true, "truth_denominator": 190, "truth_numerator": 153, "validation": "Exact rational probability from the complete population and stated random selection mechanism"}, "state": {"counts": {"marked": 18, "unmarked": 2}, "event": "both_marked", "interpretation": "Estimate the probability of the event for these two random draws, using the supplied counts.", "selection": "Select two different items uniformly at random without replacement. Every unordered pair is equally likely.", "setting": "A fictional box contains the entire population described by these counts."}, "target": 0.8052631578947368, "target_kind": "probability"}
{"criteria": {"false": "The stated event does not occur in that random selection.", "true": "The stated event occurs in the random selection described in the state."}, "experiment": "probability", "id": "expanded-probability-0113", "question": "Is the selected item marked?", "source": {"actual_rng_seed": 20261104, "batch_index": 1, "generator": "expanded/datasets.py using pinned pilot probability generator", "generator_version": 1, "kind": "synthetic_exact_probability", "pilot_overlap": false, "probability_family": "conditional_table", "seed": 20261103, "selection": "Keep first unique generated problems until 100 per family; exclude pilot-identical inputs.", "shared_pilot_templates": true, "truth_denominator": 5, "truth_numerator": 1, "validation": "Exact rational probability from the complete population and stated random selection mechanism"}, "state": {"counts": {"A": {"marked": 8, "unmarked": 32}, "B": {"marked": 52, "unmarked": 38}}, "interpretation": "Estimate the probability of the event for this one random draw, using the supplied counts.", "selected_group": "A", "selection": "Choose one item uniformly at random from group A only.", "setting": "Fictional inventory. The following table describes the entire population, not a sample."}, "target": 0.2, "target_kind": "probability"}
{"criteria": {"false": "The stated event does not occur in that random selection.", "true": "The stated event occurs in the random selection described in the state."}, "experiment": "probability", "id": "expanded-probability-0190", "question": "Is the selected item marked?", "source": {"actual_rng_seed": 20261106, "batch_index": 3, "generator": "expanded/datasets.py using pinned pilot probability generator", "generator_version": 1, "kind": "synthetic_exact_probability", "pilot_overlap": false, "probability_family": "conditional_table", "seed": 20261105, "selection": "Keep first unique generated problems until 100 per family; exclude pilot-identical inputs.", "shared_pilot_templates": true, "truth_denominator": 5, "truth_numerator": 1, "validation": "Exact rational probability from the complete population and stated random selection mechanism"}, "state": {"counts": {"A": {"marked": 7, "unmarked": 4}, "B": {"marked": 6, "unmarked": 24}}, "interpretation": "Estimate the probability of the event for this one random draw, using the supplied counts.", "selected_group": "B", "selection": "Choose one item uniformly at random from group B only.", "setting": "Fictional inventory. The following table describes the entire population, not a sample."}, "target": 0.2, "target_kind": "probability"}
{"criteria": {"false": "The stated event does not occur in that random selection.", "true": "The stated event occurs in the random selection described in the state."}, "experiment": "probability", "id": "expanded-probability-0032", "question": "Are both selected items marked?", "source": {"actual_rng_seed": 20261103, "batch_index": 0, "generator": "expanded/datasets.py using pinned pilot probability generator", "generator_version": 1, "kind": "synthetic_exact_probability", "pilot_overlap": false, "probability_family": "without_replacement", "seed": 20261102, "selection": "Keep first unique generated problems until 100 per family; exclude pilot-identical inputs.", "shared_pilot_templates": true, "truth_denominator": 30, "truth_numerator": 1, "validation": "Exact rational probability from the complete population and stated random selection mechanism"}, "state": {"counts": {"marked": 7, "unmarked": 29}, "event": "both_marked", "interpretation": "Estimate the probability of the event for these two random draws, using the supplied counts.", "selection": "Select two different items uniformly at random without replacement. Every unordered pair is equally likely.", "setting": "A fictional box contains the entire population described by these counts."}, "target": 0.03333333333333333, "target_kind": "probability"}
{"criteria": {"false": "The stated event does not occur in that random selection.", "true": "The stated event occurs in the random selection described in the state."}, "experiment": "probability", "id": "expanded-probability-0118", "question": "Is at least one of the two selected items marked?", "source": {"actual_rng_seed": 20261104, "batch_index": 1, "generator": "expanded/datasets.py using pinned pilot probability generator", "generator_version": 1, "kind": "synthetic_exact_probability", "pilot_overlap": false, "probability_family": "without_replacement", "seed": 20261103, "selection": "Keep first unique generated problems until 100 per family; exclude pilot-identical inputs.", "shared_pilot_templates": true, "truth_denominator": 1, "truth_numerator": 1, "validation": "Exact rational probability from the complete population and stated random selection mechanism"}, "state": {"counts": {"marked": 43, "unmarked": 0}, "event": "at_least_one_marked", "interpretation": "Estimate the probability of the event for these two random draws, using the supplied counts.", "selection": "Select two different items uniformly at random without replacement. Every unordered pair is equally likely.", "setting": "A fictional box contains the entire population described by these counts."}, "target": 1.0, "target_kind": "probability"}
{"criteria": {"false": "The stated event does not occur in that random selection.", "true": "The stated event occurs in the random selection described in the state."}, "experiment": "probability", "id": "expanded-probability-0143", "question": "Is the selected item marked?", "source": {"actual_rng_seed": 20261105, "batch_index": 2, "generator": "expanded/datasets.py using pinned pilot probability generator", "generator_version": 1, "kind": "synthetic_exact_probability", "pilot_overlap": false, "probability_family": "conditional_table", "seed": 20261104, "selection": "Keep first unique generated problems until 100 per family; exclude pilot-identical inputs.", "shared_pilot_templates": true, "truth_denominator": 1, "truth_numerator": 0, "validation": "Exact rational probability from the complete population and stated random selection mechanism"}, "state": {"counts": {"A": {"marked": 0, "unmarked": 20}, "B": {"marked": 5, "unmarked": 60}}, "interpretation": "Estimate the probability of the event for this one random draw, using the supplied counts.", "selected_group": "A", "selection": "Choose one item uniformly at random from group A only.", "setting": "Fictional inventory. The following table describes the entire population, not a sample."}, "target": 0.0, "target_kind": "probability"}
{"criteria": {"false": "The hypothesis does not follow from the premise; it may be unsupported or contradicted.", "true": "The hypothesis follows from the premise."}, "experiment": "rte", "id": "rte-validation-0198", "question": "Does the premise entail the hypothesis?", "source": {"configuration": "rte", "dataset": "SuperGLUE RTE", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=entailment, 1=not_entailment; target_yes = 1-label", "license": "Original dataset terms; no unified license verified", "license_status": "unresolved_original_RTE_terms", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 1, "raw_jsonl_sha256": "eec4bac5953538dc63792265f59f7df4f68c0048bcacd3d07336b7d7dc60db40", "row_index_zero_based": 198, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 198, "split": "validation", "task_source": "https://www.tensorflow.org/datasets/catalog/super_glue#super_gluerte"}, "state": {"retrieve_from": "https://huggingface.co/datasets/aps/super_glue/viewer/rte/validation", "source_idx": 198, "source_text_omitted": "RTE source text omitted from public export; retrieve the original row from the linked source."}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The hypothesis does not follow from the premise; it may be unsupported or contradicted.", "true": "The hypothesis follows from the premise."}, "experiment": "rte", "id": "rte-validation-0188", "question": "Does the premise entail the hypothesis?", "source": {"configuration": "rte", "dataset": "SuperGLUE RTE", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=entailment, 1=not_entailment; target_yes = 1-label", "license": "Original dataset terms; no unified license verified", "license_status": "unresolved_original_RTE_terms", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 0, "raw_jsonl_sha256": "eec4bac5953538dc63792265f59f7df4f68c0048bcacd3d07336b7d7dc60db40", "row_index_zero_based": 188, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 188, "split": "validation", "task_source": "https://www.tensorflow.org/datasets/catalog/super_glue#super_gluerte"}, "state": {"retrieve_from": "https://huggingface.co/datasets/aps/super_glue/viewer/rte/validation", "source_idx": 188, "source_text_omitted": "RTE source text omitted from public export; retrieve the original row from the linked source."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The hypothesis does not follow from the premise; it may be unsupported or contradicted.", "true": "The hypothesis follows from the premise."}, "experiment": "rte", "id": "rte-validation-0130", "question": "Does the premise entail the hypothesis?", "source": {"configuration": "rte", "dataset": "SuperGLUE RTE", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=entailment, 1=not_entailment; target_yes = 1-label", "license": "Original dataset terms; no unified license verified", "license_status": "unresolved_original_RTE_terms", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 1, "raw_jsonl_sha256": "eec4bac5953538dc63792265f59f7df4f68c0048bcacd3d07336b7d7dc60db40", "row_index_zero_based": 130, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 130, "split": "validation", "task_source": "https://www.tensorflow.org/datasets/catalog/super_glue#super_gluerte"}, "state": {"retrieve_from": "https://huggingface.co/datasets/aps/super_glue/viewer/rte/validation", "source_idx": 130, "source_text_omitted": "RTE source text omitted from public export; retrieve the original row from the linked source."}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The hypothesis does not follow from the premise; it may be unsupported or contradicted.", "true": "The hypothesis follows from the premise."}, "experiment": "rte", "id": "rte-validation-0143", "question": "Does the premise entail the hypothesis?", "source": {"configuration": "rte", "dataset": "SuperGLUE RTE", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=entailment, 1=not_entailment; target_yes = 1-label", "license": "Original dataset terms; no unified license verified", "license_status": "unresolved_original_RTE_terms", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 0, "raw_jsonl_sha256": "eec4bac5953538dc63792265f59f7df4f68c0048bcacd3d07336b7d7dc60db40", "row_index_zero_based": 143, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 143, "split": "validation", "task_source": "https://www.tensorflow.org/datasets/catalog/super_glue#super_gluerte"}, "state": {"retrieve_from": "https://huggingface.co/datasets/aps/super_glue/viewer/rte/validation", "source_idx": 143, "source_text_omitted": "RTE source text omitted from public export; retrieve the original row from the linked source."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The hypothesis does not follow from the premise; it may be unsupported or contradicted.", "true": "The hypothesis follows from the premise."}, "experiment": "rte", "id": "rte-validation-0167", "question": "Does the premise entail the hypothesis?", "source": {"configuration": "rte", "dataset": "SuperGLUE RTE", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=entailment, 1=not_entailment; target_yes = 1-label", "license": "Original dataset terms; no unified license verified", "license_status": "unresolved_original_RTE_terms", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 1, "raw_jsonl_sha256": "eec4bac5953538dc63792265f59f7df4f68c0048bcacd3d07336b7d7dc60db40", "row_index_zero_based": 167, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 167, "split": "validation", "task_source": "https://www.tensorflow.org/datasets/catalog/super_glue#super_gluerte"}, "state": {"retrieve_from": "https://huggingface.co/datasets/aps/super_glue/viewer/rte/validation", "source_idx": 167, "source_text_omitted": "RTE source text omitted from public export; retrieve the original row from the linked source."}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The hypothesis does not follow from the premise; it may be unsupported or contradicted.", "true": "The hypothesis follows from the premise."}, "experiment": "rte", "id": "rte-validation-0011", "question": "Does the premise entail the hypothesis?", "source": {"configuration": "rte", "dataset": "SuperGLUE RTE", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=entailment, 1=not_entailment; target_yes = 1-label", "license": "Original dataset terms; no unified license verified", "license_status": "unresolved_original_RTE_terms", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 1, "raw_jsonl_sha256": "eec4bac5953538dc63792265f59f7df4f68c0048bcacd3d07336b7d7dc60db40", "row_index_zero_based": 11, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 11, "split": "validation", "task_source": "https://www.tensorflow.org/datasets/catalog/super_glue#super_gluerte"}, "state": {"retrieve_from": "https://huggingface.co/datasets/aps/super_glue/viewer/rte/validation", "source_idx": 11, "source_text_omitted": "RTE source text omitted from public export; retrieve the original row from the linked source."}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The hypothesis does not follow from the premise; it may be unsupported or contradicted.", "true": "The hypothesis follows from the premise."}, "experiment": "rte", "id": "rte-validation-0009", "question": "Does the premise entail the hypothesis?", "source": {"configuration": "rte", "dataset": "SuperGLUE RTE", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=entailment, 1=not_entailment; target_yes = 1-label", "license": "Original dataset terms; no unified license verified", "license_status": "unresolved_original_RTE_terms", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 0, "raw_jsonl_sha256": "eec4bac5953538dc63792265f59f7df4f68c0048bcacd3d07336b7d7dc60db40", "row_index_zero_based": 9, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 9, "split": "validation", "task_source": "https://www.tensorflow.org/datasets/catalog/super_glue#super_gluerte"}, "state": {"retrieve_from": "https://huggingface.co/datasets/aps/super_glue/viewer/rte/validation", "source_idx": 9, "source_text_omitted": "RTE source text omitted from public export; retrieve the original row from the linked source."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The hypothesis does not follow from the premise; it may be unsupported or contradicted.", "true": "The hypothesis follows from the premise."}, "experiment": "rte", "id": "rte-validation-0170", "question": "Does the premise entail the hypothesis?", "source": {"configuration": "rte", "dataset": "SuperGLUE RTE", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=entailment, 1=not_entailment; target_yes = 1-label", "license": "Original dataset terms; no unified license verified", "license_status": "unresolved_original_RTE_terms", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 0, "raw_jsonl_sha256": "eec4bac5953538dc63792265f59f7df4f68c0048bcacd3d07336b7d7dc60db40", "row_index_zero_based": 170, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 170, "split": "validation", "task_source": "https://www.tensorflow.org/datasets/catalog/super_glue#super_gluerte"}, "state": {"retrieve_from": "https://huggingface.co/datasets/aps/super_glue/viewer/rte/validation", "source_idx": 170, "source_text_omitted": "RTE source text omitted from public export; retrieve the original row from the linked source."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The hypothesis does not follow from the premise; it may be unsupported or contradicted.", "true": "The hypothesis follows from the premise."}, "experiment": "rte", "id": "rte-validation-0095", "question": "Does the premise entail the hypothesis?", "source": {"configuration": "rte", "dataset": "SuperGLUE RTE", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=entailment, 1=not_entailment; target_yes = 1-label", "license": "Original dataset terms; no unified license verified", "license_status": "unresolved_original_RTE_terms", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 1, "raw_jsonl_sha256": "eec4bac5953538dc63792265f59f7df4f68c0048bcacd3d07336b7d7dc60db40", "row_index_zero_based": 95, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 95, "split": "validation", "task_source": "https://www.tensorflow.org/datasets/catalog/super_glue#super_gluerte"}, "state": {"retrieve_from": "https://huggingface.co/datasets/aps/super_glue/viewer/rte/validation", "source_idx": 95, "source_text_omitted": "RTE source text omitted from public export; retrieve the original row from the linked source."}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The hypothesis does not follow from the premise; it may be unsupported or contradicted.", "true": "The hypothesis follows from the premise."}, "experiment": "rte", "id": "rte-validation-0064", "question": "Does the premise entail the hypothesis?", "source": {"configuration": "rte", "dataset": "SuperGLUE RTE", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=entailment, 1=not_entailment; target_yes = 1-label", "license": "Original dataset terms; no unified license verified", "license_status": "unresolved_original_RTE_terms", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 1, "raw_jsonl_sha256": "eec4bac5953538dc63792265f59f7df4f68c0048bcacd3d07336b7d7dc60db40", "row_index_zero_based": 64, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 64, "split": "validation", "task_source": "https://www.tensorflow.org/datasets/catalog/super_glue#super_gluerte"}, "state": {"retrieve_from": "https://huggingface.co/datasets/aps/super_glue/viewer/rte/validation", "source_idx": 64, "source_text_omitted": "RTE source text omitted from public export; retrieve the original row from the linked source."}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The hypothesis does not follow from the premise; it may be unsupported or contradicted.", "true": "The hypothesis follows from the premise."}, "experiment": "rte", "id": "rte-validation-0148", "question": "Does the premise entail the hypothesis?", "source": {"configuration": "rte", "dataset": "SuperGLUE RTE", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=entailment, 1=not_entailment; target_yes = 1-label", "license": "Original dataset terms; no unified license verified", "license_status": "unresolved_original_RTE_terms", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 0, "raw_jsonl_sha256": "eec4bac5953538dc63792265f59f7df4f68c0048bcacd3d07336b7d7dc60db40", "row_index_zero_based": 148, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 148, "split": "validation", "task_source": "https://www.tensorflow.org/datasets/catalog/super_glue#super_gluerte"}, "state": {"retrieve_from": "https://huggingface.co/datasets/aps/super_glue/viewer/rte/validation", "source_idx": 148, "source_text_omitted": "RTE source text omitted from public export; retrieve the original row from the linked source."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The hypothesis does not follow from the premise; it may be unsupported or contradicted.", "true": "The hypothesis follows from the premise."}, "experiment": "rte", "id": "rte-validation-0220", "question": "Does the premise entail the hypothesis?", "source": {"configuration": "rte", "dataset": "SuperGLUE RTE", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=entailment, 1=not_entailment; target_yes = 1-label", "license": "Original dataset terms; no unified license verified", "license_status": "unresolved_original_RTE_terms", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 0, "raw_jsonl_sha256": "eec4bac5953538dc63792265f59f7df4f68c0048bcacd3d07336b7d7dc60db40", "row_index_zero_based": 220, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 220, "split": "validation", "task_source": "https://www.tensorflow.org/datasets/catalog/super_glue#super_gluerte"}, "state": {"retrieve_from": "https://huggingface.co/datasets/aps/super_glue/viewer/rte/validation", "source_idx": 220, "source_text_omitted": "RTE source text omitted from public export; retrieve the original row from the linked source."}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The marked occurrences use different senses, even if their topics are related.", "true": "Both marked occurrences use the target word in the same sense."}, "experiment": "wic", "id": "wic-validation-0120", "question": "Does the word marked with [TARGET] and [/TARGET] have the same meaning in both sentences?", "source": {"configuration": "wic", "dataset": "SuperGLUE WIC", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=False, 1=True; target_yes = label", "license": "CC BY-NC 4.0", "license_url": "https://creativecommons.org/licenses/by-nc/4.0/", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 1, "raw_jsonl_sha256": "dccc5a7709e2d016ed6dcab6566039da37dc2a6a1be2fd500f8b04952a5dcaec", "row_index_zero_based": 120, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 120, "split": "validation", "target_offsets": {"end1": 20, "end2": 12, "start1": 13, "start2": 4}, "task_source": "https://pilehvar.github.io/wic/", "transformations": "Inserted TARGET markers around the source character spans; inflected surface forms are retained."}, "state": {"sentence_1": "Did you ever [TARGET]lecture[/TARGET] at Harvard?", "sentence_2": "She [TARGET]lectured[/TARGET] to the class about her travels.", "target_word": "lecture"}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The marked occurrences use different senses, even if their topics are related.", "true": "Both marked occurrences use the target word in the same sense."}, "experiment": "wic", "id": "wic-validation-0223", "question": "Does the word marked with [TARGET] and [/TARGET] have the same meaning in both sentences?", "source": {"configuration": "wic", "dataset": "SuperGLUE WIC", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=False, 1=True; target_yes = label", "license": "CC BY-NC 4.0", "license_url": "https://creativecommons.org/licenses/by-nc/4.0/", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 0, "raw_jsonl_sha256": "dccc5a7709e2d016ed6dcab6566039da37dc2a6a1be2fd500f8b04952a5dcaec", "row_index_zero_based": 223, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 223, "split": "validation", "target_offsets": {"end1": 32, "end2": 25, "start1": 24, "start2": 17}, "task_source": "https://pilehvar.github.io/wic/", "transformations": "Inserted TARGET markers around the source character spans; inflected surface forms are retained."}, "state": {"sentence_1": "He played the trumps in [TARGET]sequence[/TARGET].", "sentence_2": "The doctor saw a [TARGET]sequence[/TARGET] of patients.", "target_word": "sequence"}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The marked occurrences use different senses, even if their topics are related.", "true": "Both marked occurrences use the target word in the same sense."}, "experiment": "wic", "id": "wic-validation-0134", "question": "Does the word marked with [TARGET] and [/TARGET] have the same meaning in both sentences?", "source": {"configuration": "wic", "dataset": "SuperGLUE WIC", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=False, 1=True; target_yes = label", "license": "CC BY-NC 4.0", "license_url": "https://creativecommons.org/licenses/by-nc/4.0/", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 0, "raw_jsonl_sha256": "dccc5a7709e2d016ed6dcab6566039da37dc2a6a1be2fd500f8b04952a5dcaec", "row_index_zero_based": 134, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 134, "split": "validation", "target_offsets": {"end1": 24, "end2": 40, "start1": 19, "start2": 35}, "task_source": "https://pilehvar.github.io/wic/", "transformations": "Inserted TARGET markers around the source character spans; inflected surface forms are retained."}, "state": {"sentence_1": "He needed a lot of [TARGET]power[/TARGET] to hit the ball out of the stadium.", "sentence_2": "The mysterious presence of an evil [TARGET]power[/TARGET].", "target_word": "power"}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The marked occurrences use different senses, even if their topics are related.", "true": "Both marked occurrences use the target word in the same sense."}, "experiment": "wic", "id": "wic-validation-0364", "question": "Does the word marked with [TARGET] and [/TARGET] have the same meaning in both sentences?", "source": {"configuration": "wic", "dataset": "SuperGLUE WIC", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=False, 1=True; target_yes = label", "license": "CC BY-NC 4.0", "license_url": "https://creativecommons.org/licenses/by-nc/4.0/", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 1, "raw_jsonl_sha256": "dccc5a7709e2d016ed6dcab6566039da37dc2a6a1be2fd500f8b04952a5dcaec", "row_index_zero_based": 364, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 364, "split": "validation", "target_offsets": {"end1": 11, "end2": 4, "start1": 7, "start2": 0}, "task_source": "https://pilehvar.github.io/wic/", "transformations": "Inserted TARGET markers around the source character spans; inflected surface forms are retained."}, "state": {"sentence_1": "Please [TARGET]hold[/TARGET] a table at Maxim's.", "sentence_2": "[TARGET]Hold[/TARGET] a table for us at 7:00.", "target_word": "hold"}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The marked occurrences use different senses, even if their topics are related.", "true": "Both marked occurrences use the target word in the same sense."}, "experiment": "wic", "id": "wic-validation-0297", "question": "Does the word marked with [TARGET] and [/TARGET] have the same meaning in both sentences?", "source": {"configuration": "wic", "dataset": "SuperGLUE WIC", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=False, 1=True; target_yes = label", "license": "CC BY-NC 4.0", "license_url": "https://creativecommons.org/licenses/by-nc/4.0/", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 0, "raw_jsonl_sha256": "dccc5a7709e2d016ed6dcab6566039da37dc2a6a1be2fd500f8b04952a5dcaec", "row_index_zero_based": 297, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 297, "split": "validation", "target_offsets": {"end1": 12, "end2": 18, "start1": 6, "start2": 11}, "task_source": "https://pilehvar.github.io/wic/", "transformations": "Inserted TARGET markers around the source character spans; inflected surface forms are retained."}, "state": {"sentence_1": "Don't [TARGET]fiddle[/TARGET] with the screws.", "sentence_2": "She always [TARGET]fiddles[/TARGET] with her van on the weekend.", "target_word": "fiddle"}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The marked occurrences use different senses, even if their topics are related.", "true": "Both marked occurrences use the target word in the same sense."}, "experiment": "wic", "id": "wic-validation-0455", "question": "Does the word marked with [TARGET] and [/TARGET] have the same meaning in both sentences?", "source": {"configuration": "wic", "dataset": "SuperGLUE WIC", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=False, 1=True; target_yes = label", "license": "CC BY-NC 4.0", "license_url": "https://creativecommons.org/licenses/by-nc/4.0/", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 0, "raw_jsonl_sha256": "dccc5a7709e2d016ed6dcab6566039da37dc2a6a1be2fd500f8b04952a5dcaec", "row_index_zero_based": 455, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 455, "split": "validation", "target_offsets": {"end1": 15, "end2": 16, "start1": 11, "start2": 12}, "task_source": "https://pilehvar.github.io/wic/", "transformations": "Inserted TARGET markers around the source character spans; inflected surface forms are retained."}, "state": {"sentence_1": "There's no [TARGET]help[/TARGET] for it.", "sentence_2": "I need some [TARGET]help[/TARGET] with my homework.", "target_word": "help"}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The marked occurrences use different senses, even if their topics are related.", "true": "Both marked occurrences use the target word in the same sense."}, "experiment": "wic", "id": "wic-validation-0566", "question": "Does the word marked with [TARGET] and [/TARGET] have the same meaning in both sentences?", "source": {"configuration": "wic", "dataset": "SuperGLUE WIC", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=False, 1=True; target_yes = label", "license": "CC BY-NC 4.0", "license_url": "https://creativecommons.org/licenses/by-nc/4.0/", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 0, "raw_jsonl_sha256": "dccc5a7709e2d016ed6dcab6566039da37dc2a6a1be2fd500f8b04952a5dcaec", "row_index_zero_based": 566, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 566, "split": "validation", "target_offsets": {"end1": 48, "end2": 40, "start1": 40, "start2": 32}, "task_source": "https://pilehvar.github.io/wic/", "transformations": "Inserted TARGET markers around the source character spans; inflected surface forms are retained."}, "state": {"sentence_1": "The airliner had to land with a nose-up [TARGET]attitude[/TARGET] after the incident.", "sentence_2": "The actor struck just the right [TARGET]attitude[/TARGET].", "target_word": "attitude"}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The marked occurrences use different senses, even if their topics are related.", "true": "Both marked occurrences use the target word in the same sense."}, "experiment": "wic", "id": "wic-validation-0383", "question": "Does the word marked with [TARGET] and [/TARGET] have the same meaning in both sentences?", "source": {"configuration": "wic", "dataset": "SuperGLUE WIC", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=False, 1=True; target_yes = label", "license": "CC BY-NC 4.0", "license_url": "https://creativecommons.org/licenses/by-nc/4.0/", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 1, "raw_jsonl_sha256": "dccc5a7709e2d016ed6dcab6566039da37dc2a6a1be2fd500f8b04952a5dcaec", "row_index_zero_based": 383, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 383, "split": "validation", "target_offsets": {"end1": 20, "end2": 38, "start1": 9, "start2": 27}, "task_source": "https://pilehvar.github.io/wic/", "transformations": "Inserted TARGET markers around the source character spans; inflected surface forms are retained."}, "state": {"sentence_1": "Economic [TARGET]cooperation[/TARGET].", "sentence_2": "They agreed on a policy of [TARGET]cooperation[/TARGET].", "target_word": "cooperation"}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The marked occurrences use different senses, even if their topics are related.", "true": "Both marked occurrences use the target word in the same sense."}, "experiment": "wic", "id": "wic-validation-0031", "question": "Does the word marked with [TARGET] and [/TARGET] have the same meaning in both sentences?", "source": {"configuration": "wic", "dataset": "SuperGLUE WIC", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=False, 1=True; target_yes = label", "license": "CC BY-NC 4.0", "license_url": "https://creativecommons.org/licenses/by-nc/4.0/", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 1, "raw_jsonl_sha256": "dccc5a7709e2d016ed6dcab6566039da37dc2a6a1be2fd500f8b04952a5dcaec", "row_index_zero_based": 31, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 31, "split": "validation", "target_offsets": {"end1": 6, "end2": 25, "start1": 0, "start2": 19}, "task_source": "https://pilehvar.github.io/wic/", "transformations": "Inserted TARGET markers around the source character spans; inflected surface forms are retained."}, "state": {"sentence_1": "[TARGET]Answer[/TARGET] the question.", "sentence_2": "She didn't want to [TARGET]answer[/TARGET].", "target_word": "answer"}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The marked occurrences use different senses, even if their topics are related.", "true": "Both marked occurrences use the target word in the same sense."}, "experiment": "wic", "id": "wic-validation-0214", "question": "Does the word marked with [TARGET] and [/TARGET] have the same meaning in both sentences?", "source": {"configuration": "wic", "dataset": "SuperGLUE WIC", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=False, 1=True; target_yes = label", "license": "CC BY-NC 4.0", "license_url": "https://creativecommons.org/licenses/by-nc/4.0/", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 1, "raw_jsonl_sha256": "dccc5a7709e2d016ed6dcab6566039da37dc2a6a1be2fd500f8b04952a5dcaec", "row_index_zero_based": 214, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 214, "split": "validation", "target_offsets": {"end1": 27, "end2": 37, "start1": 18, "start2": 28}, "task_source": "https://pilehvar.github.io/wic/", "transformations": "Inserted TARGET markers around the source character spans; inflected surface forms are retained."}, "state": {"sentence_1": "The enemy had the [TARGET]advantage[/TARGET] of a more elevated position.", "sentence_2": "The experience gave him the [TARGET]advantage[/TARGET] over me.", "target_word": "advantage"}, "target": 1.0, "target_kind": "label"}
{"criteria": {"false": "The marked occurrences use different senses, even if their topics are related.", "true": "Both marked occurrences use the target word in the same sense."}, "experiment": "wic", "id": "wic-validation-0057", "question": "Does the word marked with [TARGET] and [/TARGET] have the same meaning in both sentences?", "source": {"configuration": "wic", "dataset": "SuperGLUE WIC", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=False, 1=True; target_yes = label", "license": "CC BY-NC 4.0", "license_url": "https://creativecommons.org/licenses/by-nc/4.0/", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 0, "raw_jsonl_sha256": "dccc5a7709e2d016ed6dcab6566039da37dc2a6a1be2fd500f8b04952a5dcaec", "row_index_zero_based": 57, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 57, "split": "validation", "target_offsets": {"end1": 23, "end2": 17, "start1": 11, "start2": 4}, "task_source": "https://pilehvar.github.io/wic/", "transformations": "Inserted TARGET markers around the source character spans; inflected surface forms are retained."}, "state": {"sentence_1": "It was the [TARGET]deliberation[/TARGET] of his act that was insulting.", "sentence_2": "The [TARGET]deliberations[/TARGET] of the jury.", "target_word": "deliberation"}, "target": 0.0, "target_kind": "label"}
{"criteria": {"false": "The marked occurrences use different senses, even if their topics are related.", "true": "Both marked occurrences use the target word in the same sense."}, "experiment": "wic", "id": "wic-validation-0625", "question": "Does the word marked with [TARGET] and [/TARGET] have the same meaning in both sentences?", "source": {"configuration": "wic", "dataset": "SuperGLUE WIC", "hf_repository": "aps/super_glue", "hf_revision_observed": "3de24cf8022e94f4ee4b9d55a6f539891524d646", "kind": "published_human_label", "label_mapping": "HF 0=False, 1=True; target_yes = label", "license": "CC BY-NC 4.0", "license_url": "https://creativecommons.org/licenses/by-nc/4.0/", "mirror": "https://huggingface.co/datasets/aps/super_glue", "pilot_overlap": false, "published_label": 1, "raw_jsonl_sha256": "dccc5a7709e2d016ed6dcab6566039da37dc2a6a1be2fd500f8b04952a5dcaec", "row_index_zero_based": 625, "source_identity_note": "Current accessible HF mirror is aps/super_glue; do not describe it as verified Google-owned.", "source_idx": 625, "split": "validation", "target_offsets": {"end1": 12, "end2": 34, "start1": 3, "start2": 23}, "task_source": "https://pilehvar.github.io/wic/", "transformations": "Inserted TARGET markers around the source character spans; inflected surface forms are retained."}, "state": {"sentence_1": "To [TARGET]embellish[/TARGET] a story, the truth.", "sentence_2": "The old book cover was [TARGET]embellished[/TARGET] with golden letters.", "target_word": "embellish"}, "target": 1.0, "target_kind": "label"}
