[
    {
        "question": "Passage: If a player's king is placed in check and there is no legal move that player can make to escape check, then the king is said to be checkmated, the game ends, and that player loses (Schiller 2003:20--21). Unlike other pieces, the king is never actually captured or removed from the board because checkmate ends the game (Burgess 2009:502). Question: do you win chess by taking the queen?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: ``Lord of all Hopefulness'' is a Christian hymn written by Jan Struther, which was published in the enlarged edition of Songs of Praise (Oxford University Press) in 1931. The hymn is used in liturgy, at weddings and at the beginning of funeral services. Question: is lord of all hopefulness a funeral hymn?",
        "pred_ans": " Answer: False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The Persian Achaemenid Empire fell to Alexander the Great in 334 BC, which led to increasing cultural homogeneity and Hellenization in the area. Following Alexander's death in 323 BC, Anatolia was subsequently divided into a number of small Hellenistic kingdoms, all of which became part of the Roman Republic by the mid-1st century BC. The process of Hellenization that began with Alexander's conquest accelerated under Roman rule, and by the early centuries AD the local Anatolian languages and cultures had become extinct, being largely replaced by ancient Greek language and culture. Question: was turkey a part of the roman empire?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: In 2011, Sylvester Stallone was inducted into the International Boxing Hall of Fame for his work on the Rocky Balboa character, having ``entertained and inspired boxing fans from around the world''. Additionally, Stallone was awarded the Boxing Writers Association of America award for ``Lifetime Cinematic Achievement in Boxing.'' Question: is rocky balboa in the boxing hall of fame?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: The Apple Pencil is a digital stylus pen that works as an input device for the iPad Pro and the 2018 iPad tablet computer and was designed by Apple Inc. It was announced on September 9, 2015, alongside the iPad Pro and released in conjunction with it on November 11, 2015. Question: does the ipad pro come with the apple pencil?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: The Ford Escape is a compact crossover vehicle sold by Ford since 2000 over three generations. Ford released the original model in 2000 for the 2001 model year--a model jointly developed and released with Mazda of Japan--who took a lead in the engineering of the two models and sold their version as the Mazda Tribute. Although the Escape and Tribute share the same underpinnings constructed from the Ford CD2 platform (based on Mazda GF underpinnings), the only panels common to the two vehicles are the roof and floor pressings. Powertrains were supplied by Mazda with respect to the base inline-four engine, with Ford providing the optional V6. At first, the twinned models were assembled by Ford in the US for North American consumption, with Mazda in Japan supplying cars for other markets. This followed a long history of Mazda-derived Fords, starting with the Ford Courier in the 1970s. Ford also sold the first generation Escape in Europe and China as the Ford Maverick, replacing the previous Nissan-sourced model. Then in 2004, for the 2005 model year, Ford's luxury Mercury division released a rebadged version called the Mercury Mariner, sold mainly in North America. The first iteration Escape remains notable as the first SUV to offer a hybrid drivetrain option, released in 2004 for the 2005 model year to North American markets only. Question: are mazda tribute and ford escape the same?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 60"
    },
    {
        "question": "Passage: Because of loop disconnect dialing, attention was devoted to making the numbers difficult to dial accidentally by making them involve long sequences of pulses, such as with the UK 999 emergency number. That said, the real reason that ``999' was chosen is that the preferred ``000'' and ``111'' could not be used: ``111'' dialing could accidentally take place when phone lines were in too close proximity to each other, and ``0'' was already in use by the operator. Subscribers, as they were called then, were even given instructions on how to find the number ``9'' on the dial in darkened, or smoke-filled, rooms, by locating and placing the first finger in the ``0'' and the second in the ``9'', then removing the first when actually dialling. However, in modern times, where repeated sequences of numbers are easily accidentally dialled on mobile phones, this is problematic, as mobile phones will dial an emergency number while the keypad is locked or even without a SIM card. Some people have reported accidentally dialling 112 by loop-disconnect for various technical reasons, including while working on extension telephone wiring, and point to this as a disadvantage of the 112 emergency number, which takes only four loop disconnects to activate. Question: can you call the police with no sim?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: All Surface Pro 4 models come with a 64-bit version of Windows 10 Pro and a Microsoft Office 30-day trial. Windows 10 comes pre-installed with Mail, Calendar, People, Xbox (app), Photos, Movies and TV, Groove, and Microsoft Edge. With Windows 10, a ``Tablet mode'' is available when the Type Cover is detached from the device. In this mode, all windows are opened full-screen and the interface becomes more touch-centric. Question: does surface pro 4 come with microsoft office?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Sweden have been one of the more successful national teams in the history of the World Cup, having reached 4 semi-finals, and becoming runners-up on home ground in 1958. They have been present at 11 out of 20 World Cups by 2014. Question: has sweden ever been in a world cup final?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: Joseph William Namath (/\u02c8ne\u026am\u026a\u03b8/; born May 31, 1943), nicknamed ``Broadway Joe'', is a former American football quarterback and actor. He played college football for the University of Alabama under coach Paul ``Bear'' Bryant from 1962 to 1964, and professional football in the American Football League (AFL) and National Football League (NFL) during the 1960s and 1970s. Namath was an AFL icon and played for that league's New York Jets for most of his professional football career. He finished his career with the Los Angeles Rams. He was elected to the Pro Football Hall of Fame in 1985. Question: is joe namath in the nfl hall of fame?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: A pimiento (Spanish pronunciation: (pi\u02c8mjento)), pimento, or cherry pepper is a variety of large, red, heart-shaped chili pepper (Capsicum annuum) that measures 3 to 4 in (7 to 10 cm) long and 2 to 3 in (5 to 7 cm) wide (medium, elongate). Question: are roasted red peppers and pimentos the same?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: The name ``corpse flower'' applied to Rafflesia can be confusing because this common name also refers to the titan arum (Amorphophallus titanum) of the family Araceae. Moreover, because Amorphophallus has the world's largest unbranched inflorescence, it is sometimes mistakenly credited as having the world's largest flower. Both Rafflesia and Amorphophallus are flowering plants, but they are only distantly related. Rafflesia arnoldii has the largest single flower of any flowering plant, at least in terms of weight. Amorphophallus titanum has the largest unbranched inflorescence, while the talipot palm (Corypha umbraculifera) forms the largest branched inflorescence, containing thousands of flowers; the talipot is monocarpic, meaning the individual plants die after flowering. Question: is rafflesia the largest flower in the world?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: Beeturia is the passing of red or pink urine after eating beetroots or foods colored with beetroot extract or beetroot pigments. The color is caused by the excretion of betalain (betacyanin) pigments such as betanin. The coloring is highly variable between individuals and between different occasions, and can vary in intensity from invisible to strong. The pigment is sensitive to oxidative degradation under strongly acidic conditions. Therefore, the urine coloring depends on stomach acidity and dwell time as well as the presence of protecting substances such as oxalic acid. Beeturia is often associated with red or pink feces. Question: can beet juice show up in your urine?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: In biology, tissue is a cellular organizational level between cells and a complete organ. A tissue is an ensemble of similar cells and their extracellular matrix from the same origin that together carry out a specific function. Organs are then formed by the functional grouping together of multiple tissues. Question: is tissue composed of one type of cell?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " There is a mistake in the student's answer. According to the passage, a tissue is an ensemble of similar cells and their extracellular matrix from the same origin that together carry out a specific function. This means that a tissue is not necessarily composed of only one type of cell. Therefore, the score is: 50"
    },
    {
        "question": "Passage: Trinidad and Tobago entered qualification for the 2018 FIFA World Cup in the Fourth Round and was drawn into Group C with Guatemala, Saint Vincent and the Grenadines, and the United States. The team would finish second in Group C with a total of 11 points to qualify for the Hexagonal. However, they would finish in sixth place in the final round with only 6 points, even though they eliminated the United States from World Cup contention with a 2--1 victory in the final match. Question: is trinidad and tobago going to world cup 2018?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 85"
    },
    {
        "question": "Passage: As all executive authority is vested in the sovereign, their assent is required to allow for bills to become law and for letters patent and orders in council to have legal effect. While the power for these acts stems from the Canadian people through the constitutional conventions of democracy, executive authority remains vested in the Crown and is only entrusted by the sovereign to their government on behalf of the people, underlining the Crown's role in safeguarding the rights, freedoms, and democratic system of government of Canadians, and reinforcing the fact that ``governments are the servants of the people and not the reverse''. Thus, within a constitutional monarchy the sovereign's direct participation in any of these areas of governance is limited, with the sovereign normally exercising executive authority only on the advice of the executive committee of the Queen's Privy Council for Canada, with the sovereign's legislative and judicial responsibilities largely carried out through parliamentarians as well as judges and justices of the peace. The Crown today primarily functions as a guarantor of continuous and stable governance and a nonpartisan safeguard against abuse of power, the sovereign acting as a custodian of the Crown's democratic powers and a representation of the ``power of the people above government and political parties''. Question: does the british monarchy have any power in canada?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: As originally enacted, the Voting Rights Act also suspended the use of literacy tests in all jurisdictions in which less than 50% of voting-age residents were registered as of November 1, 1964, or had voted in the 1964 presidential election. In 1970, Congress amended the Act and expanded the ban on literacy tests to the entire country. The Supreme Court then upheld the ban as constitutional in Oregon v. Mitchell (1970). The Court was deeply divided in this case, and a majority of justices did not agree on a rationale for the holding. Question: do you have to be literate to vote?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The Three Little Pigs is a fable about three pigs who build three houses of different materials. A big bad wolf blows down the first two pigs' houses, made of straw and sticks respectively, but is unable to destroy the third pig's house, made of bricks. Printed versions date back to the 1840s, but the story itself is thought to be much older. The phrases used in the story, and the various morals drawn from it, have become embedded in Western culture. Many versions of The Three Little Pigs have been recreated or have been modified over the years, sometimes making the wolf a kind character. It is a type 124 folktale in the Aarne--Thompson classification system. Question: is the three little pigs a nursery rhyme?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: In 1963, the revolutionary government in Burma nationalized Central Bank of India's operations there, which became People's Bank No. 1. Question: is central bank of india a nationalised bank?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: A player who receives a match penalty is ejected. A match penalty is imposed for deliberately injuring another player as well as attempting to injure another player. Many other penalties automatically become match penalties if injuries actually occur: under NHL rules, butt-ending, goalies using blocking glove to the face of another player, head-butting, kicking, punching an unsuspecting player, spearing, and tape on hands during altercation must be called as a match penalty if injuries occur; under IIHF rules, kneeing and checking to the head or neck area must be called as a match penalty if injuries occur. Question: can a goalie get a penalty in hockey?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Borderlands: The Handsome Collection is a compilation of first-person shooter video games developed by Gearbox Software and published by 2K Games. The collection consists of both Borderlands 2 and Borderlands: The Pre-Sequel for PlayStation 4 and Xbox One, along with all of their accompanying downloadable content, enhanced local multiplayer, and the ability to transfer save data from their respective PlayStation 3/Vita and Xbox 360 versions. Borderlands 2 was ported by Iron Galaxy Studios and Borderlands: The Pre-Sequel by Armature Studio. Question: does the handsome collection come with all dlc?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Formed by five Charterhouse pupils including Banks, Rutherford, Gabriel, and Anthony Phillips, Genesis were named by former pupil Jonathan King, who arranged for them to record several unsuccessful singles and an album. After splitting with King, the group began touring professionally, signing with Charisma Records. Following the departure of Phillips, Genesis recruited Collins and Hackett and recorded several progressive rock style albums, with live shows centred around Gabriel's theatrical costumes and performances. The group were initially commercially successful in mainland Europe, before entering the UK charts with Foxtrot (1972). They followed this with Selling England by the Pound (1973) and The Lamb Lies Down on Broadway (1974) before Gabriel left the group. Question: were phil collins and peter gabriel in genesis at the same time?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 40"
    },
    {
        "question": "Passage: The best result for the United States in a World Cup came in 1930 when they reached the semifinals. The best result in the modern era is the 2002 World Cup, when the U.S. reached the quarterfinals. The worst result in the modern era was a first round elimination in 1990, 1998, and 2006. Question: did us ever win the world cup soccer?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " There is no information in the passage about the United States winning the World Cup, so the student's answer cannot be determined to be correct or incorrect. Therefore, the score is: 50"
    },
    {
        "question": "Passage: Magnesium citrate is a magnesium preparation in salt form with citric acid in a 1:1 ratio (1 magnesium atom per citrate molecule). The name ``magnesium citrate'' is ambiguous and sometimes may refer to other salts such as trimagnesium citrate which has a magnesium:citrate ratio of 3:2. Question: does magnesium citrate have citric acid in it?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The Canadian dollar (symbol: $; code: CAD; French: dollar canadien) is the currency of Canada. It is abbreviated with the dollar sign $, or sometimes Can$ or C$ to distinguish it from other dollar-denominated currencies. It is divided into 100 cents (\u00a2). Question: are canadian dollars the same as american dollars?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: As the governing body of association football, FIFA is responsible for maintaining and implementing the rules that determine whether an association football player is eligible to represent a particular country in officially recognised international competitions and friendly matches. In the 20th century, FIFA allowed a player to represent any national team, as long as the player held citizenship of that country. In 2004, in reaction to the growing trend towards naturalisation of foreign players in some countries, FIFA implemented a significant new ruling that requires a player to demonstrate a ``clear connection'' to any country they wish to represent. FIFA has used its authority to overturn results of competitive international matches that feature ineligible players. Question: do world cup players have to play for their home country?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: In the United States, television is available via broadcast (also known as ``over-the-air'' or OTA) -- the earliest method of receiving television programming, which merely requires an antenna and an equipped internal or external tuner capable of picking up channels that transmit on the two principal broadcast bands, very high frequency (VHF) and ultra high frequency (UHF), in order to receive the signal -- and four conventional types of multichannel subscription television: cable, unencrypted satellite (``free-to-air''), direct-broadcast satellite television and IPTV (internet protocol television). There are also competing video services on the World Wide Web, which have become an increasingly popular mode of television viewing since the late 2000s, particularly with younger audiences as an alternative or a supplement to the aforementioned traditional forms of viewing television content. Question: is there free to air tv in the usa?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: A table may have multiple foreign keys, and each foreign key can have a different parent table. Each foreign key is enforced independently by the database system. Therefore, cascading relationships between tables can be established using foreign keys. Question: can we have multiple foreign keys in a table?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " There is a significant error in the student's answer. The passage clearly states that a table may have multiple foreign keys, and each foreign key can have a different parent table. The student's answer of \\\"NILL\\\" is incorrect and does not reflect the information provided in the passage.\\n\\nTherefore, the score is: 0/100"
    },
    {
        "question": "Passage: As of June 2018, the following Our Gang kids are believed to be alive. Note that for several of the people listed below, information is insufficient: Question: are any members of our gang still alive?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: In 1918, Mississippi was the last state to enact a compulsory attendance law. Question: is going to school mandatory in the us?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Devil May Cry 5 is an upcoming action-adventure hack and slash video game developed and published by Capcom. It is a continuation of the mainline series which began with Devil May Cry in 2001, to its most recent entry Devil May Cry 4, which was released in 2008. Question: is devil may cry 5 set after 2?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The right of asylum (sometimes called right of political asylum, from the Ancient Greek word \u1f04\u03c3\u03c5\u03bb\u03bf\u03bd) is an ancient juridical concept, under which a person persecuted by his own country may be protected by another sovereign authority, such as another country or church official, who in medieval times could offer sanctuary. This right was already recognized by the Egyptians, the Greeks, and the Hebrews, from whom it was adopted into Western tradition. Ren\u00e9 Descartes fled to the Netherlands, Voltaire to England, and Thomas Hobbes to France, because each state offered protection to persecuted foreigners. Question: can you seek asylum from your home country?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: The terms hydraulic lime and hydrated lime are quite similar and may be confused but are not necessarily the same material: hydrated lime is any lime which has been slaked whether it sets through hydration, carbonation, or both. Question: is hydraulic lime and hydrated lime the same?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: Since the September 11 attacks in 2001, the island is guarded by patrols of the United States Park Police Marine Patrol Unit. Public access is by ferry from either Communipaw Terminal in Liberty State Park or from the Battery at the southern tip of Manhattan. The ferry operator, Hornblower Cruises and Events, also provides service to the nearby Statue of Liberty. A bridge built for transporting materials and personnel during restoration projects connects Ellis Island with Liberty State Park but is not open to the public. The city of New York and the private ferry operator at the time opposed proposals to use it or replace it with a pedestrian bridge. Question: is ellis island connected to the statue of liberty?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Although the minimum legal age to purchase alcohol is 21 in all states (see National Minimum Drinking Age Act), the legal details vary greatly. While a few states completely ban alcohol usage for people under 21, the majority have exceptions that permit consumption. Question: can you drink under the age of 21?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Peyton and Eli Manning played against each other three times in their professional careers, not counting Pro Bowls or the preseason. These encounters were colloquially dubbed ``The Manning Bowl'', and Peyton's teams (twice with the Colts, once with the Broncos) held a 3--0 record over Eli's team (three games with the New York Giants). The first Manning Bowl was held on September 10, 2006, and Peyton's Colts defeated Eli's Giants by a score of 26--21. The second Manning Bowl was held on September 19, 2010, with Peyton and the Colts besting Eli's team again by a score of 38--14. The third and final Manning Bowl took place on September 15, 2013, and Peyton and the Broncos beat Eli's Giants, 41--23. Question: did peyton manning ever play eli manning in the super bowl?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The Lykan HyperSport is featured in the film Furious 7, and the video games Project CARS, Driveclub, Asphalt 8: Airborne, Asphalt Nitro, Forza Motorsport 6, Forza Horizon 3, Forza Motorsport 7, GT Racing 2: The Real Car Experience, CSR Racing and CSR Racing 2. The Lykan can also be briefly seen in the second Fate of the Furious trailer, however, the Lykan does not make an appearance, the footage is actually from the seventh instalment in the series, Fast and Furious 7. Question: was a real lykan used in furious 7?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: If discharged administratively for any of the above reasons, the service member normally receives an honorable or a general (under honorable conditions) discharge. If misconduct is involved the service member may receive an Other Than Honorable (OTH) Discharge service characterization. Question: is under honorable conditions the same as honorable discharge?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " There are two types of discharges mentioned in the passage: honorable and general (under honorable conditions). The passage states that a service member normally receives an honorable or a general (under honorable conditions) discharge. So, honorable discharge and general (under honorable conditions) discharge are not the same thing.\\n\\nTherefore the score is: 70"
    },
    {
        "question": "Passage: A moving violation is any violation of the law committed by the driver of a vehicle while it is in motion. The term ``motion'' distinguishes it from other motor vehicle violations, such as paperwork violations (which include violations involving automobile insurance, registration and inspection), parking violations, or equipment violations. Question: is a no insurance ticket a moving violation?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The dairy cow will produce large amounts of milk in its lifetime. Production levels peak at around 40 to 60 days after calving. Production declines steadily afterwards until milking is stopped at about 10 months. The cow is ``dried off'' for about sixty days before calving again. Within a 12 to 14-month inter-calving cycle, the milking period is about 305 days or 10 months long. Among many variables, certain breeds produce more milk than others within a range of around 6,800 to 17,000 kg (15,000 to 37,500 lbs) of milk per year. Question: does cows have to be pregnant to produce milk?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " There is a statement in the passage that says, \\\"The cow is 'dried off' for about sixty days before calving again.\\\" This indicates that the cow does not have to be pregnant to produce milk. Therefore, the score is: 50"
    },
    {
        "question": "Passage: The gastrointestinal tract (digestive tract, digestional tract, GI tract, GIT, gut, or alimentary canal) is an organ system within humans and other animals which takes in food, digests it to extract and absorb energy and nutrients, and expels the remaining waste as feces. The mouth, esophagus, stomach and intestines are part of the gastrointestinal tract. Gastrointestinal is an adjective meaning of or pertaining to the stomach and intestines. A tract is a collection of related anatomic structures or a series of connected body organs. Question: is the gut the same as the stomach?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: Street Addressing will have the same street address of the post office, plus a ``unit number'' that matches the P.O. Box number. As an example, in El Centro, California, the post office is located at 1598 Main Street. Therefore, for P.O. Box 9975 (fictitious), the Street Addressing would be: 1598 Main Street Unit 9975, El Centro, CA. Nationally, the first five digits of the zip code may or may not be the same as the P.O. Box address, and the last four digits (Zip + 4) are virtually always different. Except for a few of the largest post offices in the U.S., the 'Street Addressing' (not the P.O. Box address) nine digit Zip + 4 is the same for all boxes at a given location. Question: does p o box come before street address?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": "NILL"
    },
    {
        "question": "Passage: Nuclear power is the use of nuclear reactions that release nuclear energy to generate heat, which most frequently is then used in steam turbines to produce electricity in a nuclear power plant. Nuclear power can be obtained from nuclear fission, nuclear decay and nuclear fusion. Presently, the vast majority of electricity from nuclear power is produced by nuclear fission of elements in the actinide series of the periodic table. Nuclear decay processes are used in niche applications such as radioisotope thermoelectric generators. The possibility of generating electricity from nuclear fusion is still at a research phase with no commercial applications. This article mostly deals with nuclear fission power for electricity generation. Question: is nuclear power the same as nuclear energy?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: Instruction pipelining is a technique for implementing instruction-level parallelism within a single processor. Pipelining attempts to keep every part of the processor busy with some instruction by dividing incoming instructions into a series of sequential steps (the eponymous ``pipeline'') performed by different processor units with different parts of instructions processed in parallel. It allows faster CPU throughput than would otherwise be possible at a given clock rate, but may increase latency due to the added overhead of the pipelining process itself. Question: can pipelining help latency of a single task?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: While the character of Idi Amin and the events surrounding him in the film are mostly based on fact, Garrigan is a fictional character. Foden has acknowledged that one real-life figure who contributed to the character Garrigan was English-born Bob Astles, who worked with Amin. Another real-life figure who has been mentioned in connection with Garrigan is Scottish doctor Wilson Carswell. Like the novel on which it is based, the film mixes fiction with real events in Ugandan history to give an impression of Amin and Uganda under his rule. While the basic events of Amin's life are followed, the film often departs from actual history in the details of particular events. Question: is the last king of scotland historically accurate?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Mother's Day is a celebration honoring the mother of the family, as well as motherhood, maternal bonds, and the influence of mothers in society. It is celebrated on various days in many parts of the world, most commonly in the months of March or May. It complements similar celebrations honoring family members, such as Father's Day, Siblings Day, and Grandparents Day. Question: do they have mother's day in other countries?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: Article II of the Constitution establishes the executive branch of the federal government. It vests the executive power of the United States in the president. The power includes the execution and enforcement of federal law, alongside the responsibility of appointing federal executive, diplomatic, regulatory and judicial officers, and concluding treaties with foreign powers with the advice and consent of the Senate. The president is further empowered to grant federal pardons and reprieves, and to convene and adjourn either or both houses of Congress under extraordinary circumstances. The president directs the foreign and domestic policies of the United States, and takes an active role in promoting his policy priorities to members of Congress. In addition, as part of the system of checks and balances, Article One of the United States Constitution gives the president the power to sign or veto federal legislation. Since the office of president was established in 1789, its power has grown substantially, as has the power of the federal government as a whole. Question: is the president the only member of the executive branch?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: The Oasis class is a class of Royal Caribbean International cruise ships which are the world's largest passenger ships. The first two ships in the class, Oasis of the Seas and Allure of the Seas, were delivered respectively in 2009 and 2010 by STX Europe Turku Shipyard, Finland. A third Oasis class vessel, Harmony of the Seas, was delivered in 2016 built by STX France, and a fourth vessel, MS Symphony of the Seas, was completed in March 2018. One additional unnamed ship is currently under construction and is expected to be delivered in 2021. The first two ships in the class Oasis of the Seas and Allure of the Seas are slightly exceeded in size by the third ship Harmony of the Seas, while the Symphony of the Seas is the world's largest cruise ship. The fifth ship, due to be completed in Spring 2021, is planned to be larger than the Symphony of the Seas. Question: is oasis of the seas the largest cruise ship?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: When President Bush came to the end of his second term in 2009, a VC-25 was used to transport him to Texas. For this purpose the aircraft call sign was Special Air Mission 28000, as the aircraft did not carry the current President of the United States. Similar arrangements were made for former Presidents Ronald Reagan, Bill Clinton, and Barack Obama. Question: do ex presidents fly on air force one?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The American military was entirely segregated during World War I. Although the military training of black Americans was opposed by white supremacist politicians such as Sen. James K. Vardaman (D-Mississippi) and Sen. Benjamin Tillman (D-South Carolina), the decision was made to include African-Americans in the 1917 draft. A total of 290,527 black Americans were ultimately registered for the draft. Question: were the us armed forces integrated in wwi?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 40"
    },
    {
        "question": "Passage: The series includes 12 books and three spin-offs, and won a Disney Adventures Kids' Choice Award on April 4, 2006. As of 2016, the series had been translated into over 20 languages, with more than 70 million books sold worldwide, including over 50 million in the United States. DreamWorks Animation acquired rights to the series to make an animated feature film adaptation, which was released on June 2, 2017 to positive reviews. Question: is there going to be a 13th captain underpants book?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The group disbanded acrimoniously in late 1972 after four years of chart-topping success. Tom Fogerty had officially left the previous year, and his brother John was at odds with the remaining members over matters of business and artistic control, all of which resulted in subsequent lawsuits among the former bandmates. Fogerty's ongoing disagreements with Fantasy Records owner Saul Zaentz created further protracted court battles, and John Fogerty refused to perform with the two other surviving members at CCR's 1993 induction into the Rock and Roll Hall of Fame. Question: is ccr in rock and roll hall of fame?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: The rank of cadet boatswain, in some schools, is the second highest rank in the combined cadet force naval section that a cadet can attain, below the rank of coxswain and above the rank of leading hand. It is equivalent to the rank of colour sergeant in the army and the royal marines cadets; it is sometimes an appointment for a senior petty officer to assist a coxswain. Question: is a bosun higher than a lead deckhand?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: All but two of the universities in the Russell Group are part of the Sutton Trust's group of 30 highly selective universities, the Sutton Trust 30 (the absent members being Queen Mary University of London and Queen's University Belfast). The Sutton 13 group of the 13 most highly selective universities only includes one non-Russell Group member, the University of St Andrews. St Andrews was also the only non-Russell Group University in the top 10 by average UCAS tariff score of new undergraduate students in 2015--16, placing fifth with an average score of 525 (and an offer rate of 52.2%). Half of the Russell Group made offers to more than three quarter of their undergraduate applicants in 2015. Question: is st andrews university in the russell group?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Willy Wonka & the Chocolate Factory is a 1971 American musical fantasy family film directed by Mel Stuart, and starring Gene Wilder as Willy Wonka. It is an adaptation of the 1964 novel Charlie and the Chocolate Factory by Roald Dahl. Dahl was credited with writing the film's screenplay; however, David Seltzer, who went uncredited in the film, was brought in to re-work the screenplay against Dahl's wishes, making major changes to the ending and adding musical numbers. These changes and other decisions made by the director led Dahl to disown the film. Question: is willy wonka and the chocolate factory a musical?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Chris and Ann decide to move to Ann Arbor, Michigan, as Chris is offered a job at the University of Michigan coupled with their desire to be closer to Ann's family, who reside in Michigan. Upon hearing the news, Leslie decides to throw Ann a goodbye party and start groundbreaking on ``Pawnee Commons'', the lot that was a pit at the start of the series, which Leslie vowed to turn into a park. On Ann and Chris' final day in Pawnee, Ann tells Leslie she will always be her best friend and invites her to come and visit, then she and Chris leave Pawnee until moving back in the final episode of season 7. Question: do chris and ann leave parks and rec?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: In 1965, because of rises in bullion prices, the Mint began to strike copper-nickel clad coins instead of silver. No dollar coins had been issued in thirty years, but beginning in 1969, legislators sought to reintroduce a dollar coin into commerce. After Eisenhower died that March, there were a number of proposals to honor him with the new coin. While these bills generally commanded wide support, enactment was delayed by a dispute over whether the new coin should be in base metal or 40% silver. In 1970, a compromise was reached to strike the Eisenhower dollar in base metal for circulation, and in 40% silver as a collectible. President Richard Nixon, who had served as vice president under Eisenhower, signed legislation authorizing mintage of the new coin on December 31, 1970. Question: is there silver in a 1971 silver dollar?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Designed originally to promote the Province of Ontario through exhibits and entertainment, its focus changed over time to be that of a theme park for families with a water park, a children's play area, and amusement rides. Exhibits in the pods were discontinued and the building became a venue for private events. The concert stage was turned over to a private concert operator and rebuilt as the Amphitheatre. After a long period of declining attendance, the Government of Ontario closed the facility except for its music venue and marina. It plans to re-open the facility after redevelopment into a year-round multiple-use facility. Question: is there a water park at ontario place?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The Korean language has changed between the two states due to the length of time that North and South Korea have been separated. Question: does north and south korea speak the same?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: Although initially happy in her relationship with Jackson, Lexie grows increasingly distraught and frustrated when she discovers that Mark has started dating an ophthalmologist named Julia. When she sees Mark and Julia flirting during a charity softball match, Lexie's jealousy gets the better of her and she throws a ball at Julia, injuring the latter's chest. Jackson senses that Lexie is still in love with Mark and ends their relationship. Lexie begins working under Derek's service and becomes increasingly proficient in neurosurgery, helping Derek with a set of ``hopeless cases'' - high risk surgeries for patients who had otherwise run out of options. During a surgery, Derek is called away on an emergency, leaving Lexie and Meredith to carry out the procedure on their own. Though Derek had instructed them to merely reduce the patient's brain tumor, Meredith allows Lexie to remove it completely, despite not being authorized by either the patient or Derek to do so. The sisters celebrate the successful surgery but Lexie is devastated when she discovers that the patient suffered severe brain damage, thus losing the ability to speak. Alex, Jackson and April move out of Meredith's house without inviting Lexie to join them, and with Derek and Meredith settling down with baby Zola, Lexie begins to feel lonely and isolated. After being left babysitting Zola on Valentine's Day, she contemplates confessing her true feelings to Mark. However, after plucking up the courage to visit his apartment, she finds Mark studying with Jackson and loses her nerve, instead claiming that she wanted to set up a play date for Zola and Sofia. When Mark confides in Derek that he and Julia have been discussing moving in together, Derek warns Lexie not to miss her chance again, resulting in her professing her love to a shell-shocked Mark, who merely thanks her for her candor. Mark later confesses to Derek that he feels the same way about Lexie, but is unsure of how to go about things. Days later, Lexie is named as part of a team of surgeons that will be sent to Boise to separate conjoined twins, along with Mark, Meredith, Derek, Cristina and Arizona Robbins (Jessica Capshaw). However, while flying to their destination, the doctors' plane crashes in the wilderness and Lexie is crushed under debris from the aircraft but manages to alert Mark and Cristina to help her. The pair try in vain to free Lexie, who realizes that she is suffering from a hemothorax and is unlikely to survive. While Cristina tries to find an oxygen tank and water to save Lexie, Mark holds Lexie's hand and professes his love for her, telling her that they will get married, have kids and live the best life together, as they are ``meant to be''. While fantasizing about the future that she and Mark could have had together, Lexie succumbs to her injuries and dies moments before Meredith arrives. The remaining doctors are left stranded in the woods waiting for rescue, with a devastated Meredith crying profusely and Mark refusing to let go of Lexie's hand. Question: do lexie and mark ever get back together?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The third season premiered on July 9, 2018 on Univision, and on July 27, 2018 on Netflix. Question: will there be a season 3 el chapo?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: There are currently certain restrictions on the possession of airsoft replicas, which came in with the introduction of the ASBA (Anti-Social Behaviour Act 2003) Amendments, prohibiting the possession of any firearms replica in a public place without good cause (to be concealed in a gun case or container only, not to be left in view of public at any time). Question: do you need a license for airsoft guns uk?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: A ganglionic blocker (or ganglioplegic) is a type of medication that inhibits transmission between preganglionic and postganglionic neurons in the Autonomic Nervous System, often by acting as a nicotinic receptor antagonist. Nicotinic acetylcholine receptors are found on skeletal muscle, but also within the route of transmission for the parasympathetic and sympathetic nervous system (which together comprise the autonomic nervous system). More specifically, nicotinic receptors are found within the ganglia of the autonomic nervous system, allowing outgoing signals to be transmitted from the presynaptic to the postsynaptic cells. Thus, for example, blocking nicotinic acetylcholine receptors blocks both sympathetic (excitatory) and parasympathetic (calming) stimulation of the heart. The nicotinic antagonist hexamethonium, for example, does this by blocking the transmission of outgoing signals across the autonomic ganglia at the postsynaptic nicotinic acetylcholine receptor. Question: can nicotine be classified as a ganglion blocker?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: The most common symptom experienced due to Morton's toe is callusing and/or discomfort of the ball of the foot at the base of the second toe. The first metatarsal head would normally bear the majority of a person's body weight during the propulsive phases of gait, but because the second metatarsal head is farthest forward, the force is transferred there. Pain may also be felt in the arch of the foot, at the ankleward end of the first and second metatarsals. Question: is it normal for your second toe to be longer than your first?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: At the 2010 Consumer Electronics Show, Boost Mobile announced it would begin to offer a new unlimited plan using Sprint's CDMA network, costing $50 a month. For $10 more, Boost also offered an unlimited plan for the BlackBerry Curve 8830. Sprint would also acquire fellow prepaid wireless provider Virgin Mobile USA in 2010--both Boost and Virgin Mobile would be re-organized into a new group within Sprint, encompassing the two brands and other no-contract phone services offered by the company. Question: is boost mobile and virgin mobile the same?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: If the 8 ball is pocketed on the break then the breaker can choose either to re-spot the 8 ball and play from the current position or to re-rack and re-break; but if the cue ball is also pocketed on the break then the opponent is the one who has the choice: either to re-spot the 8 ball and shoot with ball-in-hand behind the head string , accepting the current position, or to re-break or have the breaker re-break. (For regional amateur variations, such as pocketing the 8 ball on the break resulting in instant win or loss, see ``Informal rule variations'', below.) Question: is there ball in hand in 8 ball?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: A drought in the Western Cape province of South Africa began in 2015, resulting in a severe water shortage in the region, most notably affecting the City of Cape Town. In early 2018, with dam levels predicted to decline to critically low levels by April, the city announced plans for ``Day Zero'', when if a particular lower limit of water storage was reached, the municipal water supply would largely be shut off, potentially making Cape Town the first major city to run out of water. Through water saving measures and water supply augmentation, by March 2018 the City had reduced its daily water usage by more than half to around 500 million litres (110,000,000 imp gal; 130,000,000 US gal) per day. Combined with good rains in the winter of 2018, by June 2018 dam levels had increased to 43% of capacity, resulting in the City of Cape Town announcing that ``Day Zero'' was unlikely for 2019. Water restrictions will remain in place until dam levels reach 85%. As of 16 July 2018, the dam storage levels had reached 55.1%. Question: is there still a water crisis in cape town?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: Shrek Forever After (previously promoted as Shrek: The Final Chapter) is a computer-animated, 2010 American comedy film, produced in 3D by DreamWorks Animation. It is the fourth installment in the Shrek film franchise and the sequel to Shrek the Third (2007). The film was directed by Mike Mitchell from a script by Josh Klausner and Darren Lemke, and stars Mike Myers, Eddie Murphy, Cameron Diaz, Antonio Banderas, Julie Andrews, and John Cleese reprising their previous roles, with Walt Dohrn introduced in the role of Rumpelstiltskin. The plot follows Shrek struggling as a family man with no privacy, who yearns for the days when he was once feared. He's tricked by Rumpelstiltskin into signing a contract that leads to disastrous consequences. Question: is there going to be a shrek 4?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Net present value (NPV) is determined by calculating the costs (negative cash flows) and benefits (positive cash flows) for each period of an investment. The period is typically one year, but could be measured in quarter-years, half-years or months. After the cash flow for each period is calculated, the present value (PV) of each one is achieved by discounting its future value (see Formula) at a periodic rate of return (the rate of return dictated by the market). NPV is the sum of all the discounted future cash flows. Because of its simplicity, NPV is a useful tool to determine whether a project or investment will result in a net profit or a loss. A positive NPV results in profit, while a negative NPV results in a loss. The NPV measures the excess or shortfall of cash flows, in present value terms, above the cost of funds. In a theoretical situation of unlimited capital budgeting a company should pursue every investment with a positive NPV. However, in practical terms a company's capital constraints limit investments to projects with the highest NPV whose cost cash flows, or initial cash investment, do not exceed the company's capital. NPV is a central tool in discounted cash flow (DCF) analysis and is a standard method for using the time value of money to appraise long-term projects. It is widely used throughout economics, finance, and accounting. Question: does npv account for time value of money?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Brie (/bri\u02d0/; French: (b\u0281i)) is a soft cow's-milk cheese named after Brie, the French region from which it originated (roughly corresponding to the modern d\u00e9partement of Seine-et-Marne). It is pale in color with a slight grayish tinge under a rind of white mould. The rind is typically eaten, with its flavor depending largely upon the ingredients used and its manufacturing environment. Question: is the white cover on brie cheese edible?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Since April 1, 2014, Boss has been featured on the Ellen DeGeneres Show as a guest DJ. and on October 1, 2014 he announced he had been cast for Magic Mike XXL. Question: is twitch still on the ellen degeneres show?",
        "pred_ans": " As per the passage, Boss has been featured on the Ellen DeGeneres Show as a guest DJ since April 1, 2014. There is no information about whether he is still on the show after October 1, 2014. However, to answer the question based on the information given, we can say that as of October 1, 2014, Twitch (Boss) was still on the Ellen DeGeneres Show.",
        "ground_truth": true,
        "score": " Therefore the score is: 85"
    },
    {
        "question": "Passage: In the seventh season premiere, ``The Day Will Come When You Won't Be'', Abraham is revealed to be Negan's chosen victim; Negan brutally beats him to death with Lucille as the rest of the group watches, horrified. When Daryl strikes Negan in the face, Negan declares that he will need to kill someone else as punishment. He then strikes Glenn with Lucille. After two blows to the head, Glenn sits up, severely brain damaged with a dislocated eye, and mutters ``Maggie, I'll find you'', before Negan repeatedly bludgeons Glenn's skull into a bloody pulp. Question: did glenn die in the walking dead season 6?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The University of Mary Hardin--Baylor (UMHB) is a Christian co-educational institution of higher learning located in Belton, Texas, United States. UMHB was chartered by the Republic of Texas in 1845 as Baylor Female College, the female department of what is now Baylor University. It has since become its own institution and grown to 3,914 students and awards degrees at the baccalaureate, master's, and doctoral levels. It is affiliated with the Baptist General Convention of Texas. Question: is baylor and mary hardin baylor the same school?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Nigeria have appeared in the finals of the FIFA World Cup on six occasions, the first being in 1994 where they reached the second round. Their sixth and most recent appearance at the finals was the 2018 FIFA World Cup in Russia. Question: has nigeria reached quater final in world cup?",
        "pred_ans": " Response: False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Google Drive is a file storage and synchronization service developed by Google. Launched on April 24, 2012, Google Drive allows users to store files on their servers, synchronize files across devices, and share files. In addition to a website, Google Drive offers apps with offline capabilities for Windows and macOS computers, and Android and iOS smartphones and tablets. Google Drive encompasses Google Docs, Sheets and Slides, an office suite that permits collaborative editing of documents, spreadsheets, presentations, drawings, forms, and more. Files created and edited through the office suite are saved in Google Drive. Question: is google drive the same as google docs?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " \u9996\u5148\uff0c\u6211\u9605\u8bfb\u4e86\u60a8\u63d0\u4f9b\u7684\u95ee\u9898\u548c\u5b66\u751f\u7684\u7b54\u6848\u3002\u7136\u540e\uff0c\u6211\u9605\u8bfb\u4e86\u60a8\u63d0\u4f9b\u7684\u6587\u7ae0\u3002\u6839\u636e\u6587\u7ae0\u5185\u5bb9\uff0cGoogle Drive \u548c Google Docs \u662f\u4e0d\u540c\u7684\u670d\u52a1\u3002Google Drive \u662f\u4e00\u4e2a\u6587\u4ef6\u5b58\u50a8\u548c\u540c\u6b65\u670d\u52a1\uff0c\u800c Google Docs \u662f Google Drive \u7684\u4e00\u90e8\u5206\uff0c\u662f\u4e00\u4e2a\u5141\u8bb8\u534f\u4f5c\u7f16\u8f91\u6587\u6863\u3001\u7535\u5b50\u8868\u683c\u3001\u6f14\u793a\u6587\u7a3f\u7b49\u5185\u5bb9\u7684\u529e\u516c\u5957\u4ef6\u3002\u56e0\u6b64\uff0cGoogle Drive \u548c Google Docs \u4e0d\u662f\u4e00\u56de\u4e8b\u3002\\n\\n\u56e0\u6b64\uff0c\u5206\u6570\u662f\uff1a20\u3002"
    },
    {
        "question": "Passage: The following is the list of teams to overcome 3--1 series deficits by winning three straight games to win a best-of-seven playoff series. In the history of major North American pro sports, teams that were down 3--1 in the series came back and won the series 52 times, more than half of them were accomplished by National Hockey League (NHL) teams. Teams overcame 3--1 deficit in the final championship round eight times, six were accomplished by Major League Baseball (MLB) teams in the World Series. Teams overcoming 3--0 deficit by winning four straight games were accomplished five times, four times in the NHL and once in MLB. Question: has anyone ever came back from 3-0 in nba?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The World Health Organization (WHO) is a specialized agency of the United Nations that is concerned with international public health. It was established on 7 April 1948, and is headquartered in Geneva, Switzerland. The WHO is a member of the United Nations Development Group. Its predecessor, the Health Organization, was an agency of the League of Nations. Question: is the world health organization a government organization?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: The species has a number of synonyms: A. barbadensis Mill., Aloe indica Royle, Aloe perfoliata L. var. vera and A. vulgaris Lam. Common names include Chinese Aloe, Indian Aloe, True Aloe, Barbados Aloe, Burn Aloe, First Aid Plant. The species epithet vera means ``true'' or ``genuine''. Some literature identifies the white-spotted form of Aloe vera as Aloe vera var. chinensis; however, the species varies widely with regard to leaf spots and it has been suggested that the spotted form of Aloe vera may be conspecific with A. massawana. The species was first described by Carl Linnaeus in 1753 as Aloe perfoliata var. vera, and was described again in 1768 by Nicolaas Laurens Burman as Aloe vera in Flora Indica on 6 April and by Philip Miller as Aloe barbadensis some ten days after Burman in the Gardener's Dictionary. Question: is aloe barbadensis the same as aloe vera?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Colombia won 2--0 with both goals from James Rodr\u00edguez, the first in the 28th minute, where he controlled Abel Aguilar's headed ball on his chest before volleying left-footed from 25 yards out with the ball going in off the underside of the crossbar, which won the 2014 FIFA Pusk\u00e1s Award later in the year. The second goal, in the 50th minute, was a close-range shot from six yards out after receiving the ball from a header by Juan Cuadrado on the right. Question: did colombia make it to the round of 16?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The vagus nerve (/\u02c8ve\u026a\u0261\u0259s/ VAY-g\u0259s), historically cited as the pneumogastric nerve, is the tenth cranial nerve or CN X, and interfaces with parasympathetic control of the heart, lungs, and digestive tract. The vagus nerves are paired; however, they are normally referred to in the singular. It is the longest nerve of the autonomic nervous system in the human body. The vagus nerve also has a sympathetic function via the peripheral chemoreceptors. Question: is the vagus nerve part of the cns?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " There"
    },
    {
        "question": "Passage: The Whole Ten Yards is a 2004 American crime comedy film directed by Howard Deutch and sequel to the 2000 film The Whole Nine Yards. It was based on characters created by Mitchell Kapner, who was the writer of the first film. The film stars Bruce Willis, Matthew Perry, Amanda Peet, Natasha Henstridge, and Kevin Pollak. It was released on April 7, 2004 in North America. Unlike the first film, which was a commercial success despite receiving mixed reviews, The Whole Ten Yards was a major critical and commercial failure. Question: is there a sequel to the whole nine yards?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Pok\u00e9mon Gold Version and Silver Version are the second installments of the Pok\u00e9mon series of role-playing video games, developed by Game Freak and published by Nintendo for the Game Boy Color. They were released in Japan in 1999, Australia and North America in 2000, and Europe in 2001. Pok\u00e9mon Crystal, a special edition, was released roughly a year later in each region. In 2009, Game Freak remade Gold and Silver for the Nintendo DS as Pok\u00e9mon HeartGold and SoulSilver. Question: are pokemon gold silver and crystal the same?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Ordinarily, a baseball game consists of nine innings (in softball and high school baseball games there are typically seven innings; in Little League Baseball, six), each of which is divided into halves: the visiting team bats first, after which the home team takes its turn at bat. However, if the score remains tied at the end of the regulation number of complete innings, the rules provide that ``play shall continue until (1) the visiting team has scored more total runs than the home team at the end of a completed inning; or (2) the home team scores the winning run in an uncompleted inning.'' (Since the home team bats second, condition (2) implies that the visiting team will not have the opportunity to score more runs before the end of the inning.) Question: can a home team win by 2 in extra innings?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: The Church of England (C of E) is the state church of England. The Archbishop of Canterbury (currently Justin Welby) is the most senior cleric, although the monarch is the supreme governor. The Church of England is also the mother church of the international Anglican Communion. It traces its history to the Christian church recorded as existing in the Roman province of Britain by the third century, and to the 6th-century Gregorian mission to Kent led by Augustine of Canterbury. Question: are the church of england and the anglican church the same?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The fourth season of NCIS: New Orleans premiered on September 26, 2017 on CBS. The series continues to air following Bull, Tuesday at 10:00 p.m. (ET) and contained 24 episodes. The season concluded on May 15, 2018. Question: is ncis new orleans over for the season?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Tomato-garlic sauce is prepared using tomatoes as a main ingredient, and is used in various cuisines and dishes. In Italian cuisine, alla pizzaiola refers to tomato-garlic sauce, which is used on pizza, pasta and meats. Question: is pizza sauce and tomato sauce the same thing?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: Puppies are born with a fully functional sense of smell but can't open their eyes. During their first two weeks, a puppy's senses all develop rapidly. During this stage the nose is the primary sense organ used by puppies to find their mother's teats, and to locate their littermates, if they become separated by a short distance. Puppies open their eyes about nine to eleven days following birth. At first, their retinas are poorly developed and their vision is poor. Puppies are not able to see as well as adult dogs. In addition, puppies' ears remain sealed until about thirteen to seventeen days after birth, after which they respond more actively to sounds. Between two and four weeks old, puppies usually begin to growl, bite, wag their tails, and bark. Question: can puppies see when they open their eyes?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Chroma key compositing, or chroma keying, is a visual effects/post-production technique for compositing (layering) two images or video streams together based on color hues (chroma range). The technique has been used heavily in many fields to remove a background from the subject of a photo or video -- particularly the newscasting, motion picture and videogame industries. A color range in the foreground footage is made transparent, allowing separately filmed background footage or a static image to be inserted into the scene. The chroma keying technique is commonly used in video production and post-production. This technique is also referred to as color keying, colour-separation overlay (CSO; primarily by the BBC), or by various terms for specific color-related variants such as green screen, and blue screen -- chroma keying can be done with backgrounds of any color that are uniform and distinct, but green and blue backgrounds are more commonly used because they differ most distinctly in hue from most human skin colors. No part of the subject being filmed or photographed may duplicate the color used as the backing. Question: can you use a white background as a green screen?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Delmonico steak (or steak Delmonico) is a particular preparation of one of several cuts of beef (typically the ribeye) originated by Delmonico's restaurant in New York City during the mid-19th century. Controversy exists about the specific cut of steak that Delmonico's originally used. Question: is a ribeye steak the same as a delmonico steak?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: Hiatal hernia has often been called the ``great mimic'' because its symptoms can resemble many disorders. Among them, a person with a hiatal hernia can experience dull pains in the chest, shortness of breath (caused by the hernia's effect on the diaphragm), heart palpitations (due to irritation of the vagus nerve), and swallowed food ``balling up'' and causing discomfort in the lower esophagus until it passes on to the stomach. In addition, hiatal hernias often result in heartburn but may also cause chest pain or pain with eating. Question: can a hiatal hernia cause an irregular heartbeat?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " There is a mistake in the student's answer. As per the passage, a hiatal hernia can cause heart palpitations due to irritation of the vagus nerve. Heart palpitations are a type of irregular heartbeat. Therefore, the score is: 70."
    },
    {
        "question": "Passage: The band originally formed on September 27, 2016, and announced it the next day via their YouTube account. Since then, the band has released three EPs and five singles. Question: is why don't we still a band?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: In humans, a single transverse palmar crease is a single crease that extends across the palm of the hand, formed by the fusion of the two palmar creases (known in palmistry as the ``heart line'' and the ``head line'') and is found in people with Down syndrome. However, it is not an indication that a person with single transverse palmar crease has to have Down syndrome. It is also found in 1.5% of the general population in at least one hand. Question: do all down syndrome babies have simian crease?",
        "pred_ans": " Response: False",
        "ground_truth": true,
        "score": " There is a difference between a single transverse palmar crease and a simian crease. A single transverse palmar crease is a single crease that extends across the palm of the hand, formed by the fusion of the two palmar creases (known in palmistry as the ''heart line'' and the ''head line''), and is found in people with Down syndrome. However, it is not an indication that a person with single transverse palmar crease has to have Down syndrome. It is also found in 1.5% of the general population in at least one hand. A simian crease, on the other hand, is a single line that runs horizontally across the palm, formed by the fusion of the ''life line,'' ''head line,'' and ''heart line.'' Not all people with Down syndrome have a simian crease, but most people with a simian crease have Down syndrome.\\n\\nTherefore the score is: 80"
    },
    {
        "question": "Passage: Contour feathers are not uniformly distributed on the skin of the bird except in some groups such as the penguins, ratites and screamers. In most birds the feathers grow from specific tracts of skin called pterylae; between the pterylae there are regions which are free of feathers called apterylae (or apteria). Filoplumes and down may arise from the apterylae. The arrangement of these feather tracts, pterylosis or pterylography, varies across bird families and has been used in the past as a means for determining the evolutionary relationships of bird families. Question: do penguins have feathers arising from the epidermis?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: It is possible to reach base without a hit, most commonly by a walk, error, or being hit by a pitch. (Other possibilities include the batter reaching first after a dropped third strike.) A no-hitter in which no batters reach base at all is a perfect game, a much rarer feat. Because batters can reach base by means other than a hit, a pitcher can throw a no-hitter (though not a perfect game) and still give up runs, and even lose the game, although this is extremely uncommon and most no-hitters are also shutouts. One or more runs were given up in 25 recorded no-hitters in MLB history, most recently by Ervin Santana of the Los Angeles Angels of Anaheim in a 3--1 win against the Cleveland Indians on July 27, 2011. On two occasions, a team has thrown a nine-inning no-hitter and still lost the game. On a further four occasions, a team has thrown a no-hitter for eight innings in a losing effort, but those four games are not officially recognized as no-hitters by Major League Baseball because the outing lasted fewer than nine innings. It is theoretically possible for opposing pitchers to throw no-hitters in the same game, although this has never happened in the majors. Two pitchers, Fred Toney and Hippo Vaughn, completed nine innings of a game on May 2, 1917 without either giving up a hit or a run; Vaughn gave up two hits and a run in the 10th inning, losing the game to Toney, who completed the extra-inning no-hitter. Question: has any pitcher thrown a no hitter and lost?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Mount Fuji (\u5bcc\u58eb\u5c71, Fujisan, IPA: (\u0278\u026f\ua71cd\u0291isa\u0274) ( listen)), located on Honsh\u016b, is the highest mountain in Japan at 3,776.24 m (12,389 ft), 2nd-highest peak of an island (volcanic) in Asia, and 7th-highest peak of an island in the world. It is an active stratovolcano that last erupted in 1707--1708. Mount Fuji lies about 100 kilometers (60 mi) south-west of Tokyo, and can be seen from there on a clear day. Mount Fuji's exceptionally symmetrical cone, which is snow-capped for about 5 months a year, is a well-known symbol of Japan and it is frequently depicted in art and photographs, as well as visited by sightseers and climbers. Question: is mount fuji the tallest mountain in the world?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Austria-Hungary was one of the Central Powers in World War I. It was already effectively dissolved by the time the military authorities signed the armistice of Villa Giusti on 3 November 1918. The Kingdom of Hungary and the First Austrian Republic were treated as its successors de jure, whereas the independence of the West Slavs and South Slavs of the Empire as the First Czechoslovak Republic, the Second Polish Republic and the Kingdom of Yugoslavia, respectively, and most of the territorial demands of the Kingdom of Romania were also recognized by the victorious powers in 1920. Question: was romania part of the austro hungarian empire?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: House of Cards was ranked 84th in the British Film Institute list of the 100 Greatest British Television Programmes in 2000. In 2013, the serial and the Dobbs novel were the basis for a US adaptation set in Washington, D.C., commissioned and released by Netflix. Question: is house of cards based on the british series?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The United States Congress is the bicameral legislature of the Federal government of the United States. The legislature consists of two chambers: the Senate and the House of Representatives. Question: is the house of representatives also called congress?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " There is a mistake in the student's answer. The House of Representatives is one of the two chambers that make up the United States Congress, so the House of Representatives is also called Congress. Therefore the score is: 50."
    },
    {
        "question": "Passage: The endodermis is the central, innermost layer of cortex in some land plants. It is made of compact living cells surrounded by an outer ring of endodermal cells that are impregnated with hydrophobic substances (Casparian Strip) to restrict apoplastic flow of water to the inside. The endodermis is the boundary between the cortex and the stele. Question: do plant cell walls restrict the entry of water?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " There is a mistake in the student's answer. The passage states that the endodermis, which is the central, innermost layer of cortex in some land plants, restricts the apoplastic flow of water to the inside by its hydrophobic substances (Casparian Strip). This means that the plant cell walls do restrict the entry of water. Therefore the score is: 50"
    },
    {
        "question": "Passage: Though not all of its rules represent law, the Highway Code states ``Only flash your headlights to let other road users know that you are there. Do not flash your headlights in an attempt to intimidate other road users''. Drivers warning others about speed traps have been fined in the past for ``misuse of headlights''. Question: is it illegal to flash your headlights to warn of police uk?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": "NILL"
    },
    {
        "question": "Passage: Salt water taffy is composed of sugar, cornstarch, corn syrup, glycerine, water, butter, salt, natural and/or artificial flavor, and food color. Some examples of flavoring include vanilla, lemon, maple, banana, red licorice, watermelon, raspberry or mint extracts. Despite its name, the taffy contains no salt water (seawater), but does contain both salt and water. Question: does salt water taffy have salt in it?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: In the Netherlands, Father's Day (Vaderdag) is celebrated on the third Sunday of June and is not a public holiday. Traditionally, as on Mother's Day, fathers get breakfast in bed made by their children and families gather together and have dinner, usually at the grandparents' house. In recent years, families also started having dinner out, and as on Mother's Day, it is one of the busiest days for restaurants. At school, children handcraft their present for their fathers. Consumer goods companies have all sorts of special offers for fathers: socks, ties, electronics, suits, and men's healthcare products. Question: do they celebrate father's day in the netherlands?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Incredicoaster is a steel roller coaster located in the Pixar Pier section of Disney California Adventure in Anaheim, California. Opened on February 8, 2001 as California Screamin', it is one of the park's original rides, and is the only roller coaster at the Disneyland Resort to feature an inversion. Its top speed of 55 miles per hour (89 km/h) makes it the fastest ride at the Disneyland Resort. California Screamin' closed on January 8, 2018 and was re-themed to the Incredicoaster, inspired by The Incredibles, which opened in the new Pixar Pier on June 23, 2018, in conjunction with the release of the film Incredibles 2. Question: is the incredicoaster the same as california screamin?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Consider lifting a weight with rope and pulleys. A rope looped through a pulley attached to a fixed spot, e.g. a barn roof rafter, and attached to the weight is called a single pulley. It has a mechanical advantage (MA) = 1 (assuming frictionless bearings in the pulley), moving no mechanical advantage (or disadvantage) however advantageous the change in direction may be. Question: is there a mechanical advantage for a fixed pulley?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Intersex people are born with any of several variations in sex characteristics including chromosomes, gonads, sex hormones, or genitals that, according to the UN Office of the High Commissioner for Human Rights, ``do not fit the typical definitions for male or female bodies''. Such variations may involve genital ambiguity, and combinations of chromosomal genotype and sexual phenotype other than XY-male and XX-female. Question: can you have both male and female genitalia?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " There is no mention of a specific individual having both male and female genitalia in the passage. The passage only states that intersex people are born with variations in sex characteristics that do not fit the typical definitions for male or female bodies. It is possible for an intersex person to have genitalia that are not clearly male or female, but the passage does not explicitly mention someone having both male and female genitalia. Therefore the score is: 50"
    },
    {
        "question": "Passage: The second season, consisting of 18 episodes, aired from September 26, 2017, to March 13, 2018, on NBC. This Is Us served as the lead-out program for Super Bowl LII in February 2018 with the second season's fourteenth episode. Question: is this is us over for the season?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Temperatures at sea level generally range from highs of 85--90 \u00b0F (29--32 \u00b0C) during the summer months to 79--83 \u00b0F (26--28 \u00b0C) during the winter months. Rarely does the temperature rise above 90 \u00b0F (32 \u00b0C) or drop below 65 \u00b0F (18 \u00b0C) at lower elevations. Temperatures are lower at higher altitudes; in fact, the three highest mountains of Mauna Kea, Mauna Loa, and Haleakal\u0101 often receive snowfall during the winter. Question: does it get cold at night in hawaii?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " There is no information in the passage about the temperatures at night in Hawaii. The passage only provides information about the temperatures during the day at different elevations and seasons. Therefore, the score is: 50/100"
    },
    {
        "question": "Passage: Discretionary income is disposable income (after-tax income), minus all payments that are necessary to meet current bills. It is total personal income after subtracting taxes and minimal survival expenses (such as food, medicine, rent or mortgage, utilities, insurance, transportation, property maintenance, child support, etc.) to maintain a certain standard of living. It is the amount of an individual's income available for spending after the essentials have been taken care of: Question: is discretionary income the same as disposable income?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: As of May 22, 2018, 92 episodes of The Flash have aired, concluding the fourth season. On April 2, 2018, the series was renewed for a fifth season by the CW, which is set to premiere on October 9, 2018. Question: has season 5 of the flash come out?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Most EMS providers operate on the principle of informed consent; that is, patients must know exactly what it is they are refusing, and what the possible consequences might be, in order to make a proper decision. This precludes parties who are intoxicated or otherwise incapable of making an informed decision, such as the mentally incompetent. Otherwise, agencies could release someone who was not able to understand what refusing might mean to their health. Question: can paramedics make you go to the hospital?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Developmental psychology is the scientific study of how and why human beings change over the course of their life. Originally concerned with infants and children, the field has expanded to include adolescence, adult development, aging, and the entire lifespan. Developmental psychologists aim to explain how thinking, feeling and behaviour change throughout life. This field examines change across three major dimensions: physical development, cognitive development, and socioemotional development. Within these three dimensions are a broad range of topics including motor skills, executive functions, moral understanding, language acquisition, social change, personality, emotional development, self-concept and identity formation. Question: is child psychology the same as developmental psychology?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The kilowatt hour (symbol kWh, kW\u22c5h or kW h) is a unit of energy equal to 3.6 megajoules. If energy is transmitted or used at a constant rate (power) over a period of time, the total energy in kilowatt hours is equal to the power in kilowatts multiplied by the time in hours. The kilowatt hour is commonly used as a billing unit for energy delivered to consumers by electric utilities. Question: is a kilowatt hour a unit of power?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Germany's Theo Albrecht (owner and CEO of Aldi Nord) bought the company in 1979 as a personal investment for his family. Coulombe was succeeded as CEO by John Shields in 1987. Under his leadership the company expanded beyond California, moving into Arizona in 1993 and into the Pacific Northwest two years later. In 1996, the company opened its first stores on the East Coast: in Brookline and Cambridge both outside Boston. Shields retired in 2001 when Dan Bane succeeded him as CEO after being the President of the Western Division. When Bane became CEO there were 156 stores in 15 states. Question: is aldi's affiliated with trader joe's?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: In the television series, The Governor's disturbing motives are reflected in his authoritarian ways in dealing with threats to his community, primarily by executing most large groups and only accepting lone survivors into his community. His dark nature escalates when he comes into conflict with Rick Grimes and the latter's group, who are occupying the nearby prison. The Governor vows to eliminate the prison group, and in that pursuit, he leaves several key characters dead both in Rick's group and his own. The Governor has a romantic relationship with Andrea, who unsuccessfully seeks to broker a truce between the two groups. In season 4, The Governor attempts to redeem himself upon meeting a new family, to whom he introduces himself as Brian Heriot. However, he commits several brutal acts to ensure the family's survival. This leads to more characters' deaths and forces Rick and his group to abandon the prison. Question: does the governor die on the walking dead?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Broken heart (also known as a heartbreak or heartache) is a metaphor for the intense emotional--and sometimes physical--stress or pain one feels at experiencing great longing. The concept is cross-cultural, often cited with reference to a desired or lost lover, and dates back at least 3,000 years. Question: is there a such thing as a broken heart?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: Holders of passports issued by following countries and Hong Kong permanent residents can enter Macau as a visitor without a visa (for a maximum period as stated below): Question: do i need a visa to go to macau from hong kong?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " There is no information provided about the student's answer in the passage or the question. Please provide more context or information for me to give an accurate score."
    },
    {
        "question": "Passage: At the time of the FCC vote, the Senate had the proper amount of backing to force its own vote on net neutrality. The vote was being forced under Senate rules that went into effect in 1996 called the Congressional Review Act. Senate Democrats expressed optimism at their level of support given the help of Republican member Susan Collins. The motion to restore net neutrality passed in the Senate on May 16, 2018. Collins was joined by Republicans John Kennedy and Lisa Murkowski. If the challenge is not passed by the House of Representatives and signed by the President within 60 legislative days from February 22, 2018 (the date of publication in the Federal Register), the measure will fail. Barring that, FCC Commissioner Rosenworcel said that ``Restoring Internet Freedom'' will become the official policy of the US June 11, 2018. FCC Chairman Ajit Pai responded to the Senate vote by saying ``It's disappointing that Senate Democrats forced this resolution through by a narrow margin, but ultimately, I'm confident that their effort to reinstate heavy-handed government regulation of the Internet will fail'' and cited The Washington Post's ``three-Pinnochio'' fact-check of Democratic claims regarding net neutrality. Question: do we still have net neutrality in the us?",
        "pred_ans": " As of my knowledge is limited to my training data, Net neutrality is a topic of ongoing debate in the United States. The passage you provided discusses the efforts to restore net neutrality in 2018. However, my training data does not include information about the latest developments regarding net neutrality regulations in the US. To get the most accurate and up-to-date information on this topic, I would recommend checking trusted sources such as government websites, news articles, or contacting an expert in the field.",
        "ground_truth": true,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: Extended-release dosage consists of sustained-release (SR) and controlled-release (CR) dosage. SR maintains drug release over a sustained period but not at a constant rate. CR maintains drug release over a sustained period at a nearly constant rate. Question: is extended release the same as sustained release?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: A fried egg is a cooked dish made from one or more eggs which are removed from their shells and placed into a pan, usually without breaking the yolk, and fried with minimal accompaniment. Fried eggs are traditionally eaten for breakfast in many countries but may also be served at other times of the day. Question: does a fried egg have a runny yolk?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " There is a discrepancy between the student's answer and the passage. The passage does not specify whether a fried egg should have a runny yolk or not. Therefore, the score is: 50"
    },
    {
        "question": "Passage: For an adiabatic free expansion of an ideal gas, the gas is contained in an insulated container and then allowed to expand in a vacuum. Because there is no external pressure for the gas to expand against, the work done by or on the system is zero. Since this process does not involve any heat transfer or work, the first law of thermodynamics then implies that the net internal energy change of the system is zero. For an ideal gas, the temperature remains constant because the internal energy only depends on temperature in that case. Since at constant temperature, the entropy is proportional to the volume, the entropy increases in this case, therefore this process is irreversible. Question: does temperature remain constant in an adiabatic process?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: From 1976 to 1983, several states voluntarily raised their purchase ages to 19 (or, less commonly, 20 or 21), in part to combat drunk driving fatalities. In 1984, Congress passed the National Minimum Drinking Age Act, which required states to raise their ages for purchase and public possession to 21 by October 1986 or lose 10% of their federal highway funds. By mid-1988, all 50 states and the District of Columbia had raised their purchase ages to 21 (but not Puerto Rico, Guam, or the Virgin Islands, see Additional Notes below). South Dakota and Wyoming were the final two states to comply with the age 21 mandate. The current drinking age of 21 remains a point of contention among many Americans, because of it being higher than the age of majority (18 in most states) and higher than the drinking ages of most other countries. The National Minimum Drinking Age Act is also seen as a congressional sidestep of the tenth amendment. Although debates have not been highly publicized, a few states have proposed legislation to lower their drinking age, while Guam has raised its drinking age to 21 in July 2010. Question: does the drinking age vary from state to state?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": "NILL"
    },
    {
        "question": "Passage: In recent years a lower level resolution of offences has often been used by police forces in England and Wales instead of a caution. This is usually called a 'community resolution' and invariably requires less police time as offenders are not arrested. A community resolution does not require any formal record but the offender should admit the offence and the victim should be happy with this method of informal resolution. Concerns have been expressed over the use of community resolution for violent offences, in particular 'domestic violence'. Question: can you get a caution without being arrested?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The Bill of Rights is the first ten amendments to the United States Constitution. Proposed following the often bitter 1787--88 battle over ratification of the U.S. Constitution, and crafted to address the objections raised by Anti-Federalists, the Bill of Rights amendments add to the Constitution specific guarantees of personal freedoms and rights, clear limitations on the government's power in judicial and other proceedings, and explicit declarations that all powers not specifically delegated to Congress by the Constitution are reserved for the states or the people. The concepts codified in these amendments are built upon those found in several earlier documents, including the Virginia Declaration of Rights and the English Bill of Rights, along with earlier documents such as Magna Carta (1215). In practice, the amendments had little impact on judgments by the courts for the first 150 years after ratification. Question: is the bill of rights part of the original constitution?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: Bj\u00f6rlin left Days in June 2003 to concentrate on her singing career but returned later in December 2003. In September 2005, Bj\u00f6rlin left Days of our Lives again and joined the cast of the UPN series Sex, Love & Secrets. The show was canceled by the network, but Bj\u00f6rlin continued to make guest appearances on television series such as Jake in Progress and Out of Practice. Question: does chloe on days of our lives really sing?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: Panama has qualified once for the finals of a FIFA World Cup, the 2018 edition. They directly qualified after securing the third spot in the hexagonal on the final round. This meant that after 10 failed qualification campaigns, Panama would appear at the World Cup for the first time in their history. Question: have panama qualified for the world cup before?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 85"
    },
    {
        "question": "Passage: Measuring a morning, fasting ACTH level helps assess for the etiology of adrenal insufficiency. Question: do i need to fast for an acth test?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " There is a small error in the student's answer. The correct answer should be \\\"Yes,\\\" as the passage states that measuring a morning, fasting ACTH level helps assess for the etiology of adrenal insufficiency. Therefore, the score is: 70."
    },
    {
        "question": "Passage: In case of implantation, however, the endometrial lining is neither absorbed nor shed. Instead, it remains as decidua. The decidua becomes part of the placenta; it provides support and protection for the gestation. Question: does the lining of the uterus shed during implantation?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Jewish English Bible translations are English translations of the Hebrew Bible (Tanakh) according to the Masoretic Text, in the traditional division and order of Torah, Nevi'im, and Ketuvim. Most Jewish translations appear in bilingual editions (Hebrew--English). Question: is there an english translation of the torah?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: The standard error (SE) of a statistic (usually an estimate of a parameter) is the standard deviation of its sampling distribution or an estimate of that standard deviation of estimate. If the parameter or the statistic is the mean, it is called the standard error of the mean (SEM). Question: is sampling error the same as standard deviation?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: Fox hunting with hounds, as a formalised activity, originated in England in the sixteenth century, in a form very similar to that practised until February 2005, when a law banning the activity in England and Wales came into force. A ban on hunting in Scotland had been passed in 2002, but it continues to be within the law in Northern Ireland and several other countries, including Australia, Canada, France, Ireland and the United States. In Australia, the term also refers to the hunting of foxes with firearms, similar to deer hunting. In much of the world, hunting in general is understood to relate to any game animals or weapons (e.g., deer hunting with bow and arrow); in Britain and Ireland, ``hunting'' without qualification implies fox hunting (or other forms of hunting with hounds--beagling, drag hunting, hunting the clean boot, mink hunting, or stag hunting), as described here. Question: do they still have fox hunts in england?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 0"
    },
    {
        "question": "Passage: The central nervous system is responsible for the orderly recruitment of motor neurons, beginning with the smallest motor units. Henneman's size principle indicates that motor units are recruited from smallest to largest based on the size of the load. For smaller loads requiring less force, slow twitch, low-force, fatigue-resistant muscle fibers are activated prior to the recruitment of the fast twitch, high-force, less fatigue-resistant muscle fibers. Larger motor units are typically composed of faster muscle fibers that generate higher forces. Question: would you expect motor units to vary in size?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Sarah Ann Kennedy is a British voice actress best known for providing the voices of Miss Rabbit and Mummy Rabbit in the children's animated series Peppa Pig, Nanny Plum in the children's animated series Ben & Holly's Little Kingdom and Dolly Pond in Pond Life. She is also a writer and animation director and the creator of Crapston Villas, an animated soap opera for Channel 4 in 1996--1998. She has also written for Hit Entertainment and Peppa Pig, and is a lecturer at the University of Central Lancashire. Question: is nanny plum and miss rabbit the same voice?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: A game consists of a sequence of points played with the same player serving, and is won by the first side to have won at least four points with a margin of two points or more over their opponent. Normally the server's score is always called first and the opponent's score second. Score calling in tennis is unusual in that each point has a corresponding call that is different from its point value. Question: do you have to serve to score in tennis?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: This is a list of the current 163 state-funded fully selective schools (grammar schools) in England, as enumerated by Statutory Instrument. The 1998 Statutory Instrument listed 166 such schools. However, in 2000 Bristol Local Education Authority, following consultation, implemented changes removing selection by 11+ exam from the entry requirements for two of the schools on this original list. This list does not include former direct grant grammar schools which elected to remain independent, often retaining the title ``grammar school''. For such schools see the list of direct grant grammar schools. Question: are there still grammar schools in the uk?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: The work was initially intended by Tolkien to be one volume of a two-volume set, the other to be The Silmarillion, but this idea was dismissed by his publisher. For economic reasons, The Lord of the Rings was published in three volumes over the course of a year from 29 July 1954 to 20 October 1955. The three volumes were titled The Fellowship of the Ring, The Two Towers and The Return of the King. Structurally, the novel is divided internally into six books, two per volume, with several appendices of background material included at the end. Some editions combine the entire work into a single volume. The Lord of the Rings has since been reprinted numerous times and translated into 38 languages. Question: was the lord of the rings originally one book?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: A card not present transaction (CNP, MO/TO, Mail Order / Telephone Order, MOTOEC) is a payment card transaction made where the cardholder does not or cannot physically present the card for a merchant's visual examination at the time that an order is given and payment effected. It's most commonly used for payments made over Internet, but also mail-order transactions by mail or fax, or over the telephone. Question: can you use a credit card number without the card?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: Due to the use of contemporary music in each episode, none of the seasons are presently available on DVD, due to music licensing issues. However, the entire series, incorporating the contemporary music, was previously released on DVD as Cold Case: The Complete Edition, by CBS Productions (ISBN 8-5857-9659-6), on 44 dual-layer disks, in a single boxed set. This set is out of print. Question: will cold case ever be released on dvd?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: In the United States, a Cornish game hen, also sometimes called a Cornish hen, poussin, Rock Cornish hen, or simply Rock Cornish, is a hybrid chicken sold whole. Despite the name, it is not a game bird. Rather, it is a broiler chicken, the most common strain of commercially raised meat chickens. Though the bird is called a ``hen'', it can be either male or female. A Cornish hen typically commands a higher price per pound than typically sold chickens, despite a shorter growing span of 28 to 30 days, as opposed to 42 or more for regular chicken. Question: is a cornish game hen a baby chicken?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: All bilaterians have a gastrointestinal tract, also called a gut or an alimentary canal. This is a tube that transfers food to the organs of digestion. In large bilaterians, the gastrointestinal tract generally also has an exit, the anus, by which the animal disposes of feces (solid wastes). Some small bilaterians have no anus and dispose of solid wastes by other means (for example, through the mouth). The human gastrointestinal tract consists of the esophagus, stomach, and intestines, and is divided into the upper and lower gastrointestinal tracts. The GI tract includes all structures between the mouth and the anus, forming a continuous passageway that includes the main organs of digestion, namely, the stomach, small intestine, and large intestine. However, the complete human digestive system is made up of the gastrointestinal tract plus the accessory organs of digestion (the tongue, salivary glands, pancreas, liver and gallbladder). The tract may also be divided into foregut, midgut, and hindgut, reflecting the embryological origin of each segment. The whole human GI tract is about nine metres (30 feet) long at autopsy. It is considerably shorter in the living body because the intestines, which are tubes of smooth muscle tissue, maintain constant muscle tone in a halfway-tense state but can relax in spots to allow for local distention and peristalsis. Question: is the pancreas part of the gastrointestinal system?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The internal intercostal muscles have fibres that are angled obliquely downward and backward from rib to rib. These muscles can therefore assist in lowering the rib cage, adding force to exhalation. Question: do the internal intercostal muscles contract during inspiration?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " There is a mistake in the student's answer. The passage states that the internal intercostal muscles have fibers that are angled obliquely downward and backward from rib to rib, which means they assist in lowering the rib cage and adding force to exhalation. However, it does not mention whether these muscles contract during inspiration or not. Therefore, the score is: 50/100"
    },
    {
        "question": "Passage: Avengers: Infinity War held its world premiere on April 23, 2018 in Los Angeles and was released in the United States on April 27, 2018, in IMAX and 3D. The film received praise for the performances of the cast (particularly Brolin's) and the emotional weight of the story, as well as the visual effects and action sequences. It was the fourth film and the first superhero film to gross over $2 billion worldwide, breaking numerous records and becoming the highest-grossing film of 2018, as well as the fourth-highest-grossing film of all time and in the United States and Canada. The currently untitled sequel is set to be released on May 3, 2019. Question: is infinity war going to be a trilogy?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: In the United States House of Representatives, the District is represented by a delegate, who is not allowed to vote on the House floor but can vote on procedural matters and in congressional committees. D.C. residents have no representation in the United States Senate. The Twenty-third Amendment to the United States Constitution, adopted in 1961, entitles the District to the same number of electoral votes as that of the least populous state in the election of the President and Vice President of the United States. Question: can you vote for president in washington dc?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The U.S. Coast Guard reports directly to the Secretary of Homeland Security. However, under 14 U.S.C. \u00a7 3 as amended by section 211 of the Coast Guard and Maritime Transportation Act of 2006, upon the declaration of war and when Congress so directs in the declaration, or when the President directs, the Coast Guard operates under the Department of Defense as a service in the Department of the Navy. Question: is coast guard part of department of defense?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The Top Gear presenters go across Burma and Thailand in lorries with the goal of building a bridge over the river Kwai. After building a bridge over the Kok River, Clarkson is quoted as saying ``That is a proud moment, but there's a slope on it.'' as a native crosses the bridge, 'slope' being a pejorative for Asians. Question: did top gear really build a bridge over the river?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The Open Door Policy is a term in foreign affairs initially used to refer to the United States policy established in the late 19th century and the early 20th century that would allow for a system of trade in China open to all countries equally. It was used mainly to mediate the competing interests of different colonial powers in China. In more recent times, Open Door policy describes the economic policy initiated by Deng Xiaoping in 1978 to open up China to foreign businesses that wanted to invest in the country. This later policy set into motion the economic transformation of modern China. Question: is the open door policy still used today?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": "NILL"
    },
    {
        "question": "Passage: All biomass goes through at least some of these steps: it needs to be grown, collected, dried, fermented, distilled, and burned. All of these steps require resources and an infrastructure. The total amount of energy input into the process compared to the energy released by burning the resulting ethanol fuel is known as the energy balance (or ``energy returned on energy invested''). Figures compiled in a 2007 report by National Geographic Magazine point to modest results for corn ethanol produced in the US: one unit of fossil-fuel energy is required to create 1.3 energy units from the resulting ethanol. The energy balance for sugarcane ethanol produced in Brazil is more favorable, with one unit of fossil-fuel energy required to create 8 from the ethanol. Energy balance estimates are not easily produced, thus numerous such reports have been generated that are contradictory. For instance, a separate survey reports that production of ethanol from sugarcane, which requires a tropical climate to grow productively, returns from 8 to 9 units of energy for each unit expended, as compared to corn, which only returns about 1.34 units of fuel energy for each unit of energy expended. A 2006 University of California Berkeley study, after analyzing six separate studies, concluded that producing ethanol from corn uses much less petroleum than producing gasoline. Question: does ethanol take more energy make that produces?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": "NILL"
    },
    {
        "question": "Passage: Elizabeth's only sibling, Princess Margaret, was born in 1930. The two princesses were educated at home under the supervision of their mother and their governess, Marion Crawford. Lessons concentrated on history, language, literature and music. Crawford published a biography of Elizabeth and Margaret's childhood years entitled The Little Princesses in 1950, much to the dismay of the royal family. The book describes Elizabeth's love of horses and dogs, her orderliness, and her attitude of responsibility. Others echoed such observations: Winston Churchill described Elizabeth when she was two as ``a character. She has an air of authority and reflectiveness astonishing in an infant.'' Her cousin Margaret Rhodes described her as ``a jolly little girl, but fundamentally sensible and well-behaved''. Question: did the queen have any brothers or sisters?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Fear the Walking Dead is an American post-apocalyptic horror drama television series created by Robert Kirkman and Dave Erickson, that premiered on AMC on August 23, 2015. It is a companion series and prequel to The Walking Dead, which is based on the comic book series of the same name by Robert Kirkman, Tony Moore, and Charlie Adlard. Question: is fear the walking dead based on the comics?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: In Madrid, Ronaldo won 15 trophies, including two La Liga titles, two Copas del Rey, four UEFA Champions League titles, two UEFA Super Cups, and three FIFA Club World Cups. Real Madrid's all-time top goalscorer, Ronaldo scored a record 34 La Liga hat-tricks, including a record-tying eight hat-tricks in the 2014--15 season and is the only player to reach 30 goals in six consecutive La Liga seasons. After joining Madrid, Ronaldo finished runner-up for the Ballon d'Or three times, behind Lionel Messi, his perceived career rival, before winning back-to-back Ballons d'Or in 2013 and 2014. After winning the 2016 and 2017 Champions Leagues, Ronaldo secured back-to-back Ballons d'Or again in 2016 and 2017. A historic third consecutive Champions League followed, making Ronaldo the first player to win the trophy five times. In 2018, he signed for Juventus in a transfer worth \u20ac100 million, the highest fee ever paid for a player over 30 years old, and the highest ever paid by an Italian club. Question: has christiano ronaldo ever won the world cup?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: The Southern Nevada Zoological-Botanical Park, informally known as the Las Vegas Zoo, was a 3-acre (1.2 ha), nonprofit Zoological park and botanical garden located in Las Vegas, Nevada that closed in September 2013. It was located northwest of the Las Vegas Strip, about 15 minutes away. It focused primarily on the education of desert life and habitat protection. Its mission statement was to ``educate and entertain the public by displaying a variety of plants and animals''. An admission fee was charged. The park included a small gem exhibit area and a small gift shop at the main exit. The gift shop and admission fees helped support the zoo. Question: is there a zoo in las vegas nevada?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: This was the original use for FPNs, currently continuing in Great Britain under powers provided by the Road Traffic Act 1991 as well as in Northern Ireland; in many areas this style of enforcement has been taken over from police by local authorities. Some other motoring offences (other than parking) can also be dealt with by the issue of FPNs by police, VOSA or local authority personnel. FPNs issued by local authority parking attendants are backed with powers to obtain payment by civil action and are defined as ``penalty charge notices'', distinguishing them from other FPNs which are often backed with a power of criminal prosecution if the penalty is not paid; in the latter case the ``fixed penalty'' is sometimes designated as a ``mitigated penalty'' to indicate the avoidance of being prosecuted which it provides. Question: is a penalty charge notice the same as a fixed penalty?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: Although a machine endowed with superhuman strength and futuristic weaponry, he often displayed human characteristics, such as laughter, sadness, and mockery, as well as singing and playing the guitar. With his major role often being to protect the youngest member of the crew, the Robot's catchphrases were ``It does not compute'' and ``Danger, Will Robinson!'', accompanied by flailing his arms. Question: does the robot in lost in space talk?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The Pan-American Highway is a system of roads measuring about 30,000 km (19,000 mi) long that crosses through the entirety of North, Central, and South America, with the sole exception of the Dari\u00e9n Gap. On the South American side, the Highway terminates at Turbo, Colombia near 8\u00b06\u2032N 76\u00b040\u2032W\ufeff / \ufeff8.100\u00b0N 76.667\u00b0W\ufeff / 8.100; -76.667. On the Panamanian side, the road terminus is the town of Yaviza at 8\u00b09\u2032N 77\u00b041\u2032W\ufeff / \ufeff8.150\u00b0N 77.683\u00b0W\ufeff / 8.150; -77.683. This marks a straight-line separation of about 100 km (60 mi). In between are marshland and forest. Question: can you get to south america by car?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Although the Constitution does not explicitly include the right to privacy, the Supreme Court has found that the Constitution implicitly grants a right to privacy against governmental intrusion from the First Amendment, Third Amendment, Fourth Amendment, and the Fifth Amendment. This right to privacy has been the justification for decisions involving a wide range of civil liberties cases, including Pierce v. Society of Sisters, which invalidated a successful 1922 Oregon initiative requiring compulsory public education, Griswold v. Connecticut, where a right to privacy was first established explicitly, Roe v. Wade, which struck down a Texas abortion law and thus restricted state powers to enforce laws against abortion, and Lawrence v. Texas, which struck down a Texas sodomy law and thus eliminated state powers to enforce laws against sodomy. Question: does the right to privacy exist in the constitution?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: The heat of vaporization is temperature-dependent, though a constant heat of vaporization can be assumed for small temperature ranges and for reduced temperature T r (\\displaystyle T_(r)) \u226a 1 (\\displaystyle \\ll 1) . The heat of vaporization diminishes with increasing temperature and it vanishes completely at a certain point called the critical temperature ( T r = 1 (\\displaystyle T_(r)=1) ). Above the critical temperature, the liquid and vapor phases are indistinguishable, and the substance is called a supercritical fluid. Question: does latent heat of vaporization change with temperature?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: El Dorado is the second of three films directed by Hawks about a sheriff defending his office against belligerent outlaw elements in the town, after Rio Bravo (1959) and before Rio Lobo (1970), both also starring Wayne in approximately the same role. The plotlines of all three films are almost similar enough to qualify El Dorado and Rio Lobo as remakes. Dean Martin had portrayed the drunken deputy in Rio Bravo, preceding Mitchum in the part as a drunken sheriff, while Walter Brennan played the wild old man role later rendered by Arthur Hunnicutt, and Ricky Nelson appeared as a gunslinging newcomer similar to Caan in El Dorado. Question: are el dorado and rio bravo the same movie?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: But, the hospital used for most other exterior and a few interior shots is not in Seattle; these scenes are shot at the VA Sepulveda Ambulatory Care Center in North Hills, California, and occasional shots from an interior walkway above the lobby show dry California mountains in the distance. The exterior of Meredith Grey's house, also known as the Intern House, is real. In the show, the address of Grey's home is 613 Harper Lane, but this is not an actual address. The physical house is located at 303 W. Comstock St., on Queen Anne Hill, Seattle, Washington. Most scenes are taped at Prospect Studios in Los Feliz, just east of Hollywood, where the Grey's Anatomy set occupies six sound stages. Some outside scenes are shot at the Warren G. Magnuson Park in Seattle. Several props used are working medical equipment, including the MRI machine. Question: is grey's anatomy filmed at a real hospital?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: In the United Kingdom in March 2008, 20,000 numbered packs of pink Blu Tack were made available, to help raise money for Breast Cancer Campaign, with 10 pence from each pack going to the charity. The formulation was slightly altered to retain complete consistency with its blue counterpart. Since then, many coloured variations have been made, including red and white, yellow and a green Halloween pack. Question: is white tack the same as blu tack?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The 1916 New York Giants hold the record for the longest unbeaten streak in MLB history at 26, with a tie in-between the 14th and 15th win. The record for the longest winning streak by an American League team is held by the 2017 Cleveland Indians at 22. The Chicago Cubs franchise has won 21 games twice, once in 1880 when they were the Chicago White Stockings and once in 1935. Question: has any major league baseball team gone undefeated?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The crank sensor can be used in combination with a similar camshaft position sensor to monitor the relationship between the pistons and valves in the engine, which is particularly important in engines with variable valve timing. This method is also used to ``synchronise'' a four stroke engine upon starting, allowing the management system to know when to inject the fuel. It is also commonly used as the primary source for the measurement of engine speed in revolutions per minute. Question: is an engine speed sensor the same as a crankshaft sensor?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: This article lists in chronology and provides additional details of incidents in which a firearm was discharged at a school infrastructure or campus in the United States, including incidents of shootings on a school bus. This list contains school shooting incidents that occurred on the campuses of K-12 public schools and private schools as well as colleges and universities. It excludes incidents that occurred during wars or police actions as well as murder-suicides by rejected suitors or estranged spouses and suicides or suicide attempts involving only one person. Mass shootings by staff of schools that involve only other employees are covered at workplace killings. This list does not include bombings such as the Bath School disaster. Question: has a school shooting happened in a private school?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Though historically ``Blue Cross'' was used for hospital coverage while ``Blue Shield'' was used for medical coverage, today that split only exists for traditional health insurance plans in Pennsylvania. Two independent companies operate in central Pennsylvania, Highmark Blue Shield (Pittsburgh) and Capital Blue Cross (Central Pennsylvania) . In southeastern Pennsylvania, Independence Blue Cross (Philadelphia) has a joint marketing agreement with Highmark Blue Shield (Pittsburgh) for their separate hospital and medical plans. However, Independence Blue Cross, like most of its sister Blue Cross-Blue Shield companies, cover most of their customers under managed care plans such as HMOs and PPOs which provide hospital and medical care in one policy. Question: is blue cross blue shield a managed care organization?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: In the United States, an honours degree (or honors degree in US spelling) requires a thesis or project work beyond that needed for the normal bachelor's program. Honours programs in the US are taken alongside the rest of the degree and often have a minimum GPA requirement for entry, which can vary between institutions. Some institutions do not have a separate honours program, but instead refer to bachelor's degrees awarded with Latin honours, which may be based either on GPA or class position, as honours degrees. Question: is a master's degree higher than an honours degree?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " There is no direct comparison between a master's degree and an honours degree in the passage. An honours degree is a type of bachelor's degree that requires additional thesis or project work and often has a minimum GPA requirement. Some institutions may refer to bachelor's degrees awarded with Latin honours as honours degrees. Therefore, the score is: 50"
    },
    {
        "question": "Passage: Roxanne Roxanne is a 2017 American drama film written and directed by Michael Larnell. It stars Chant\u00e9 Adams, Mahershala Ali, Nia Long, Elvis Nolasco, Kevin Phillips and Shenell Edmonds. The film revolves around the life of rapper Roxanne Shant\u00e9. It was screened in the U.S. Dramatic Competition section of the 2017 Sundance Film Festival. Question: is the movie roxanne roxanne a true story?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Eastbound vehicles must pay a toll to cross the bridge; as with all Hudson River crossings along the North River, westbound vehicles cross for free. As of December 6, 2015, the cash tolls going from New Jersey to New York are $15 for both cars and motorcycles. E-ZPass users are charged $10.50 for cars and $9.50 for motorcycles during off-peak hours, and $12.50 for cars and $11.50 for motorcycles during peak hours. Trucks are charged cash tolls of $20.00 per axle, with discounted peak, off-peak, and overnight E-ZPass tolls. A discounted carpool toll ($6.50) is available at all times for cars with three or more passengers using NY or NJ E-ZPass, who proceed through a staffed toll lane (provided they have registered with the free ``Carpool Plan''). There is an off-peak toll of $7.00 for qualified low-emission passenger vehicles, which have received a Green E-ZPass based on registering for the Port Authority Green Pass Discount Plan. Question: do you have to pay both ways on the george washington bridge?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: In the U.S. in 2010, the bottling size was reduced from a typical 12 oz. per serving to 11.2 oz. per serving which is equivalent to the typical metric serving of 0.33L. Question: is red stripe light sold in the us?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " There is no information in the passage about Red Stripe Light being sold in the US. The student's answer cannot be determined to be true or false based on the information provided. Therefore the score is: 50"
    },
    {
        "question": "Passage: The Waterloo & City line (colloquially known as The Drain) is a London Underground line that runs between Waterloo and Bank with no intermediate stops. Its primary traffic consists of commuters from south-west London, Surrey and Hampshire arriving at Waterloo main line station and travelling forward to the City of London financial district, and for this reason the line does not normally operate on Sundays. Question: does the waterloo and city line run on sundays?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The Welland Canal is a ship canal in Ontario, Canada, connecting Lake Ontario and Lake Erie. It forms a key section of the St. Lawrence Seaway. Traversing the Niagara Peninsula from Port Weller to Port Colborne, it enables ships to ascend and descend the Niagara Escarpment and bypass Niagara Falls. Question: can you boat from lake erie to lake ontario?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: Black garlic can be eaten alone, on bread, or used in soups, sauces, crushed into a mayonnaise or simply tossed into a vegetable dish. A vinaigrette can be made with black garlic, sherry vinegar, soy, a neutral oil, and Dijon mustard. Its softness increases with water content. Question: is japanese black garlic supposed to be mushy?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: In the Northern Hemisphere, the winter sun (December, January, February) rises in the southeast, transits the celestial meridian at a low angle in the south (more than 43\u00b0 above the southern horizon in the tropics), and then sets in the southwest. It is on the south (equator) side of the house all day long. A vertical window facing south (equator side) is effective for capturing solar thermal energy. For comparison, the winter sun in the Southern Hemisphere (June, July, August) rises in the northeast, peaks out at a low angle in the north (more than halfway up from the horizon in the tropics), and then sets in the northwest. There, the north-facing window would let in plenty of solar thermal energy to the house. Question: does the sun ever shine from the north?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The galactic year, also known as a cosmic year, is the duration of time required for the Sun to orbit once around the center of the Milky Way Galaxy. Estimates of the length of one orbit range from 225 to 250 million terrestrial years. The Solar System is traveling at an average speed of 828,000 km/h (230 km/s) or 514,000 mph (143 mi/s) within its trajectory around the galactic center, a speed at which an object could circumnavigate the Earth's equator in 2 minutes and 54 seconds; that speed corresponds to approximately one 1300th of the speed of light. Question: does the sun orbit around the milky way?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 0"
    },
    {
        "question": "Passage: Maze Runner: The Death Cure (also known simply as The Death Cure) is a 2018 American dystopian science fiction action film directed by Wes Ball and written by T.S. Nowlin, based on the novel The Death Cure written by James Dashner. It is the sequel to the 2015 film Maze Runner: The Scorch Trials and the third and final installment in the Maze Runner film series. The film stars Dylan O'Brien, Kaya Scodelario, Thomas Brodie-Sangster, Dexter Darden, Nathalie Emmanuel, Giancarlo Esposito, Aidan Gillen, Walton Goggins, Ki Hong Lee, Jacob Lofland, Katherine McNamara, Barry Pepper, Will Poulter, Rosa Salazar, and Patricia Clarkson. Question: is the maze runner death cure the last movie?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: In the United States education system, social studies is the integrated study of multiple fields of social science and the humanities, including history, geography, and political science. The term was first coined by American educators around the turn of the twentieth century as a catch-all for these subjects, as well as others which did not fit into the traditional models of lower education in the United States, such as philosophy and psychology. Question: is social studies and social science the same?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: In December of 2015, Bad Lip Reading simultaneously released three new videos, one for each of the three films in the original Star Wars trilogy. These videos found BLR using guest voices for the first time, featuring Jack Black as Darth Vader, Maya Rudolph as Princess Leia, and Bill Hader in multiple roles. The Empire Strikes Back BLR video featured a scene of Yoda singing to Luke about an unfortunate encounter with a seagull on the beach. BLR would later expand this scene into a full-length standalone song known as ``Seagulls! (Stop It Now)'', which was released in November 2016 (eventually hitting #1 on the Billboard Comedy Digital Tracks chart.) As of late 2017, the ``Seagulls!'' video is Bad Lip Reading's second most viewed YouTube upload, and most popular musical production. In the song, Yoda sings to Luke Skywalker about the dangers posed by vicious seagulls if one dares to go to the beach. Mark Hamill, who played Luke Skywalker in the Star Wars films, publicly praised ``Seagulls!'' (and Bad Lip Reading in general) while speaking at Star Wars Celebration in 2017: ``I love them, and I showed Carrie (Fisher) the Yoda one... we were dying. I showed it to her in her trailer. She loved it. I retweeted it... and (BLR) contacted me and said 'Do you want to do Bad Lip Reading?' And I said, 'I'd love to...'''. Hamill and Bad Lip Reading would go on to collaborate on Bad Lip Reading's version of The Force Awakens, with Hamill providing the voice of Han Solo. Question: is seagulls stop it now a real song?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: Citizens of countries in the European Economic Area (other than British and Irish citizens) and Swiss citizens obtain permanent residence status automatically after five years' residence in the United Kingdom exercising Treaty rights rather than ILR. The rights of EEA citizens are not governed by UK Immigration Regulations but rather the EEA Regulations. Question: do eu citizens have indefinite leave to remain?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": "NILL"
    },
    {
        "question": "Passage: A bagel and cream cheese (also known as bagel with cream cheese) is a common food pairing in American cuisine, the cuisine of New York City, and American Jewish cuisine, consisting in its basic form of an open-faced sandwich made of a bagel spread with cream cheese. The bagel is typically sliced into two pieces, and can be served as-is or toasted. The basic bagel with cream cheese serves as the base for other sandwiches such as the ``lox and schmear'', a staple of delicatessens in the New York area, and across the U.S. Question: is a bagel with cream cheese a sandwich?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: The eurozone ( pronunciation (help info)), officially called the euro area, is a monetary union of 19 of the 28 European Union (EU) member states which have adopted the euro (\u20ac) as their common currency and sole legal tender. The monetary authority of the eurozone is the Eurosystem. The other nine members of the European Union continue to use their own national currencies, although most of them are obliged to adopt the euro in the future. Question: do european countries still have their own currency?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: A player is in an 'offside position' if they are in the opposing team's half of the field and also ``nearer to the opponents' goal line than both the ball and the second-last opponent.'' The 2005 edition of the Laws of the Game included a new IFAB decision that stated, ``In the definition of offside position, 'nearer to his opponents' goal line' means that any part of their head, body or feet is nearer to their opponents' goal line than both the ball and the second last opponent. The arms are not included in this definition''. By 2017, the wording had changed to say that, in judging offside position, ``The hands and arms of all players, including the goalkeepers, are not considered.'' In other words, a player is in an offside position if two conditions are met: Question: does the goalkeeper count in the offside rule?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: International Blind Sports Federation rules require that any time during a game in which one team has scored ten (10) more goals than the other team that game is deemed completed. In US high school soccer, most states use a mercy rule that ends the game if one team is ahead by 10 or more goals at any point from halftime onward. Youth soccer leagues use variations on the rule. Question: is there a mercy rule in professional soccer?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: For a standard game of Klondike, drawing three cards at a time and placing no limit on the number of re-deals, the number of possible hands is over 7067800000000000000\u26608\u00d710, or an 8 followed by 67 zeros. About 79% of the games are theoretically winnable, but in practice, human players do not win 79% of games played, due to wrong moves that cause the game to become unwinnable. If one allows cards from the foundation to be moved back to the tableau, then between 82% and 91.5% are theoretically winnable. Note that these results depend on complete knowledge of the positions of all 52 cards, which a player does not possess. Another recent study has found the Draw 3, Re-Deal Infinite to have a 83.6% win rate after 1000 random games were solved by a computer solver. The issue is that a wrong move cannot be known in advance whenever more than one move is possible. The number of games a skilled player can probabilistically expect to win is at least 43%. In addition, some games are ``unplayable'' in which no cards can be moved to the foundations even at the start of the game; these occur in only 0.25% (1 in 400) of hands dealt. Question: is there always a way to win solitaire?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Rule 6.2 of the 2008--09 Official NHL Rulebook indicates that ``(only) when the captain is not in uniform, the coach shall have the right to designate three alternate captains. This must be done prior to the start of the game.'' Many NHL teams with a named captain select more than two alternate captains and rotate the ``A'' among these players throughout the season. There are currently seven teams without captains: the Arizona Coyotes, the Buffalo Sabres, the New York Islanders, the New York Rangers, the Toronto Maple Leafs, the Vancouver Canucks, and the Vegas Golden Knights. Question: does the vegas golden knights have a captain?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: In a common U.S. butchery, the steak is cut from the rear back portion of the animal, continuing off the short loin from which T-bone, porterhouse, and club steaks are cut. The sirloin is actually divided into several types of steak. The top sirloin is the most prized of these and is specifically marked for sale under that name. The bottom sirloin, which is less tender and much larger, is typically marked for sale simply as ``sirloin steak''. The bottom sirloin in turn connects to the sirloin tip roast. Question: is top sirloin steak the same as sirloin steak?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: By March 2015, after not playing any matches, India reached their lowest FIFA ranking position of 173. A couple months prior, Stephen Constantine was re-hired as the head coach after first leading India more than a decade before. Constantine's first major assignment back as the India head coach were the 2018 FIFA World Cup qualifiers. After making it through the first round of qualifiers, India crashed out during the second round, losing seven of their eight matches and thus, once again, failed to qualify for the World Cup. Question: did india qualify for football world cup 2018?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: The following is the list of teams to overcome 3--1 series deficits by winning three straight games to win a best-of-seven playoff series. In the history of major North American pro sports, teams that were down 3--1 in the series came back and won the series 52 times, more than half of them were accomplished by National Hockey League (NHL) teams. Teams overcame 3--1 deficit in the final championship round eight times, six were accomplished by Major League Baseball (MLB) teams in the World Series. Teams overcoming 3--0 deficit by winning four straight games were accomplished five times, four times in the NHL and once in MLB. Question: has anyone come back from 3-0 in the nba finals?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Fans of author J.R.R. Tolkien have drawn attention to the similarities between his novel The Lord of the Rings and the Harry Potter series; specifically Tolkien's Wormtongue and Rowling's Wormtail, Tolkien's Shelob and Rowling's Aragog, Tolkien's Gandalf and Rowling's Dumbledore, Tolkien's Nazg\u00fbl and Rowling's Dementors, Old Man Willow and the Whomping Willow and the similarities between both authors' antagonists, Tolkien's Dark Lord Sauron and Rowling's Lord Voldemort (both of whom are sometimes within their respective continuities unnamed due to intense fear surrounding their names; both often referred to as 'The Dark Lord'; and both of whom are, during the time when the main action takes place, seeking to recover their lost power after having been considered dead or at least no longer a threat). Several reviews of Harry Potter and the Deathly Hallows noted that the locket used as a horcrux by Voldemort bore comparison to Tolkien's One Ring, as it negatively affects the personality of the wearer. Rowling maintains that she had not read The Hobbit until after she completed the first Harry Potter novel (though she had read The Lord of the Rings as a teenager) and that any similarities between her books and Tolkien's are ``Fairly superficial. Tolkien created a whole new mythology, which I would never claim to have done. On the other hand, I think I have better jokes.'' Tolkienian scholar Tom Shippey has maintained that ``no modern writer of epic fantasy has managed to escape the mark of Tolkien, no matter how hard many of them have tried''. Question: was harry potter inspired by lord of the rings?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: A new film of the musical was planned in 2008 with a screenplay by Emma Thompson but the project did not materialize. Keira Knightley, Carey Mulligan, and Colin Firth were among those in consideration for the lead roles. Question: is there a remake of my fair lady?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 30"
    },
    {
        "question": "Passage: Per capita income is often used c measure an area's average income. This is used to see the wealth of the population with those of others. Per capita income is often used to measure a country's standard of living. It is usually expressed in terms of a commonly used international currency such as the euro or United States dollar, and is useful because it is widely known, is easily calculable from readily available gross domestic product (GDP) and population estimates, and produces a useful statistic for comparison of wealth between sovereign territories. This helps to ascertain a country's development status. It is one of the three measures for calculating the Human Development Index of a country. Question: is gdp per capita same as per capita income?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: A press release, news release, media release, press statement or video release is a written or recorded communication directed at members of the news media for the purpose of announcing something ostensibly newsworthy. Typically, they are mailed, faxed, or e-mailed to assignment editors and journalists at newspapers, magazines, radio stations, online media, television stations or television networks. Question: is news release and press release the same?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: In 1866, at the behest of Chief Justice Chase, Congress passed an act providing that the next three justices to retire would not be replaced, which would thin the bench to seven justices by attrition. Consequently, one seat was removed in 1866 and a second in 1867. In 1869, however, the Circuit Judges Act returned the number of justices to nine, where it has since remained. Question: can we have more than 9 supreme court justices?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: At the end of season 10, she says goodbye to her fellow co-workers she has come to know and love including Owen and Meredith. Cristina and Meredith share special moments together reminiscing about all the horrors they went through and dancing it out one last time. Cristina leaves for Zurich with surgical intern Shane Ross, who chooses to leave in order to study under her in Switzerland. Question: does christina yang die in season 10 episode 24?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: The Isle of Man (Manx: Ellan Vannin (\u02c8\u025blj\u0259n \u02c8van\u026an)), sometimes referred to simply as Mann (/m\u00e6n/; Manx: Mannin (\u02c8man\u026an)), is a self-governing British Crown dependency, an island in the Irish Sea between Great Britain and Ireland. The head of state is Queen Elizabeth II, who holds the title of Lord of Mann and is represented by a Lieutenant Governor. Defence is the responsibility of the United Kingdom. Question: is the isle of man part of the uk?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Regan, who was not allowed in the basement previously, sees her father's notes on the creatures and on his experimentation with several different implants. When the creature returns to invade the basement, Regan places the boosted implant on a nearby microphone, magnifying the feedback to ward off the creature. Painfully disoriented, the creature exposes the flesh beneath its armored head, and Evelyn shoots the creature in the head with a shotgun, destroying its head and killing it. The family views a CCTV monitor, showing two creatures attracted by the noise of the shotgun blast approaching the house. With their newly acquired knowledge of the creatures' weakness, the members of the family arm themselves and prepare to fight back. Question: do they live at the end of a quiet place?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Years later, Millard is in Lakeside high school and dating Shannon. Hoping to impress his father, he begins playing football. However, he is injured, breaking both ankles and ending his career. In order to make up the credits he would miss from football, he signs up for music class, the only available class left. Initially, Millard is assigned to be a sound technician. After the director catches him singing in the empty auditorium of Lakeside high school, she casts him as Curly, the lead role in the school production of Oklahoma. He doesn't tell his father of his role in the play, and while Bart has risen to the singing demands of the part, Arthur subsequently collapses with severe abdominal pain, but refuses to tell Bart or Shannon about his cancer diagnosis. The following morning, Millard voices his frustrations with his father and is assaulted by his father, who smashes a plate over his head. Shannon presses Bart to open up, but he responds by breaking up with her and leaving to seek his fortune in the city after graduation. Question: did bart millard sing in movie i can only imagine?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: The deed in lieu of foreclosure offers several advantages to both the borrower and the lender. The principal advantage to the borrower is that it immediately releases him/her from most or all of the personal indebtedness associated with the defaulted loan. The borrower also avoids the public notoriety of a foreclosure proceeding and may receive more generous terms than he/she would in a formal foreclosure. Another benefit to the borrower is that it hurts his/her credit less than a foreclosure does. Advantages to a lender include a reduction in the time and cost of a repossession, lower risk of borrower revenge (metal theft and vandalism of the property before sheriff eviction), and additional advantages if the borrower subsequently files for bankruptcy. Question: is a deed in lieu considered a foreclosure?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Tonic water (or Indian tonic water) is a carbonated soft drink in which quinine is dissolved. Originally used as a prophylactic against malaria, tonic water usually now has a significantly lower quinine content and is consumed for its distinctive bitter flavor, which is similar to a sour grapefruit. It is often used in mixed drinks, particularly in gin and tonic. Question: are tonic water and quinine water the same thing?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: There are normally two broods, with the original nest being reused for the second brood and being repaired and reused in subsequent years. The female lays two to seven, but typically four or five, reddish-spotted white eggs. The clutch size is influenced by latitude, with clutch sizes of northern populations being higher on average than southern populations. The eggs are 20 mm \u00d7 14 mm (0.79 in \u00d7 0.55 in) in size, and weigh 1.9 g (0.067 oz), of which 5% is shell. In Europe, the female does almost all the incubation, but in North America the male may incubate up to 25% of the time. The incubation period is normally 14--19 days, with another 18--23 days before the altricial chicks fledge. The fledged young stay with, and are fed by, the parents for about a week after leaving the nest. Occasionally, first-year birds from the first brood will assist in feeding the second brood. Compared to those from early broods, juvenile barn swallows from late broods have been found to migrate at a younger age, fuel less efficiently during migration and have lower return rates the following year. Question: do barn swallows lay eggs more than once a year?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Humans have a four-chambered heart consisting of the right atrium, left atrium, right ventricle, and left ventricle. The atria are the two upper chambers. The right atrium receives and holds deoxygenated blood from the superior vena cava, inferior vena cava, anterior cardiac veins and smallest cardiac veins and the coronary sinus, which it then sends down to the right ventricle (through the tricuspid valve) which in turn sends it to the pulmonary artery for pulmonary circulation. The left atrium receives the oxygenated blood from the left and right pulmonary veins, which it pumps to the left ventricle (through the mitral valve) for pumping out through the aorta for systemic circulation. Question: is there a difference in structure of the two atria?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: High-performance sailing is achieved with low forward surface resistance--encountered by catamarans, sailing hydrofoils, iceboats or land sailing craft--as the sailing craft obtains motive power with its sails or aerofoils at speeds that are often faster than the wind. Question: can a boat sail faster than the wind?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " There is no clear answer in the student's response, so I cannot give a score. The student should provide a clear answer, either Yes or No, to the question based on the information provided in the passage."
    },
    {
        "question": "Passage: Arm span or reach (sometimes referred to as wingspan) is the physical measurement of the length from one end of an individual's arms (measured at the fingertips) to the other when raised parallel to the ground at shoulder height at a 90\u00b0 angle. The average reach correlates to the person's height. Age and sex have to be taken into account to best predict height from arm span. Question: is it true that your arm span is your height?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " There is a possibility that the student's answer might be correct, but it is not definite. To determine the score, we need to consider the context and the information provided in the passage.\\n\\nThe passage states that the average reach correlates to a person's height. However, it also mentions that age and sex need to be taken into account to best predict height from arm span.\\n\\nTherefore, the score is: 60"
    },
    {
        "question": "Passage: The Fire Tablet, formerly called the Kindle Fire, is a tablet computer developed by Amazon.com. Built with Quanta Computer, the Kindle Fire was first released in November 2011, featuring a color 7-inch multi-touch display with IPS technology and running a custom version of Google's Android operating system called Fire OS. The Kindle Fire HD followed in September 2012, and the Kindle Fire HDX in September 2013. In September 2014, when the fourth generation was introduced, the name ``Kindle'' was dropped. In September 2015, the fifth generation Fire 7 was released, followed by the sixth generation Fire HD 8, in September 2016. The seventh generation Fire 7 was released in June 2017. Question: is a fire 7 the same as a kindle?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " There is no mention of whether a Fire 7 is the same as a Kindle in the passage. The passage mainly provides information about the history of the Fire Tablet, including its various versions and name changes. Therefore, the score is: 50"
    },
    {
        "question": "Passage: In mathematics, and more specifically set theory, the empty set or null set is the unique set having no elements; its size or cardinality (count of elements in a set) is zero. Some axiomatic set theories ensure that the empty set exists by including an axiom of empty set; in other theories, its existence can be deduced. Many possible properties of sets are vacuously true for the empty set. Question: is an empty set an element of an empty set?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: Russia has participated in 4 FIFA World Cups since its independence in December 1991. The Russian Federation played their first international match against Mexico on 16 August 1992 winning 2-0. Their first participation in a World Cup was the United States of America in 1994 and they achieved 18th place. In 1946 the Soviet Union was accepted by FIFA and played their first World Cup in Sweden 1958. The Soviet Union represented 15 Socialist republics and various football federations, and the majority of players came from the Dynamo Kyiv team of the Ukrainian SSR. The Soviet Union national football team played in 7 World Cups. Their best performance was reaching 4th place in England 1966. However Soviet football was dissolved in 1991 when Belarus, Russia and Ukraine declared independence under the Belavezha Accords. The CIS national football team (Commonwealth of Independent States) was formed with other independent nations in 1992 but did not participate in any World Cups. Question: has russia ever made it to the world cup finals?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Tanjore painting is an important form of classical South Indian painting native to the town of Tanjore in Tamil Nadu. The art form dates back to the early 9th century, a period dominated by the Chola rulers, who encouraged art and literature. These paintings are known for their elegance, rich colours, and attention to detail. The themes for most of these paintings are Hindu Gods and Goddesses and scenes from Hindu mythology. In modern times, these paintings have become a much sought-after souvenir during festive occasions in South India. Question: is tanjore a traditional indian folk art form?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: At any time, police may approach a person and ask questions. The objective may simply be a friendly conversation; however, the police also may suspect involvement in a crime, but lack ``specific and articulable facts'' that would justify a detention or arrest, and hope to obtain these facts from the questioning. The person approached is not required to identify himself or answer any other questions, and may leave at any time. Police are not usually required to tell a person that he is free to decline to answer questions and go about his business; however, a person can usually determine whether the interaction is consensual by asking, ``Am I free to go?'' Question: do i have to give my name to a police officer?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: A function f is said to be continuously differentiable if the derivative f\u2032(x) exists and is itself a continuous function. Though the derivative of a differentiable function never has a jump discontinuity, it is possible for the derivative to have an essential discontinuity. For example, the function Question: is the derivative of a continuous function always continuous?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The Air Force usually does not have fighter aircraft escort the presidential aircraft over the United States but it has occurred, for example during the attack on the World Trade Center. Question: does air force one fly with fighter escort?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " \u9996\u5148\uff0c\u6211\u9700\u8981\u6f84\u6e05\u4e00\u70b9\uff0c\u4f5c\u4e3a\u4e00\u4e2a\u4eba\u5de5\u667a\u80fd\u52a9\u624b\uff0c\u6211\u65e0\u6cd5\u9605\u8bfb\u6216\u7406\u89e3\u5177\u4f53\u7684\u6587\u672c\u5185\u5bb9\uff0c\u4e5f\u65e0\u6cd5\u5bf9\u5b66\u751f\u7684\u7b54\u6848\u8fdb\u884c\u8bc4\u5206\u3002\u4f46\u662f\uff0c\u6211\u53ef\u4ee5\u5e2e\u52a9\u4f60\u7406\u89e3\u95ee\u9898\u548c\u7b54\u6848\u7684\u5185\u5bb9\u3002\\n\\n\u95ee\u9898\u6d89\u53ca\u5230\u7a7a\u519b\u4e00\u53f7\u662f\u5426\u5728\u98de\u884c\u65f6\u6709\u6218\u6597\u673a\u62a4\u822a\u3002\u6839\u636e\u4f60\u63d0\u4f9b\u7684\u4fe1\u606f\uff0c\u5b66\u751f\u7684\u7b54\u6848\u662f\u201cTrue\u201d\uff0c\u4e5f\u5c31\u662f\u8bf4\u7a7a\u519b\u4e00\u53f7\u5728\u98de\u884c\u65f6\u662f\u6709\u6218\u6597\u673a\u62a4\u822a\u7684\u3002\\n\\n\u6839\u636e\u4f60\u63d0\u4f9b\u7684\u6587\u7ae0\u5185\u5bb9\uff0c\u201c\u7a7a\u519b\u901a\u5e38\u4e0d\u4f1a\u5728\u7f8e\u56fd\u5883\u5185\u4e3a\u603b\u7edf\u4e13\u673a\u63d0\u4f9b\u6218\u6597\u673a\u62a4\u822a\uff0c\u4f46\u662f\u8fd9\u79cd\u60c5\u51b5\u4e5f\u53d1\u751f\u8fc7\uff0c\u6bd4\u5982\u5728\u88ad\u51fb\u4e16\u8d38\u4e2d\u5fc3\u7684\u65f6\u5019\u3002\u201d\u8fd9\u53e5\u8bdd\u5e76\u6ca1\u6709\u660e\u786e\u8bf4\u660e\u7a7a\u519b\u4e00\u53f7\u662f\u5426\u603b\u662f\u6709\u6218\u6597\u673a\u62a4\u822a\uff0c\u53ea\u662f\u8bf4\u8fd9\u79cd\u60c5\u51b5\u66fe\u7ecf\u53d1\u751f\u8fc7\u3002\\n\\n\u56e0\u6b64\uff0c\u5982\u679c\u5b66\u751f\u7684\u7b54\u6848\u662f\u57fa\u4e8e\u5bf9\u6587\u7ae0\u7684\u51c6\u786e\u7406\u89e3\uff0c\u90a3\u4e48\u4ed6\u7684\u7b54\u6848\u5e94\u8be5\u662f\u6b63\u786e\u7684\u3002\u4f46\u662f\uff0c\u5177\u4f53\u7684\u8bc4\u5206\u9700\u8981\u7531\u8001\u5e08\u6839\u636e\u5b66\u751f\u7684\u7b54\u6848\u548c\u7406\u89e3\u7a0b\u5ea6\u6765\u51b3\u5b9a\u3002"
    },
    {
        "question": "Passage: European Summer Time is the variation of standard clock time that is applied in most European countries, not including Iceland, Georgia, Azerbaijan, Belarus, Turkey and Russia -- in the period between spring and autumn, during which clocks are advanced by one hour from the time observed in the rest of the year, in order to make the most efficient use of seasonal daylight. It corresponds to the notion and practice of ``daylight saving time'' to be found in many other parts of the world. Question: are england the only country to change their clocks?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: Some tax protesters such as Edward Brown and tax protester organizations such as the We the People Foundation have used the phrase ``show me the law'' to argue that the Internal Revenue Service refuses to disclose the laws that impose the legal obligation to file Federal income tax returns or pay Federal income taxes--and to argue that there must be no law imposing Federal income taxes. Question: is there a law that says we have to pay taxes?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: In basketball, a flagrant foul is a personal foul that involves excessive or violent contact that could injure the fouled player. A flagrant foul may be unintentional or purposeful; the latter type is also called an ``intentional foul'' in the NBA. However, most intentional fouls are not considered flagrant and fouling intentionally is an accepted tactic to regain possession of the ball with minimal time off the game clock. Question: does a flagrant foul count as a personal foul?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The 21 World Cup tournaments have been won by eight national teams. Brazil have won five times, and they are the only team to have played in every tournament. The other World Cup winners are Germany and Italy, with four titles each; Argentina, France and inaugural winner Uruguay, with two titles each; and England and Spain with one title each. Question: has the u s ever won a world cup?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: The Territory of Hawaii or Hawaii Territory was an organized incorporated territory of the United States that existed from August 12, 1898, until August 21, 1959, when most of its territory, excluding Palmyra Island and the Stewart Islands, was admitted to the Union as the fiftieth U.S. state, the State of Hawaii. The Hawaii Admission Act specified that the State of Hawaii would not include the distant Palmyra Island, the Midway Islands, Kingman Reef, and Johnston Atoll, which includes Johnston (or Kalama) Island and Sand Island, and the Act was silent regarding the Stewart Islands. Question: is hawaii part of the united states territory?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Most competitions only allow each team to make a maximum of three substitutions during a game and a fourth substitute during extra time, although more substitutions are often permitted in non-competitive fixtures such as friendlies. A fourth substitution in extra time was first implemented in recent tournaments, including the 2016 Summer Olympic Games, the 2017 FIFA Confederations Cup and the 2017 CONCACAF Gold Cup final. A fourth substitute in extra time has been approved for use in the elimination rounds at the 2018 FIFA World Cup, the UEFA Champions League and the UEFA Europa League. Each team nominates a number of players (typically between five and seven, depending on the competition) who may be used as substitutes; these players typically sit in the technical area with the coaches, and are said to be ``on the bench''. When the substitute enters the field of play it is said they have come on or have been brought on, while the player they are substituting is coming off or being brought off. Question: can a player be substituted twice in football?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: Toys ``R'' Us expanded as a chain, becoming predominant in its niche field of toy retail. Represented by cartoon mascot Geoffrey the Giraffe from 1969, Toys ``R'' Us eventually branched out into launching the stores Babies ``R'' Us, Toys ``R'' Us Express, and the now-defunct Kids ``R'' Us. Question: is babies r us and toys r us the same?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: The central nervous system is responsible for the orderly recruitment of motor neurons, beginning with the smallest motor units. Henneman's size principle indicates that motor units are recruited from smallest to largest based on the size of the load. For smaller loads requiring less force, slow twitch, low-force, fatigue-resistant muscle fibers are activated prior to the recruitment of the fast twitch, high-force, less fatigue-resistant muscle fibers. Larger motor units are typically composed of faster muscle fibers that generate higher forces. Question: does the size of a motor unit vary?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " There is no information provided in the passage to determine if the size of a motor unit varies or not. Therefore, the score is: 50"
    },
    {
        "question": "Passage: In most jurisdictions, secondary education in the United States refers to the last four years of statutory formal education (grade nine through grade twelve) either at high school or split between a final year of 'junior high school' and three in high school. Question: is secondary school the same as high school in the united states?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Maple syrup is a syrup usually made from the xylem sap of sugar maple, red maple, or black maple trees, although it can also be made from other maple species. In cold climates, these trees store starch in their trunks and roots before winter; the starch is then converted to sugar that rises in the sap in late winter and early spring. Maple trees are tapped by drilling holes into their trunks and collecting the exuded sap, which is processed by heating to evaporate much of the water, leaving the concentrated syrup. Question: does maple syrup come straight from the tree?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": "NILL"
    },
    {
        "question": "Passage: In mathematics, a ratio is a relationship between two numbers indicating how many times the first number contains the second. For example, if a bowl of fruit contains eight oranges and six lemons, then the ratio of oranges to lemons is eight to six (that is, 8:6, which is equivalent to the ratio 4:3). Similarly, the ratio of lemons to oranges is 6:8 (or 3:4) and the ratio of oranges to the total amount of fruit is 8:14 (or 4:7). Question: does it matter which number comes first in a ratio?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: The principal bridesmaid, if one is so designated, may be called the chief bridesmaid or maid of honor if she is unmarried, or the matron of honor if she is married. A junior bridesmaid is a girl who is clearly too young to be married, but who is included as an honorary bridesmaid. In the United States, typically only the maid/matron of honor and the best man are the official witnesses for the wedding license. Question: do you have to call a married woman matron of honor?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The turkey vulture (Cathartes aura), also known in some North American regions as the turkey buzzard (or just buzzard), and in some areas of the Caribbean as the John crow or carrion crow, is the most widespread of the New World vultures. One of three species in the genus Cathartes of the family Cathartidae, the turkey vulture ranges from southern Canada to the southernmost tip of South America. It inhabits a variety of open and semi-open areas, including subtropical forests, shrublands, pastures, and deserts. Question: is a turkey vulture and a buzzard the same thing?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: A key difference between the two sports is that in rugby union both sets of forwards try to push the opposition backwards whilst competing for the ball and thus the team that did not throw the ball into the scrum have some minimal chance of winning the possession. In practice, however, the team with the 'put-in' usually keeps possession (92% of the time with the feed) and put-ins are not straight. Forwards in rugby league do not usually push in the scrum, scrum-halves often feed the ball directly under the legs of their own front row rather than into the tunnel, and the team with the put-in usually retains possession (thereby making the 40/20 rule workable). Question: can you push in a rugby league scrum?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: On March 7, 2013, Variety confirmed that Disney has already approved plans for a sequel with Mitchell Kapner and Joe Roth returning as screenwriter and producer respectively. Mila Kunis said during an interview with E! News, ``We're all signed on for sequels.'' On March 8, 2013, Sam Raimi told Bleeding Cool that he has no plans to direct the sequel, saying, ``I did leave some loose ends for another director if they want to make the picture,'' and that ``I was attracted to this story but I don't think the second one would have the thing I would need to get me interested.'' On March 11, 2013, Kapner and Roth have said to the Los Angeles Times that the sequel will ``absolutely not'' involve Dorothy Gale, with Kapner pointing out that there are twenty years between the events of the first film and Dorothy's arrival, and ``a lot can happen in that time.'' Question: is there a sequel to oz the great and powerful?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: On October 20, 1977 -- three days after the release of the band's fifth studio album Street Survivors -- a chartered plane on which the members and crew were travelling crashed in Gillsburg, Mississippi. Six people died in the accident, including band members Ronnie Van Zant, Steve Gaines and Cassie Gaines; many of the other passengers onboard were seriously injured, including Wilkeson who was left in a critical condition and reportedly declared dead three times. The group disbanded after the crash. In 1978, a collection of previously unreleased recordings from 1971 and 1972 was released as Skynyrd's First and... Last. The following year, the surviving members (with the exception of Wilkeson) reunited at Volunteer Jam for a performance of ``Free Bird'' with Charlie Daniels and his band. Question: are any original members of lynyrd skynyrd still alive?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": "NILL"
    },
    {
        "question": "Passage: The film was shot in black and white in the styles and motifs of German Expressionism (bizarre shadows, stylized dialogue, distorted perspectives, surrealistic sets, odd camera angles) to create a simplified and disturbing mood that reflects the sinister character of Powell, the nightmarish fears of the children, and the sweetness of their savior Rachel. Due to the film's visual style and themes, it is also often categorized as a film noir. Question: is night of the hunter a film noir?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Ross-on-Wye railway station is a former junction railway station on the Hereford, Ross and Gloucester Railway constructed just to the north of the Herefordshire town of Ross-on-Wye. It was the terminus of the Ross and Monmouth Railway which joined the Hereford, Ross and Gloucester Railway just south of the station. Question: does ross on wye have a train station?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The sixth and final season premiered on 19 August 2018. Question: is a place to call home finished for good?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Stop & Shop/Giant-Landover was a combined supermarket chain owned by the American subsidiary of the Dutch retailer Ahold. The company took its form in 2004, after Ahold decided to combine the operations of its New England-based Stop & Shop chain with its DMV-based Giant Food chain to create the largest supermarket company in the Mid-Atlantic States. Giant's headquarters relocated in Landover, Maryland, and Stop & Shop kept their headquarters in Quincy, Massachusetts. This combination failed, as Mid-Atlantic market area shoppers grocery needs did not align with those of Stop & Shop's offerings. In 2011 the two companies were separated and now operate independently. The separation of Stop & Shop/Giant-Landover, also brought the separation of the Stop & Shop Supermarket into two separate operating divisions, Stop & Shop-New England and Stop & Shop-New York. Both Giant Food and Stop & Shop's two divisions continue to share the same Fruit Basket Logo even though they all operate independently. Question: are stop and shop and giant owned by the same company?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: In transfusions of packed red blood cells, individuals with type O Rh D negative blood are often called universal donors. Those with type AB Rh D positive blood are called universal recipients. However, these terms are only generally true with respect to possible reactions of the recipient's anti-A and anti-B antibodies to transfused red blood cells, and also possible sensitization to Rh D antigens. One exception is individuals with hh antigen system (also known as the Bombay phenotype) who can only receive blood safely from other hh donors, because they form antibodies against the H antigen present on all red blood cells. Question: is blood type o positive a universal donor?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " There is a small error in the student's answer. Type O Rh D negative blood is considered a universal donor, NOT type O positive blood. Therefore the score is: 80"
    },
    {
        "question": "Passage: There is ample evidence that the areas surrounding the Amazon River were home to complex and large-scale indigenous societies, mainly chiefdoms who developed large towns and cities. Archaeologists estimate that by the time the Spanish conquistador De Orellana traveled across the Amazon in 1541, more than 3 million indigenous people lived around the Amazon. These pre-Columbian settlements created highly developed civilizations. For instance, pre-Columbian indigenous people on the island of Maraj\u00f3 may have developed social stratification and supported a population of 100,000 people. In order to achieve this level of development, the indigenous inhabitants of the Amazon rainforest altered the forest's ecology by selective cultivation and the use of fire. Scientists argue that by burning areas of the forest repetitiously, the indigenous people caused the soil to become richer in nutrients. This created dark soil areas known as terra preta de \u00edndio (``Indian dark earth''). Because of the terra preta, indigenous communities were able to make land fertile and thus sustainable for the large-scale agriculture needed to support their large populations and complex social structures. Further research has hypothesized that this practice began around 11,000 years ago. Some say that its effects on forest ecology and regional climate explain the otherwise inexplicable band of lower rainfall through the Amazon basin. Question: does the amazon river run through the amazon rainforest?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: Shower gels for men may contain the ingredient menthol, which gives a cooling and stimulating sensation on the skin, and some men's shower gels are also designed specifically for use on hair and body. Shower gels contain milder surfactant bases than shampoos, and some also contain gentle conditioning agents in the formula. This means that shower gels can also double as an effective and perfectly acceptable substitute to shampoo, even if they are not labelled as a hair and body wash. Washing hair with shower gel should give approximately the same result as using a moisturising shampoo. Question: is it bad to wash your hair with shower gel?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 85"
    },
    {
        "question": "Passage: Surgical treatments of ingrown toenails include a number of different options. If conservative treatment of a minor ingrown toenail does not succeed or if the ingrown toenail is severe, surgical management is recommended by a podiatrist. The initial surgical approach is typically a partial avulsion of the nail plate known as a wedge resection or a complete removal of the toenail. If the ingrown toenail reoccurs despite this treatment, destruction of the germinal matrix with phenol is recommended. Antibiotics are not needed if surgery is performed. Question: can ingrown toenails come back after being removed?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " There is a possibility of ingrown toenails coming back after being removed, as stated in the passage, \\\"If the ingrown toenail reoccurs despite this treatment, destruction of the germinal matrix with phenol is recommended.\\\" Therefore, the score is: 50"
    },
    {
        "question": "Passage: Marble Falls is located in southern Burnet County at 30\u00b034\u2032N 98\u00b017\u2032W\ufeff / \ufeff30.567\u00b0N 98.283\u00b0W\ufeff / 30.567; -98.283 (30.5741, -98.2782), on the banks of Lake Marble Falls. According to the Handbook of Texas website, the former falls were flooded by the lake, which was created by a shelf of limestone running diagonally across the Colorado River from northeast to southwest. The upper layer of limestone, brownish on the exterior but a deep blue inside, was so hard and cherty it was mistaken for marble. The falls were actually three distinct formations at the head of a canyon 1.25 miles (2.01 km) long, with a drop of some 50 feet (15 m) through the limestone strata. The natural lake and waterfall were covered when the Colorado River was dammed with the completion of Max Starcke Dam in 1951. A photo of the falls as they once existed can be seen at the website for the Wallace Guest House, a local bed and breakfast. Lake Marble Falls sits between Lake Lyndon B. Johnson to the north and Lake Travis to the south. The falls for which the city is named are now underwater but are revealed every few years when the lake is lowered. Question: is there a waterfall in marble falls tx?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Humans have a four-chambered heart consisting of the right atrium, left atrium, right ventricle, and left ventricle. The atria are the two upper chambers. The right atrium receives and holds deoxygenated blood from the superior vena cava, inferior vena cava, anterior cardiac veins and smallest cardiac veins and the coronary sinus, which it then sends down to the right ventricle (through the tricuspid valve) which in turn sends it to the pulmonary artery for pulmonary circulation. The left atrium receives the oxygenated blood from the left and right pulmonary veins, which it pumps to the left ventricle (through the mitral valve) for pumping out through the aorta for systemic circulation. Question: does the right atrium receive blood from the lungs?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " There is a part in the passage that states, \\\"The right atrium receives and holds deoxygenated blood from the superior vena cava, inferior vena cava, anterior cardiac veins and smallest cardiac veins and the coronary sinus.\\\" This information shows that the right atrium does receive blood, but it is deoxygenated blood, not blood from the lungs. The student's answer is incorrect.\\n\\nTherefore, the score is: 50"
    },
    {
        "question": "Passage: The Blue Ridge Parkway is a National Parkway and All-American Road in the United States, noted for its scenic beauty. The parkway, which is America's longest linear park, runs for 469 miles (755 km) through 29 Virginia and North Carolina counties, linking Shenandoah National Park to Great Smoky Mountains National Park. It runs mostly along the spine of the Blue Ridge, a major mountain chain that is part of the Appalachian Mountains. Its southern terminus is at U.S. 441 on the boundary between Great Smoky Mountains National Park and the Cherokee Indian Reservation in North Carolina, from which it travels north to Shenandoah National Park in Virginia. The roadway continues through Shenandoah as Skyline Drive, a similar scenic road which is managed by a different National Park Service unit. Both Skyline Drive and the Virginia portion of the Blue Ridge Parkway are part of Virginia State Route 48, though this designation is not signed. Question: is skyline drive part of the blue ridge parkway?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The Northern Ireland national football team have appeared in the finals of the FIFA World Cup on three occasions. Question: does northern ireland have a world cup team?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 70"
    },
    {
        "question": "Passage: The eighth season of the American legal drama Suits was ordered on January 30, 2018, and began airing on USA Network in the United States July 18, 2018. Question: are they making a season 8 of suits?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: While Life Unexpected received mostly positive reviews, it struggled in the ratings and was cancelled by The CW in 2011. The show has since been released on DVD, and it is available on Netflix as well as Amazon Video streaming services. Question: are they making a season 3 of life unexpected?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The first FA Cup Final to go to extra time and a replay was the 1875 final, between the Royal Engineers and the Old Etonians. The initial tie finished 1--1 but the Royal Engineers won the replay 2--0 in normal time. The last replayed final was the 1993 FA Cup Final, when Arsenal and Sheffield Wednesday fought a 1--1 draw. The replay saw Arsenal win the FA Cup, 2--1 after extra time. Question: can the fa cup final end in a tie?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Fetal surgery also known as Fetal reconstructive surgery antenatal surgery, prenatal surgery. is a growing branch of maternal-fetal medicine that covers any of a broad range of surgical techniques that are used to treat birth defects in fetuses who are still in the pregnant uterus. There are three main types: open fetal surgery, which involves completely opening the uterus to operate on the fetus; minimally invasive fetoscopic surgery, which uses small incisions and is guided by fetoscopy and sonography; and percutaneous fetal therapy, which involves placing a catheter under continuous ultrasound guidance. Question: can you do surgery on a baby in utero?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Citric acid can be added to ice cream as an emulsifying agent to keep fats from separating, to caramel to prevent sucrose crystallization, or in recipes in place of fresh lemon juice. Citric acid is used with sodium bicarbonate in a wide range of effervescent formulae, both for ingestion (e.g., powders and tablets) and for personal care (e.g., bath salts, bath bombs, and cleaning of grease). Citric acid sold in a dry powdered form is commonly sold in markets and groceries as ``sour salt'', due to its physical resemblance to table salt. It has use in culinary applications, as an alternative to vinegar or lemon juice, where a pure acid is needed. Question: can lemon juice be used as citric acid?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " There is no information in the passage to determine if lemon juice can be used as citric acid, therefore the score is: 50"
    },
    {
        "question": "Passage: Robert Westbrook adapted the screenplay to novel form, which was published by Alex in May 2002. Question: was the movie insomnia based on a book?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: A fair catch is a feature of American football and several other codes of football, in which a player attempting to catch a ball kicked by the opposing team -- either on a kickoff or punt -- is entitled to catch the ball without interference from any member of the kicking team. A ball caught in this manner becomes dead once caught, i.e., the player catching the ball is not entitled to run with the ball in an attempt to gain yardage, and the receiving team begins its drive at the spot where the ball was caught. A player wishing to make a fair catch signals his intent by extending one arm above his head and waving it while the kicked ball is in flight. The kicking team must allow the player an opportunity to make the catch without interference. Question: can you call fair catch on a kickoff?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " There is a clear mention in the passage that a fair catch can be called on a kickoff. The student's answer is incorrect.\\n\\nTherefore the score is: 0"
    },
    {
        "question": "Passage: An unlisted public company is a public company that is not listed on any stock exchange. Though the criteria vary somewhat between jurisdictions, a public company is a company that is registered as such and generally has a minimum share capital and a minimum number of shareholders. Each stock exchange has its own listing requirements which a company (or other entity) wishing to be listed must meet. Besides not qualifying to be listed, a public company may choose not to be listed on a stock exchange for a number of reasons, including because it is too small to qualify for a stock exchange listing, does not seek public investors, or there are too few shareholders for a listing. There is a cost to the listed entities, in the listing process and ongoing costs as well as in compliance costs such as the maintenance of a company register. Question: are all public companies listed on the stock exchange?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The yellow-billed cuckoo (Coccyzus americanus) is a cuckoo. Common folk-names for this bird in the southern United States are rain crow and storm crow. These likely refer to the bird's habit of calling on hot days, often presaging rain or thunderstorms. Question: is there a bird called a rain crow?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: The Commonwealth was first officially formed in 1931 when the Statute of Westminster gave legal recognition to the sovereignty of dominions. Known as the ``British Commonwealth'', the original members were the United Kingdom, Canada, Australia, New Zealand, South Africa, Irish Free State, and Newfoundland, although Australia and New Zealand did not adopt the statute until 1942 and 1947 respectively. In 1949, the London Declaration was signed and marked the birth of the modern Commonwealth and the adoption of its present name. The newest member is Rwanda, which joined on 29 November 2009. The most recent departure was the Maldives, which severed its connection with the Commonwealth on 13 October 2016. Question: is canada part of the commonwealth of england?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: A California roll or California maki is a makizushi sushi roll, usually made inside-out, containing cucumber, crab meat or imitation crab, and avocado. Sometimes crab salad is substituted for the crab stick, and often the outer layer of rice in an inside-out roll (uramaki) is sprinkled with toasted sesame seeds, tobiko or masago (capelin roe). Question: does a california roll have fish in it?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The Juris Doctor degree (J.D. or JD), also known as the Doctor of Jurisprudence degree (J.D., JD, D.Jur. or DJur), is a graduate-entry professional degree in law and one of several Doctor of Law degrees. It is earned by completing law school in Australia, Canada and the United States, and some other common law countries. It has the academic standing of a professional doctorate in the United States, a master's degree in Australia, and a second-entry, baccalaureate degree in Canada, (in all three jurisdictions the same as other professional degrees such as M.D. or D.D.S., the degrees required to be a practicing physician or dentist, respectively). Question: is a jd the same as a doctorate?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: In the 1930s, the show began hiring professionals and expanded to four hours. Broadcasting by then at 50,000 watts, WSM made the program a Saturday night musical tradition in nearly 30 states. In 1939, it debuted nationally on NBC Radio. The Opry moved to a permanent home, the Ryman Auditorium, in 1943. As it developed in importance, so did the city of Nashville, which became America's ``country music capital.'' The Grand Ole Opry holds such significance in Nashville that its name is included on the city/county line signs on all major roadways. The signs read ``Music City Metropolitan Nashville Davidson County Home of the Grand Ole Opry.'' Question: are the ryman and grand ole opry the same thing?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The game takes place in the same fictional world as the comic, with events occurring shortly after the onset of the zombie apocalypse in Georgia. However, most of the characters are original to the game, which centers on university professor and convicted criminal Lee Everett, who helps to rescue and subsequently care for a young girl named Clementine. Kirkman provided oversight for the game's story to ensure it corresponded to the tone of the comic, but allowed Telltale to handle the bulk of the developmental work and story specifics. Some characters from the original comic book series also make in-game appearances. Question: is the walking dead game the same as the show?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: On June 1, 2017, it was announced that Seyfried would return as Sophie. Later that month, Dominic Cooper confirmed that he would return for the sequel, along with Streep, Firth and Brosnan as Sky, Donna, Harry, and Sam, respectively. In July 2017, Baranski was also confirmed to return as Tanya. On July 12, 2017, Lily James was cast to play the role of young Donna. On August 3, 2017, Jeremy Irvine and Alexa Davies were also cast in the film, with Irvine playing Brosnan's character Sam in a past era, and Hugh Skinner to play Young Harry, Davies as a young Rosie, played by Julie Walters. On August 16, 2017, it was announced that Jessica Keenan Wynn had been cast as a young Tanya, who is played by Baranski. Julie Walters and Stellan Skarsg\u00e5rd also reprised their roles as Rosie and Bill, respectively. On October 16, 2017, it was announced that singer and actress Cher had joined the cast, in her first on-screen film role since 2010, and her first film with Streep since Silkwood. Question: is the cast of mama mia the same?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: UPC (technically refers to UPC-A) consists of 12 numeric digits, that are uniquely assigned to each trade item. Along with the related EAN barcode, the UPC is the barcode mainly used for scanning of trade items at the point of sale, per GS1 specifications. UPC data structures are a component of GTINs and follow the global GS1 specification, which is based on international standards. But some retailers (clothing, furniture) do not use the GS1 system (rather other barcode symbologies or article number systems). On the other hand, some retailers use the EAN/UPC barcode symbology, but without using a GTIN (for products sold in their own stores only). Question: is a upc code the same as a barcode?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: College football players who are considering entering the NFL draft but who still have eligibility to play football can request an expert opinion from the NFL-created Draft Advisory Board. The Board, composed of scouting experts and team executives, makes a prediction as to the likely round in which a player would be drafted. This information, which has proven to be fairly accurate, can help college players determine whether to enter the draft or to continue playing and improving at the college level. There are also many famous reporting scouts, such as Mel Kiper Jr. Question: do college football players have to enter the draft?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: A new trophy design was created for the 1977 NBA Finals, although it retained the Walter A. Brown title. Unlike the original championship trophy, the new trophy was given permanently to the winning team and a new one was made every year. Question: is a new nba trophy made every year?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: A perfect game is defined by Major League Baseball as a game in which a pitcher (or combination of pitchers) pitches a victory that lasts a minimum of nine innings in which no opposing player reaches base. Thus, the pitcher (or pitchers) cannot allow any hits, walks, hit batsmen, or any opposing player to reach base safely for any other reason and the fielders cannot make an error that allows an opposing player to reach a base; in short, ``27 up, 27 down.'' The feat has been achieved 23 times in MLB history -- 21 times since the modern era began in 1900, most recently by F\u00e9lix Hern\u00e1ndez of the Seattle Mariners on August 15, 2012. A perfect game is also a no-hitter and a shutout. A fielding error that does not allow a batter to reach base, such as a misplayed foul ball, does not spoil a perfect game. Weather-shortened contests in which a team has no baserunners and games in which a team reaches first base only in extra innings do not qualify as perfect games under the present definition. Question: can there be an error in a perfect game?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: Durant was selected as the second overall pick in the 2007 NBA draft by the Seattle SuperSonics. In his first regular season game, the 19-year-old Durant registered 18 points, 5 rebounds, and 3 steals against the Denver Nuggets. On November 16, he made the first game-winning shot of his career in a game against the Atlanta Hawks. At the conclusion of the season, he was named the NBA Rookie of the Year behind averages of 20.3 points, 4.4 rebounds, and 2.4 assists per game. He joined Carmelo Anthony and LeBron James as the only teenagers in league history to average at least 20 points per game over an entire season. Question: did kevin durant play for the seattle supersonics?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 0"
    },
    {
        "question": "Passage: Senate cloture rules historically required a two-thirds affirmative vote to advance nominations to a vote; this was changed to a three-fifths supermajority in 1975. In November 2013, the then-Democratic Senate majority eliminated the filibuster for executive branch nominees and judicial nominees except for Supreme Court nominees by invoking the so called nuclear option. In April 2017, the Republican Senate majority applied the nuclear option to Supreme Court nominations as well, enabling the nominations of Trump nominees Neil Gorsuch and Brett Kavanaugh to proceed to a vote. Question: can a filibuster stop a supreme court nominee?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: JPMorgan Chase Bank, N.A., doing business as Chase Bank, is a national bank headquartered in Manhattan, New York City, that constitutes the consumer and commercial banking subsidiary of the U.S. multinational banking and financial services holding company, JPMorgan Chase & Co. The bank was known as Chase Manhattan Bank until it merged with J.P. Morgan & Co. in 2000. Chase Manhattan Bank was formed by the merger of the Chase National Bank and The Manhattan Company in 1955. The bank has been headquartered in Columbus, Ohio since its merger with Bank One Corporation in 2004. The bank acquired the deposits and most assets of Washington Mutual. Question: is jpmorgan chase the same as chase bank?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: A batsman may not be given out bowled, leg before wicket, caught, stumped or hit wicket off a no-ball. A batsman may be given out run out, hit the ball twice, or obstructing the field. Thus the call of no-ball protects the batsman against losing his wicket in ways that are attributed to the bowler, but not in ways that are attributed to running, or to the batsman's own conduct. Question: can a batsman be run out on a no ball?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: ``Mine Would Be You'' is a song written by Jessi Alexander, Connie Harrington, and Deric Ruttan and recorded by American country music artist Blake Shelton. It was released in July 2013 as the third single from Shelton's 2013 album Based on a True Story.... The song reached number one the US Billboard Country Airplay chart. Question: did blake shelton wrote mine would be you?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The Aquarium's notable specimens include whale sharks, beluga whales, California sea lions, bottlenose dolphins, and manta rays. Question: does the georgia aquarium still have whale sharks?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The likelihood that a family will use a payday loan increases if they are unbanked or underbanked, or lack access to a traditional deposit bank account. In an American context the families who will use a payday loan are disproportionately either of black or Hispanic descent, recent immigrants, and/or under-educated. These individuals are least able to secure normal, lower-interest-rate forms of credit. Since payday lending operations charge higher interest-rates than traditional banks, they have the effect of depleting the assets of low-income communities. The Insight Center, a consumer advocacy group, reported in 2013 that payday lending cost U.S communities $774 million a year. Question: is a payday lender a type of bank?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Area codes in the North American Numbering Plan area may not contain 0 or 1 as the first digit. Question: are there any area codes that start with 1?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " There is a mistake in the student's answer. As per the passage, area codes in the North American Numbering Plan area may not contain 0 or 1 as the first digit. So, there cannot be any area codes that start with 1. Therefore the score is: 50"
    },
    {
        "question": "Passage: A round is over when either one player plays the last domino in their hand or no players can make a legal play. The latter situation can occur if someone plays a double that no longer has three remaining free dominoes to play on it and the boneyard is exhausted. Question: can you go out on a double in chicken foot?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: In 1962, the first appearance of a space-faring Robinson family occurred in a comic book published by Gold Key Comics. The Space Family Robinson, who were scientists aboard Earth's ``Space Station One'', are swept away in a cosmic storm in the comic's second issue. These Robinsons were scientist father Craig, scientist mother June, early teens Tim (son) and Tam (daughter), along with pets Clancy (dog) and Yakker (parrot). Space Station One also boasted two spacemobiles for ship-to-planet travel. Question: is lost in space based on a book?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The Battle of the Alamo (February 23 -- March 6, 1836) was a pivotal event in the Texas Revolution. Following a 13-day siege, Mexican troops under President General Antonio L\u00f3pez de Santa Anna launched an assault on the Alamo Mission near San Antonio de B\u00e9xar (modern-day San Antonio, Texas, United States), killing the Texian defenders. Santa Anna's cruelty during the battle inspired many Texians--both Texas settlers and adventurers from the United States--to join the Texian Army. Buoyed by a desire for revenge, the Texians defeated the Mexican Army at the Battle of San Jacinto, on April 21, 1836, ending the revolution. Question: was the battle of the alamo part of the mexican american war?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The Bill of Rights is the first ten amendments to the United States Constitution. Proposed following the often bitter 1787--88 battle over ratification of the U.S. Constitution, and crafted to address the objections raised by Anti-Federalists, the Bill of Rights amendments add to the Constitution specific guarantees of personal freedoms and rights, clear limitations on the government's power in judicial and other proceedings, and explicit declarations that all powers not specifically delegated to Congress by the Constitution are reserved for the states or the people. The concepts codified in these amendments are built upon those found in several earlier documents, including the Virginia Declaration of Rights and the English Bill of Rights, along with earlier documents such as Magna Carta (1215). In practice, the amendments had little impact on judgments by the courts for the first 150 years after ratification. Question: was the bill of rights in the constitution?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: While mentioned in passing throughout later seasons, Burke officially returns in the tenth season in order to conclude Cristina Yang's departure from the series. Question: does dr burke come back after season 3?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: A spark plug (sometimes, in British English, a sparking plug, and, colloquially, a plug) is a device for delivering electric current from an ignition system to the combustion chamber of a spark-ignition engine to ignite the compressed fuel/air mixture by an electric spark, while containing combustion pressure within the engine. A spark plug has a metal threaded shell, electrically isolated from a central electrode by a porcelain insulator. The central electrode, which may contain a resistor, is connected by a heavily insulated wire to the output terminal of an ignition coil or magneto. The spark plug's metal shell is screwed into the engine's cylinder head and thus electrically grounded. The central electrode protrudes through the porcelain insulator into the combustion chamber, forming one or more spark gaps between the inner end of the central electrode and usually one or more protuberances or structures attached to the inner end of the threaded shell and designated the side, earth, or ground electrode(s). Question: does a spark plug keep an engine running?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: The southernmost flight route with plausible airports would be between Buenos Aires and Perth. With a 175\u00b0 (S) heading, the route's great circle exceeds 85 \u00b0S and would be within 500 kilometres (270 nmi) from the South Pole. Currently, no commercial airliners operates this 6,800 nautical miles (12,600 km) route. However, in February, 2018, it was stated that Norwegian Air Argentina is considering this ``less than 15 hours'' trans-polar flight between South America and Asia, with a stop-over in Perth enroute Singapore. They will not fly over the South Pole, but around Antarctica taking advantage of the strong winds which circle that continent in an easterly direction. Hence, the ``westbound'' flight from Buenos Aires would actually travel south-east south of Cape Town, over the southern Indian Ocean and on to Perth, while the true ``eastbound'' flight would also head south-east south of Tasmania and New Zealand, over the South Pacific and on to South America. If this route becomes operational, a Buenos Aires - Singapore return flight would possibly be the fastest circumnavigation available with commercial airliners, although Perth - Buenos Aires return would be faster but without passing the Equator. Question: do commercial aircraft fly over the north pole?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: A legislator (or lawmaker) is a person who writes and passes laws, especially someone who is a member of a legislature. Legislators are usually politicians and are often elected by the people of the state. Legislatures may be supra-national (for example, the European Parliament), national (for example, the United States Congress), regional (for example, the National Assembly for Wales), or local (for example, local authorities). Question: is a legislator the same as a senator?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: How to Train Your Dragon: The Hidden World is an upcoming 2019 American 3D computer-animated action fantasy film produced by DreamWorks Animation and distributed by Universal Pictures, loosely based on the book series of the same name by Cressida Cowell. It is a sequel to 2010's How to Train Your Dragon and 2014's How to Train Your Dragon 2, and is the third and final installment in the How to Train Your Dragon trilogy. Question: is there really a how to train your dragon 3?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 0"
    },
    {
        "question": "Passage: A rooster, also known as a gamecock, cockerel or cock, is an adult male gallinaceous bird, usually a male chicken (Gallus gallus domesticus). Question: is a chicken and rooster the same thing?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " There is a difference between a chicken and a rooster. A rooster is an adult male chicken, while a chicken can be either male or female. Therefore the score is: 70"
    },
    {
        "question": "Passage: As of September, 2017, Destination Maternity operates over 1,000 retail locations in North America, including 512 stores, predominantly under the trade-names Motherhood Maternity\u00ae, A Pea in the Pod\u00ae, and Destination Maternity\u00ae, and sells on the web through DestinationMaternity.com, Motherhood.com and APeainthePod.com; Destination Maternity brands are offered at retailers such as Macy's and Boscov's. Question: is motherhood maternity and destination maternity the same?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The United States does not have a nationwide soda tax, but a few of its cities have passed their own tax and the U.S. has seen a growing debate around taxing soda in various cities, states and even in Congress in recent years. A few states impose excise taxes on bottled soft drinks or on wholesalers, manufacturers, or distributors of soft drinks. Question: is there a sugar tax in the us?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: As of 11 March 2018, Cutmore-Scott dons an American accent to play disgraced illusionist/magician-turned-FBI consultant Cameron Black following an illusion that goes horribly wrong in the new ABC murder-mystery series Deception. Cutmore-Scott also portrays Cameron's incarcerated, identical-twin brother Jonathan. Deception began airing the same evening in Canada on CTV. Question: is the guy from deception really a twin?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: In addition to domestic units, industrial dishwashers are available for use in commercial establishments such as hotels and restaurants, where a large number of dishes must be cleaned. Washing is conducted with temperatures of 65--71 \u00b0C (149--160 \u00b0F) and sanitation is achieved by either the use of a booster heater that will provide an 82 \u00b0C (180 \u00b0F) ``final rinse'' temperature or through the use of a chemical sanitizer. Question: does the dishwasher make its own hot water?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Though The Big Cypress is the largest growth of cypress swamps in South Florida, cypress swamps can be found near the Atlantic Coastal Ridge and between Lake Okeechobee and the Eastern flatwoods, as well as in sawgrass marshes. Cypresses are deciduous conifers that are uniquely adapted to thrive in flooded conditions, with buttressed trunks and root projections that protrude out of the water, called ``knees''. Bald cypress trees grow in formations with the tallest and thickest trunks in the center, rooted in the deepest peat. As the peat thins out, cypresses grow smaller and thinner, giving the small forest the appearance of a dome from the outside. They also grow in strands, slightly elevated on a ridge of limestone bordered on either side by sloughs. Other hardwood trees can be found in cypress domes, such as red maple, swamp bay, and pop ash. If cypresses are removed, the hardwoods take over, and the ecosystem is recategorized as a mixed swamp forest. Question: is the everglades the largest swamp in north america?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: In chess, the king (\u2654,\u265a) is the most important piece. The object of the game is to threaten the opponent's king in such a way that escape is not possible (checkmate). If a player's king is threatened with capture, it is said to be in check, and the player must remove the threat of capture on the next move. If this cannot be done, the king is said to be in checkmate, resulting in a loss for that player. Although the king is the most important piece, it is usually the weakest piece in the game until a later phase, the endgame. Players cannot make any move that places their own king in check. Question: can you take out the king in chess?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The Thirteenth Amendment (Amendment XIII) to the United States Constitution abolished slavery and involuntary servitude, except as punishment for a crime. In Congress, it was passed by the Senate on April 8, 1864, and by the House on January 31, 1865. The amendment was ratified by the required number of states on December 6, 1865. On December 18, 1865, Secretary of State William H. Seward proclaimed its adoption. It was the first of the three Reconstruction Amendments adopted following the American Civil War. Question: was the 13th amendment after the civil war?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 0"
    },
    {
        "question": "Passage: Each legislator shall be at least twenty-one years of age, an elector and resident of the District from which elected and shall have resided in the state for a period of two years prior to election. Question: does the florida constitution give a minimum age for legislators?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The Train is a 1964 war film directed by John Frankenheimer from a story and screenplay by Franklin Coen and Frank Davis, inspired by the non-fiction book Le front de l'art by Rose Valland, who documented the works of art placed in storage that had been looted by the Germans from museums and private art collections. It stars Burt Lancaster, Paul Scofield and Jeanne Moreau. Question: is the movie the train a true story?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The crisis was documented by photographers, musicians, and authors, many hired during the Great Depression by the federal government. For instance, the Farm Security Administration hired numerous photographers to document the crisis. Artists such as Dorothea Lange were aided by having salaried work during the Depression. She captured what have become classic images of the dust storms and migrant families. Among her most well-known photographs is Destitute Pea Pickers in California. Mother of Seven Children, which depicted a gaunt-looking woman, Florence Owens Thompson, holding three of her children. This picture expressed the struggles of people caught by the Dust Bowl and raised awareness in other parts of the country of its reach and human cost. Decades later, Thompson disliked the boundless circulation of the photo and resented the fact she did not receive any money from its broadcast. Thompson felt it gave her the perception as a Dust Bowl ``Okie.'' Question: did the dust bowl happen during the great depression?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Following the success of the .17 HMR, the .17 Hornady Mach 2 was introduced in early 2004. The .17 HM2 is based on the .22 LR (slightly longer in case dimensions) case necked down to .17 caliber using the same bullet as the HMR but at a velocity of approximately 2,100 feet per second (640 m/s) in the 17-grain (1.1 g) polymer tip loading. Question: is a 17 hmr bigger than a 22lr?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Additionally, although fees for debit card ATM usage are very rare in countries such as the UK, where cashback originated , this is not the case in some other countries. In Canada and the United States, fees of $1~2 are typical when using an ATM from a different bank than the one with which the customer has an account. The fees in some other countries are even higher. In Germany, for instance, usual fees are \u20ac4~5 when using an ATM of another bank network than the one of his bank. This gives rise to another potential cashback advantage for the consumer: by making use of the cashback procedure, this ATM fee can be avoided for the cardholder. Question: does it cost money to get cash back?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Christopher Robert Evans (born June 13, 1981) is an American actor. Evans is known for his superhero roles as the Marvel Comics characters Captain America in the Marvel Cinematic Universe and Human Torch in Fantastic Four (2005) and its 2007 sequel. Question: is the human torch the same guy as captain america?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: An athlete who is awarded a letter (or letters in multiple sports) is said to have ``lettered'' when they receive their letter. Question: can you get more than one varsity letter?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The Hatch Act of 1939, officially An Act to Prevent Pernicious Political Activities, is a United States federal law whose main provision prohibits employees in the executive branch of the federal government, except the president, vice-president, and certain designated high-level officials, from engaging in some forms of political activity. It went into law on August 2, 1939. The law was named for Senator Carl Hatch of New Mexico. It was most recently amended in 2012. Question: does the hatch act apply to elected officials?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: Access courses are generally tailored as pathways; that is, they prepare students with the necessary skills and imbue the appropriate knowledge required for a specific undergraduate career. For example, there are 'access to law', 'access to medicine' and 'access to nursing' pathways that prepare students to study law, medicine and nursing at undergraduate level, respectively. Question: is an access course classed as higher education?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: Pirates of the Caribbean is a dark ride attraction at Disneyland, Magic Kingdom, Tokyo Disneyland, and Disneyland Park in Paris. The original version at Disneyland, which opened in 1967, was the last attraction whose construction was overseen by Walt Disney; he died three months before it opened. The ride, which tells the story of a band of pirates and their troubles and exploits, was replicated at the Magic Kingdom in 1973, at Tokyo Disneyland in 1983, and at Disneyland Paris in 1992. Each of the initial four versions of the ride has a different fa\u00e7ade but a similar ride experience. A reimagined version of the ride, Pirates of the Caribbean: Battle for the Sunken Treasure, opened at the Shanghai Disneyland Park in 2016. Question: did the pirates of the caribbean ride come first?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: In Topeka, I-70 intersects a child route, I-470, twice. The second time it is intersected, the Kansas Turnpike merges, making I-70 into a toll road. This is one of only two sections of I-70 that are tolled (the other is along the Pennsylvania Turnpike), with the maximum toll distance costing $17.50 as of 2016. I-70 carries this designation from Topeka to Bonner Springs. It is the eastern terminus of the turnpike, and from there to 18th Street and extending on to the Kansas eastern border, the highway is free. Question: is interstate 70 in kansas a toll road?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: In the US, semolina (specifically farina) is boiled to produce a porridge; a popular brand of this is Cream of Wheat. Question: is semolina flour the same as cream of wheat?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " There is no clear-cut answer to this question, as both semolina flour and Cream of Wheat are made from different components of the wheat plant. Semolina flour is made from the endosperm of durum wheat, while Cream of Wheat is made from the germ and bran of soft red winter wheat. However, they are often used interchangeably in cooking, and they have a similar appearance and texture when cooked."
    },
    {
        "question": "Passage: The Xbox One gaming console has received updates from Microsoft since its launch in 2013 that enable it to play select games from its two predecessor consoles, Xbox and Xbox 360. On June 15, 2015, backward compatibility with supported Xbox 360 games became available to eligible Xbox Preview program users with a beta update to the Xbox One system software. The dashboard update containing backward compatibility was released publicly on November 12, 2015. On October 24, 2017, another such update added games from the original Xbox library. The following is a list of all backward compatible games on Xbox One under this functionality. Question: can you use xbox 360 games in the xbox one?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: This has led to categorizing trucks similarly, even if their payload is different. Therefore, the Toyota Tacoma, Dodge Dakota, Ford Ranger, Honda Ridgeline, Chevrolet S-10, GMC S-15 and Nissan Frontier are called quarter-tons (1\u20444-ton). The Ford F-150, Chevrolet C10/K10, Chevrolet/GMC 1500, Dodge 1500, Toyota Tundra, and Nissan Titan are half-tons (1\u20442-ton). The Ford F-250, Chevrolet C20/K20, Chevrolet/GMC 2500, and Dodge 2500 are three-quarter-tons (3\u20444-ton). Chevrolet/GMC's 3\u20444-ton suspension systems were further divided into light and heavy-duty, differentiated by 5-lug and 6 or 8-lug wheel hubs depending on year, respectively. The Ford F-350, Chevrolet C30/K30, Chevrolet/GMC 3500, and Dodge 3500 are one tons (1-ton). Question: is a dodge 3500 a 1 ton truck?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: The President's Guest House is one of several residences owned by the United States government for use by the President and Vice President of the United States; other such residences include the White House, Camp David, One Observatory Circle, the Presidential Townhouse, and Trowbridge House. The President's Guest House has been called ``the world's most exclusive hotel'' because it is primarily used to host visiting dignitaries and other guests of the president. It is larger than the White House and closed to the public. Question: do foreign dignitaries stay at the white house?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: In the 1990s, Clark proposes marriage to Lois and reveals his identity as Superman to her. They began a long engagement, which was complicated by the death of Superman, a breakup, and several other problems. The couple finally married in Superman: The Wedding Album (Dec. 1996). Clark and Lois' biological child in DC Comics canon was born in Convergence: Superman #2 (July 2015), a son named Jonathan Samuel Kent, who eventually becomes Superboy. Question: do superman and lois lane end up together?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Scar makes a brief cameo appearance in the film in Simba's nightmare. In the nightmare, Simba runs down the cliff where his father died, attempting to rescue him. Scar intervenes, however, and then turns into Kovu and throws Simba off the cliff. Scar makes another cameo appearance in a pool of water, as a reflection, after Kovu is exiled from Pride Rock. Question: is scar alive in the lion king 2?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Retail sale of beer and wine is prohibited on Sundays between 2:00 a.m. and 1:00 p.m. and between 2:00 a.m. and 7:00 a.m. on weekdays and Saturdays. Retail sale of liquor is prohibited on Sundays, Christmas Day, and between 12:00 midnight and 8:00 a.m on all other days. Question: can i buy liquor on sunday in wv?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: At the Liberation of France in the summer of 1944, Metropolitan France kept GMT+2 as it was the time then used by the Allies (British Double Summer Time). In the winter of 1944--1945, Metropolitan France switched to GMT+1, same as in the United Kingdom, and switched again to GMT+2 in April 1945 like its British ally. In September 1945, Metropolitan France returned to GMT+1 (pre-war summer time), which the British had already done in July 1945. Metropolitan France was officially scheduled to return to GMT+0 on November 18, 1945 (the British returned to GMT+0 in on October 7, 1945), but the French government canceled the decision on November 5, 1945, and GMT+1 has since then remained the official time of Metropolitan France. Question: is france the same timezone as the uk?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: The first commercially available device that could be properly referred to as a ``smartphone'' began as a prototype called ``Angler'' developed by Frank Canova in 1992 while at IBM and demonstrated in November of that year at the COMDEX computer industry trade show. A refined version was marketed to consumers in 1994 by BellSouth under the name Simon Personal Communicator. In addition to placing and receiving cellular calls, the touchscreen-equipped Simon could send and receive faxes and emails. It included an address book, calendar, appointment scheduler, calculator, world time clock, and notepad, as well as other visionary mobile applications such as maps, stock reports and news. The term ``smart phone'' or ``smartphone'' was not coined until a year after the introduction of the Simon, appearing in print as early as 1995, describing AT&T's PhoneWriter Communicator. Question: was the iphone the first touch screen phone?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: A person who has indefinite leave to remain, the right of abode or Irish citizenship has settled status if resident in the United Kingdom (all full British citizens have the right of abode). Question: is right of abode the same as indefinite leave to remain?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Three species of maple trees are predominantly used to produce maple syrup: the sugar maple (Acer saccharum), the black maple (A. nigrum), and the red maple (A. rubrum), because of the high sugar content (roughly two to five percent) in the sap of these species. The black maple is included as a subspecies or variety in a more broadly viewed concept of A. saccharum, the sugar maple, by some botanists. Of these, the red maple has a shorter season because it buds earlier than sugar and black maples, which alters the flavour of the sap. Question: does maple syrup come out of the tree sweet?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 40"
    },
    {
        "question": "Passage: The Terror is an American anthology horror drama television series that premiered on AMC on March 25, 2018. The series is named after Dan Simmons' 2007 best-selling novel, a fictionalized account of Captain Sir John Franklin's lost expedition to the Arctic in 1845--1848, which serves as the basis for the series' first season. On June 22, 2018, it was announced that AMC had renewed the series for a ten-episode second season set to premiere in 2019. Question: is there only 1 season of the terror?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The term remote keyless system (RKS), also called keyless entry or remote central locking, refers to a lock that uses an electronic remote control as a key which is activated by a handheld device or automatically by proximity. Question: is keyless entry the same as remote start?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Cannabis in Connecticut is illegal for recreational use, but possession of small amounts is decriminalized. Medical usage is permitted. Question: is it illegal to smoke weed in ct?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Some legal scholars have argued that because countries have constantly invoked the Declaration for more than 50 years, it has become binding as a part of customary international law. However, in the United States, the Supreme Court in Sosa v. Alvarez-Machain (2004), concluded that the Declaration ``does not of its own force impose obligations as a matter of international law.'' Courts of other countries have also concluded that the Declaration is not in and of itself part of domestic law. Question: does the us follow the universal declaration of human rights?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Destin--Fort Walton Beach Airport (IATA: VPS, ICAO: KVPS, FAA LID: VPS) is an airport located within Eglin Air Force Base, near Destin and Fort Walton Beach in Okaloosa County, Florida. No private aircraft are allowed, so Destin Executive Airport is used instead for non-commercial operations by general aviation and business aircraft. The airport was previously named Northwest Florida Regional Airport until February 17, 2015 and Okaloosa Regional Airport until September 2008. Question: is fort walton beach the same as destin?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: Most seat belt laws in the United States are left to the states. However, the first seat belt law was a federal law, Title 49 of the United States Code, Chapter 301, Motor Vehicle Safety Standard, which took effect on January 1, 1968, that required all vehicles (except buses) to be fitted with seat belts in all designated seating positions. This law has since been modified to require three-point seat belts in outboard-seating positions, and finally three-point seat belts in all seating positions. Initially, seat belt use was voluntary. New York was the first state to pass a law which required vehicle occupants to wear seat belts, a law that came into effect on December 1, 1984. Officer Nicholas Cimmino of the Westchester County Department of Public Safety wrote the nation's first ticket for such violation. New Hampshire is the only state that has no enforceable laws for the wearing of seat belts in a vehicle. Question: are there any states that do not have a seat belt law?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: With eleven films released, the X-Men film series is the sixth highest-grossing film series, having grossed over US$5.7 billion worldwide. It is set to continue with the releases of Dark Phoenix and The New Mutants in 2019. Question: is there a new x-men movie coming out?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: Second-year player Michael Jordan put on a record-setting performance in Game 2 of the Bulls' first-round series against the Celtics, scoring 63 points in a 2-OT loss, which surpassed Elgin Baylor's 61-point performance from the 1962 NBA Finals and still stands as the NBA Playoff scoring record. Jordan averaged 43.7 points per game in the series, but was unable to prevent the Bulls from being swept by a more experienced, more talented Celtics team. The Bulls set a dubious mark by posting the second worst record for a playoff-qualifying team in history, going just 30--52 during the season. Game 2, where the record was set, was ranked by TV Guide as the 26th Most Memorable Moment in Television History, and is credited with boosting the NBA's popularity surge and eventual rise to near the top of the United States television sports market, trailing only football by the mid-90s. Question: did the bulls get swept by the celtics?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Brake fluid is a type of hydraulic fluid used in hydraulic brake and hydraulic clutch applications in automobiles, motorcycles, light trucks, and some bicycles. It is used to transfer force into pressure, and to amplify braking force. It works because liquids are not appreciably compressible. Question: can i use hydraulic fluid for brake fluid?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " There is a difference between hydraulic fluid and brake fluid, although they are both used in hydraulic systems. Hydraulic fluid is a general term used to describe any fluid used in hydraulic systems, while brake fluid is a specific type of hydraulic fluid designed specifically for braking systems.\\n\\nBrake fluid is specially designed to withstand high temperatures and pressures generated during braking and to provide consistent performance over a wide range of temperatures. It is also designed to resist moisture absorption, which can lead to brake system corrosion and performance issues.\\n\\nIn summary, while hydraulic fluid can be used in a brake system, it is not recommended due to its lower boiling point and less resistance to moisture absorption compared to brake fluid. Therefore, the score is: 80/100."
    },
    {
        "question": "Passage: An AC adapter, AC/DC adapter, or AC/DC converter is a type of external power supply, often enclosed in a case similar to an AC plug. Other common names include plug pack, plug-in adapter, adapter block, domestic mains adapter, line power adapter, wall wart, power brick, and power adapter. Adapters for battery-powered equipment may be described as chargers or rechargers (see also battery charger). AC adapters are used with electrical devices that require power but do not contain internal components to derive the required voltage and power from mains power. The internal circuitry of an external power supply is very similar to the design that would be used for a built-in or internal supply. Question: is a power adapter the same as a charger?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 60"
    },
    {
        "question": "Passage: DMSI applied the brand to a new online and catalog-based retailing operation, with no physical stores, headquartered in Cedar Rapids, Iowa. DMSI then began operating under the Montgomery Ward branding and managed to get it up and running in three months. The new firm began operations in June 2004, selling essentially the same categories of products as the former brand, but as a new, smaller catalog. Question: are there any montgomery ward stores still open?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: After the defeat in the 2016 Olympics, the USWNT underwent a year of experimentation which saw them losing 3 home games. If not for a comeback win against Brazil, the USWNT was on the brink of losing 4 home games in one year, a low never before seen by the USWNT. 2017 saw the USWNT play 12 games against teams ranked in the top-15 in the world. The USWNT heads into World Cup Qualifying in fall of 2018. Question: is the us womens soccer team in the world cup?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Canadian citizenship is typically obtained by birth in Canada on the principle of jus soli, or birth abroad when at least one parent is a Canadian citizen or by adoption by at least one Canadian citizen under the rules of jus sanguinis. It can also be granted to a permanent resident who has lived in Canada for a period of time through naturalization. Immigration, Refugees and Citizenship Canada (IRCC, formerly known as Citizenship and Immigration Canada, or CIC) is the department of the federal government responsible for citizenship-related matters, including confirmation, grant, renunciation and revocation of citizenship. Question: can a child born in canada get citizenship?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The Capitals were founded in 1974 as an expansion franchise, alongside the Kansas City Scouts. Since purchasing the team in 1999, Leonsis revitalized the franchise by drafting star players such as Alexander Ovechkin, Nicklas Backstrom, Mike Green and Braden Holtby. The 2009--10 Capitals won the franchise's first-ever Presidents' Trophy for being the team with the most points at the end of the regular season. They won it a second time in 2015--16, and did so for a third time the following season in 2016--17. In addition to eleven division titles and three Presidents' Trophies, the Capitals have reached the Stanley Cup Finals twice (in 1998 and 2018), winning in 2018. Question: have the capitals ever win the stanley cup?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 0"
    },
    {
        "question": "Passage: Depending on the context, a dependent variable is sometimes called a ``response variable'', ``regressand'', ``criterion'', ``predicted variable'', ``measured variable'', ``explained variable'', ``experimental variable'', ``responding variable'', ``outcome variable'', ``output variable'' or ``label''. Question: is outcome variable the same as dependent variable?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Lynyrd Skynyrd is a Southern rock band from Jacksonville, Florida. Formed in 1964, the group originally included vocalist Ronnie Van Zant, guitarists Gary Rossington and Allen Collins, bassist Larry Junstrom and drummer Bob Burns. The current lineup features Rossington, guitarist and vocalist Rickey Medlocke (from 1971 to 1972, and since 1996), lead vocalist Johnny Van Zant (since 1987), drummer Michael Cartellone (since 1999), guitarist Mark Matejka (since 2006), keyboardist Peter Keys (since 2009) and bassist Keith Christopher (since 2017). The band also tours with two backing vocalists, currently Dale Krantz-Rossington (since 1987) and Carol Chase (since 1996). Question: is there any original members of lynyrd skynyrd?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The character appears in various Marvel Cinematic Universe films, including The Avengers (2012), portrayed by Damion Poitier, and Guardians of the Galaxy (2014), Avengers: Age of Ultron (2015), Avengers: Infinity War (2018), and the fourth Avengers film (2019), portrayed by Josh Brolin through voice and motion capture. The character has appeared in various comic adaptations, including animated television series, arcade, and video games. Question: was thanos in the first guardians of the galaxy?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Brown rice is whole-grain rice with the inedible outer hull removed; white rice is the same grain with the hull, bran layer, and cereal germ removed. Red rice, gold rice, and black rice (also called purple rice) are all whole rices, but with differently pigmented outer layers. Question: is whole wheat rice the same as brown rice?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": "NILL"
    },
    {
        "question": "Passage: Proxy voting is automatically prohibited in organizations that have adopted Robert's Rules of Order Newly Revised (RONR) or The Standard Code of Parliamentary Procedure (TSC) as their parliamentary authority, unless it is provided for in its bylaws or charter or required by the laws of its state of incorporation. Robert's Rules says, ``If the law under which an organization is incorporated allows proxy voting to be prohibited by a provision of the bylaws, the adoption of this book as parliamentary authority by prescription in the bylaws should be treated as sufficient provision to accomplish that result''. Demeter says the same thing, but also states that ``if these laws do not prohibit voting by proxy, the body can pass a law permitting proxy voting for any purpose desired.'' RONR opines, ``Ordinarily it should neither be allowed nor required, because proxy voting is incompatible with the essential characteristics of a deliberative assembly in which membership is individual, personal, and nontransferable. In a stock corporation, on the other hand, where the ownership is transferable, the voice and vote of the member also is transferable, by use of a proxy.'' While Riddick opines that ``proxy voting properly belongs in incorporate organizations that deal with stocks or real estate, and in certain political organizations,'' it also states, ``If a state empowers an incorporated organization to use proxy voting, that right cannot be denied in the bylaws.'' Riddick further opines, ``Proxy voting is not recommended for ordinary use. It can discourage attendance, and transfers an inalienable right to another without positive assurance that the vote has not been manipulated.'' Question: does robert's rules of order allow proxy voting?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Persons driving into Canada must have their vehicle's registration document and proof of insurance. Question: can u drive in canada with us license?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Colocasia esculenta is thought to be native to Southern India and Southeast Asia, but is widely naturalised. It is a perennial, tropical plant primarily grown as a root vegetable for its edible starchy corm, and as a leaf vegetable. It is a food staple in African, Oceanic and Indian cultures and is believed to have been one of the earliest cultivated plants. Colocasia is thought to have originated in the Indomalaya ecozone, perhaps in East India, Nepal, and Bangladesh, and spread by cultivation eastward into Southeast Asia, East Asia and the Pacific Islands; westward to Egypt and the eastern Mediterranean Basin; and then southward and westward from there into East Africa and West Africa, where it spread to the Caribbean and Americas. It is known by many local names and often referred to as ``elephant ears'' when grown as an ornamental plant. At around 3.3 million metric tons per year, Nigeria is the largest producer of taro in the world. Question: is taro root the same as elephant ears?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The statue was administered by the United States Lighthouse Board until 1901 and then by the Department of War; since 1933 it has been maintained by the National Park Service. Public access to the balcony around the torch has been barred for safety since 1916. Question: does the us own the statue of liberty?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Triamcinolone acetonide as an intra-articular injectable has been used to treat a variety of musculoskeletal conditions. When applied as a topical ointment, applied to the skin, it is used to mitigate blistering from poison ivy, oak, and sumac, . When combined with Nystatin, it is used to treat skin infections with discomfort from fungus, though it should not be used on the eyes, mouth, or genital area. It provides relatively immediate relief and is used before using oral prednisone. Oral and dental paste preparations are used for treating aphthous ulcers. Question: can triamcinolone acetonide cream be used to treat poison ivy?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The World Cup is a gold trophy that is awarded to the winners of the FIFA World Cup association football tournament. Since the advent of the World Cup in 1930, two trophies have been used: the Jules Rimet Trophy from 1930 to 1970, and the FIFA World Cup Trophy from 1974 to the present day. Question: is it the same world cup trophy every year?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Charles B. McVay III (July 30, 1898 -- November 6, 1968) was an American naval officer and the commanding officer of USS Indianapolis (CA-35) when it was lost in action in 1945, resulting in a massive loss of life. Of all captains in the history of the United States Navy, he is the only one to have been subjected to court-martial for losing a ship sunk by an act of war, despite the fact that he was on a top secret mission maintaining radio silence (the testimony of the Japanese commander who sank his ship also seemed to exonerate McVay). After years of mental health problems, he committed suicide. Following years of efforts by some survivors and others to clear his name, McVay was posthumously exonerated by the 106th United States Congress and President Bill Clinton on October 30, 2000. Question: did the captain of the uss indianapolis live?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Sebastian Philip Bierk (born April 3, 1968), known professionally as Sebastian Bach, is a Canadian heavy metal singer who achieved mainstream success as frontman of Skid Row from 1987 to 1996. He continues a solo career, acted on Broadway, and has made appearances in film and television. Question: is the lead singer of skid row dead?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Damon Albarn OBE (/\u02c8de\u026am\u0259n \u02c8\u00e6lb\u0251\u02d0rn/; born 23 March 1968) is an English musician, singer, songwriter and record producer. He is best known as the lead singer of the British rock band Blur as well as the co-founder, lead vocalist, instrumentalist, and principal songwriter of the virtual band Gorillaz. Question: is the singer from blur in the gorillaz?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: All branches of the U.S. Military currently prohibit beards for a vast majority of recruits, although some mustaches are still allowed, based on policies that were initiated during the period of World War I. Question: can i have a beard in the military?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Leatherface is a 2017 American horror film directed by Julien Maury and Alexandre Bustillo, written by Seth M. Sherwood, and starring Stephen Dorff, Vanessa Grasse, Sam Strike, and Lili Taylor. It is the eighth film in the Texas Chainsaw Massacre franchise (TCM), and works as a prequel to 1974's The Texas Chain Saw Massacre, explaining the origin of the series' lead character. Question: is leatherface in texas chainsaw massacre the same?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Croatia national football team have appeared in the FIFA World Cup on five occasions (in 1998, 2002, 2006, 2014 and 2018) since gaining independence in 1991. Before that, from 1930 to 1990 Croatia was part of Yugoslavia. For World Cup records and appearances in that period, see Yugoslavia national football team and Serbia at the FIFA World Cup. Their best result thus far was silver position at the 2018 final, where they lost 4-2 to France. Question: has croatia ever been in a world cup final?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Corinthian leather is a term coined by the advertising agency Bozell to describe the upholstery used in certain Chrysler luxury vehicles. The term first appeared in advertising in 1974. Although the term suggests that the product has a relationship to or origination from Corinth, there is no relationship; the term is merely a marketing concept. Question: is there such a thing as corinthian leather?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: The Indiana Toll Road, officially the Indiana East--West Toll Road, is a toll road that runs for 156.28 miles (251.51 km) east--west across northern Indiana from the Illinois state line to the Ohio state line. It has been advertised as the ``Main Street of the Midwest''. The entire toll road is designated as part of Interstate 90, and the segment from Lake Station east to the Ohio state line is a concurrency with Interstate 80. The toll road is owned by the Indiana Finance Authority and operated by the Indiana Toll Road Concession Company, which is owned by IFM Investors. Question: is i 80 in indiana a toll road?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: A split-phase or single-phase three-wire system is a type of single-phase electric power distribution. It is the AC equivalent of the original Edison three-wire direct-current system. Its primary advantage is that it saves conductor material over a single-ended single-phase system, while only requiring a single phase on the supply side of the distribution transformer. Question: is split phase the same as single phase?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The Fate of the Furious premiered in Berlin on April 4, 2017, and was theatrically released in the United States on April 14, 2017, playing in 3D, IMAX 3D and 4DX internationally. The film received mixed reviews from critics, many of whom praised the action sequences and acting performances but criticized the storyline. The film grossed over $1.2 billion worldwide, making it the thirtieth film (and the second in the franchise, after Furious 7) to gross over $1 billion, the third-highest-grossing film of 2017 and the fifteenth-highest-grossing film of all time. The film grossed $542 million worldwide during its opening weekend, which is the second highest-grossing worldwide opening of all time behind Avengers: Infinity War (2018). A spinoff film starring Johnson and Statham's characters is scheduled for release in August 2019, while the ninth and tenth films are scheduled for releases on the years 2020 and 2021. Question: is fate and the furious the last movie?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: Citizens of member nations of the Gulf Cooperation Council may travel to Oman without visa limits. Nationals of 71 other countries and territories can apply for visas online which are valid for a period of 30 days. All visitors must hold a passport valid for 6 months. Question: do you need a visa to visit oman?",
        "pred_ans": " Passage: Citizens of member nations of the Gulf Cooperation Council may travel to Oman without visa limits. Nationals of 71 other countries and territories can apply for visas online which are valid for a period of 30 days. All visitors must hold a passport valid for 6 months.\\n\\nResponse: False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The Los Angeles Chargers are a professional American football team based in the Greater Los Angeles Area. The Chargers compete in the National Football League (NFL) as a member club of the league's American Football Conference (AFC) West division. The team was founded on August 14, 1959, and began play on September 10, 1960, as a charter member of the American Football League (AFL), and spent its first season in Los Angeles, before moving to San Diego in 1961 to become the San Diego Chargers. The Chargers joined the NFL as result of the AFL--NFL merger in 1970, and played their home games at SDCCU Stadium. The return of the Chargers to Los Angeles was announced for the 2017 season, just one year after the Rams had moved back to the city from St. Louis. The Chargers will play their home games at the StubHub Center until the opening in 2020 of the Los Angeles Stadium at Hollywood Park, which they will share with the Rams. Question: do the chargers still play in san diego?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 25"
    },
    {
        "question": "Passage: When the equilateral pentagon is dissected into triangles, two of them appear as isosceles (triangles in orange and blue) while the other one is more general (triangle in green). We assume that we are given the adjacent angles \u03b1 (\\displaystyle \\alpha ) and \u03b2 (\\displaystyle \\beta ) . Question: is a pentagon made of 5 equilateral triangles?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The United States Marine Corps (USMC), also referred to as the United States Marines, is a branch of the United States Armed Forces responsible for conducting amphibious operations with the United States Navy. The U.S. Marine Corps is one of the four armed service branches in the U.S. Department of Defense (DoD) and one of the seven uniformed services of the United States. Question: is the marines a part of the navy?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: As a British prince, William does not use a surname for everyday purposes. For formal and ceremonial purposes, the children of the Prince of Wales use the title of ``prince'' or ``princess'' before their Christian name and their father's territorial designation after it. Thus, Prince William was styled as ``Prince William of Wales''. Such territorial designations are discarded by women when they marry and by men if they are given a peerage of their own, such as when Prince William was given his dukedom. Question: is a prince the same as a duke?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: In the Time of the Butterflies is a historical novel by Julia Alvarez, relating an account of the Mirabal sisters during the time of the Trujillo dictatorship in the Dominican Republic. The book is written in the first and third person, by and about the Mirabal sisters. First published in 1994, the story was adapted into a feature film in 2001. Question: is in the time of the butterflies a true story?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The turkey vulture received its common name from the resemblance of the adult's bald red head and its dark plumage to that of the male wild turkey, while the name ``vulture'' is derived from the Latin word vulturus, meaning ``tearer'', and is a reference to its feeding habits. The word buzzard is used by North Americans to refer to this bird, yet in the Old World that term refers to members of the genus Buteo. The generic term Cathartes means ``purifier'' and is the Latinized form from the Greek kathart\u0113s/\u03ba\u03b1\u03b8\u03b1\u03c1\u03c4\u03b7\u03c2. The turkey vulture was first formally described by Linnaeus as Vultur aura in his Systema Naturae in 1758, and characterised as V. fuscogriseus, remigibus nigris, rostro albo (``brown-gray vulture, with black wings and a white beak''). It is a member of the family Cathartidae, along with the other six species of New World vultures, and included in the genus Cathartes, along with the greater yellow-headed vulture and the lesser yellow-headed vulture. Like other New World vultures, the turkey vulture has a diploid chromosome number of 80. Question: is a vulture the same as a buzzard?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: In management accounting or managerial accounting, managers use the provisions of accounting information in order to better inform themselves before they decide matters within their organizations, which aids their management and performance of control functions. Question: is managerial accounting and management accounting the same?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: First aid treatment is pressure on the wound and artificial respiration once the paralysis has disabled the victim's respiratory muscles, which often occurs within minutes of being bitten. Because the venom primarily kills through paralysis, victims are frequently saved if artificial respiration is started and maintained before marked cyanosis and hypotension develop. Efforts should be continued even if the victim appears not to be responding. Respiratory support until medical assistance arrives ensures the victims will generally recover. Question: can you survive a blue ringed octopus bite?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " There is a misconception in the student's answer. The passage states that victims can be saved if artificial respiration is started and maintained before marked cyanosis and hypotension develop. This means that with appropriate first aid treatment, it is possible to survive a blue-ringed octopus bite. Therefore, the score is: 50/100"
    },
    {
        "question": "Passage: In response to the National Minimum Drinking Age Act in 1984, which reduced by up to 10% the federal highway funding of any state which did not have a minimum purchasing age of 21, the New York Legislature raised the drinking age from 19 to 21, effective December 1, 1985. (The drinking age had been 18 for many years before the first raise on December 4th, 1982, to 19.) Persons under 21 are prohibited from purchasing alcohol or possessing alcohol with the intent to consume, unless the alcohol was given to that person by their parent or legal guardian. There is no law prohibiting where people under 21 may possess or consume alcohol that was given to them by their parents. Persons under 21 are prohibited from having a blood alcohol level of 0.02% or higher while driving. Question: can minors drink with parents in new york?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Physically and functionally, the auditory system of an absolute listener does not appear to be different from that of a non-absolute listener. Rather, ``it reflects a particular ability to analyze frequency information, presumably involving high-level cortical processing.'' Absolute pitch is an act of cognition, needing memory of the frequency, a label for the frequency (such as ``B-flat''), and exposure to the range of sound encompassed by that categorical label. Absolute pitch may be directly analogous to recognizing colors, phonemes (speech sounds), or other categorical perception of sensory stimuli. Just as most people have learned to recognize and name the color blue by the range of frequencies of the electromagnetic radiation that are perceived as light, it is possible that those who have been exposed to musical notes together with their names early in life will be more likely to identify, for example, the note C. Absolute pitch may also be related to certain genes, possibly an autosomal dominant genetic trait, though it ``might be nothing more than a general human capacity whose expression is strongly biased by the level and type of exposure to music that people experience in a given culture.'' Question: do you have to be born with perfect pitch?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The liver detoxifies and breaks down chemicals, poisons and other toxins that enter the body. For example, the liver transforms ammonia (which is poisonous) into urea in fish, amphibians and mammals, and into uric acid in birds and reptiles. Urea is filtered by the kidney into urine or through the gills in fish and tadpoles. Uric acid is paste-like and expelled as a semi-solid waste (the ``white'' in bird excrements). The liver also produces bile, and the body uses bile to break down fats into usable fats and unusable waste. Question: is the liver part of the excretory system?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Lowe's is the second-largest hardware chain in the United States behind The Home Depot and ahead of Menards. Globally, Lowe's is also the second-largest hardware chain, again behind The Home Depot but ahead of the European stores B&Q and OBI. Question: is home depot and lowes the same company?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 0"
    },
    {
        "question": "Passage: Ladies may wear a long (over the shoulders or to ankles) cloak usually called a cape, or a full-length cloak. Gentlemen wear an ankle-length or full-length cloak. Formal cloaks often have expensive, colored linings and trimmings such as silk, satin, velvet and fur. Question: is a cape and a cloak the same?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " There is a difference between a cape and a cloak according to the passage in the question. A cape is usually called a cloak, but it is different from a full-length cloak. Therefore the score is: 70"
    },
    {
        "question": "Passage: Six Flags New Orleans (SFNO) is a 140-acre, abandoned theme park in New Orleans that has been closed since Hurricane Katrina struck the state in August 2005. It is owned by the Industrial Development Board (IDB) of New Orleans. Six Flags had leased the park from 2002 until 2009, when the lease was terminated during its bankruptcy proceedings. The former park is located in New Orleans East, off Interstate 10. Question: is there a six flags in new orleans?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 10"
    },
    {
        "question": "Passage: Tully is a 2018 American comedy-drama film directed by Jason Reitman and written by Diablo Cody. The film stars Charlize Theron, Mackenzie Davis, Mark Duplass, and Ron Livingston, and follows the friendship between a mother of three and her nanny. It is the fourth collaboration between Reitman and Cody, following Juno (2007), Jennifer's Body (2009), and Young Adult (2011), the last of which also starred Theron. Question: is the film tully based on the book?",
        "pred_ans": "NILL",
        "ground_truth": false,
        "score": " Therefore the score is: 20"
    },
    {
        "question": "Passage: Mount Evans is the highest summit of the Chicago Peaks in the Front Range of the Rocky Mountains of North America. The prominent 14,271-foot (4350 m) fourteener is located in the Mount Evans Wilderness, 13.4 miles (21.6 km) southwest by south (bearing 214\u00b0) of the City of Idaho Springs in Clear Creek County, Colorado, United States, on the drainage divide between Arapaho National Forest and Pike National Forest. Question: is mt evans part of rocky mountain national park?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: The Federal Reserve began taking high-denomination currency out of circulation (destroying large bills received by banks) in 1969. As of May 30, 2009, only 336 $10,000 bills were known to exist; 342 remaining $5,000 bills; and 165,372 remaining $1,000 bills. Due to their rarity, collectors often pay considerably more than the face value of the bills to acquire them. Some are in museums in other parts of the world. Question: are there any thousand dollar bills in circulation?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: Uterine transplantation starts with the uterus retrieval surgery on the donor. Working techniques for this exist for animals, including primates and more recently humans. The recovered uterus may need to be stored, for example for transportation to the location of the recipient. Studies on cold-ischemia/eperfusion indicate an ischemic tolerance of more than 24 hours. Question: is it possible to do a uterus transplant?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " There is no specific information given in the passage to determine if the student's answer is correct or not. The passage only provides an overview of the process of uterine transplantation and does not mention if it is possible or not. Therefore, the score is: 50"
    },
    {
        "question": "Passage: A pitaya (/p\u026a\u02c8ta\u026a.\u0259/) or pitahaya (/\u02ccp\u026at\u0259\u02c8ha\u026a.\u0259/) is the fruit of several different cactus species indigenous to the Americas. Pitaya usually refers to fruit of the genus Stenocereus, while pitahaya or dragon fruit refers to fruit of the genus Hylocereus, both in the Cactaceae family. The dragon fruit is cultivated in Southeast Asia, Florida, the Caribbean, Australia, and throughout tropical and subtropical world regions. Question: is dragon fruit and pitaya the same thing?",
        "pred_ans": " True",
        "ground_truth": false,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: Arm span or reach (sometimes referred to as wingspan) is the physical measurement of the length from one end of an individual's arms (measured at the fingertips) to the other when raised parallel to the ground at shoulder height at a 90\u00b0 angle. The average reach correlates to the person's height. Age and sex have to be taken into account to best predict height from arm span. Question: is the width of your arms your height?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 80"
    },
    {
        "question": "Passage: Dalmatian puppies are born with plain white coats and their first spots usually appear within 3 to 4 weeks after birth, however spots are visible on their skin. After about a month, they have most of their spots, although they continue to develop throughout life at a much slower rate. Spots usually range in size from 30 to 60 mm, and are most commonly black or brown (liver) on a white background. Other, more rare colors, include blue (a blue-grayish color), brindle, mosaic, tricolor-ed (with tan spotting on the eyebrows, cheeks, legs, and chest), and orange or lemon (dark to pale yellow). Patches of color may appear anywhere on the body, mostly on the head or ears, and usually, consist of a solid color. Patches are visible at birth and are not a group of connected spots and are identifiable by the smooth edge of the patch. Question: do dalmatians get more spots as they grow?",
        "pred_ans": " False",
        "ground_truth": true,
        "score": " Therefore the score is: 50"
    },
    {
        "question": "Passage: Film rights were optioned by Paramount Pictures in a seven-figure deal in 2014, well before the book's publication, and Mark Johnson (Breaking Bad, Chronicles of Narnia) has signed on to produce the film. The book's success led to the acquisition of a sequel almost immediately after its release. The sequel, entitled A Torch Against the Night, was released in August 2017. The third book, A Reaper at the Gates, was published on June 12, 2018. Question: is an ember in the ashes a trilogy?",
        "pred_ans": "NILL",
        "ground_truth": true,
        "score": " Therefore the score is: 20"
    }
]