Εμφάνιση αναρτήσεων με ετικέτα Library of Congress. Εμφάνιση όλων των αναρτήσεων
Εμφάνιση αναρτήσεων με ετικέτα Library of Congress. Εμφάνιση όλων των αναρτήσεων

Δευτέρα, Σεπτεμβρίου 19, 2016

Νέες ιδέες στη βιβλιοθήκη Κογκρέσου


Τη νύχτα που κάηκε η Βαλτιμόρη, τον Απρίλιο του 2015, η Κάρλα Χέιντεν, η βιβλιοθηκάριος της πόλης, είχε εντολή να κλείσει ερμητικά τη βιβλιοθήκη και να μείνει ασφαλής μέχρις ότου θα καταλαγιάζαν τα βίαια επεισόδια που ξέσπασαν εξαιτίας του θανάτου του Φρέντι Γκρέι στα χέρια της αστυνομίας.

Η ίδια είχε διαφορετική γνώμη. Τη στιγμή που ο κυβερνήτης του Μέριλαντ κήρυττε κατάσταση έκτακτης ανάγκης, η Χέιντεν και οι συνάδελφοί της αποφάσισαν να ανοίξουν τις πόρτες της βιβλιοθήκης την επόμενη ημέρα και να δεχθούν τους ανήσυχους συμπολίτες τους. Για τη δόκτορα Χέιντεν, που ορκίστηκε την Τετάρτη 14η βιβλιοθηκάριος της Βιβλιοθήκης του Κογκρέσου, οι ταραχές αποτέλεσαν μια δοκιμή για τις αξίες της και την έκαναν να συνειδητοποιήσει ότι οι βιβλιοθήκες ήταν κάτι πολύ περισσότερο από τα βιβλία τους.

«Οι κάτοικοι της περιοχής προστάτευσαν τη βιβλιοθήκη», δήλωσε η δρ Χέιντεν σε πρόσφατη συνέντευξή της. «Οι νεαροί άνδρες που στέκονταν απέξω ήταν ένα σύμβολο».

Στα 64 της χρόνια η δρ Χέιντεν είναι η πρώτη Αφροαμερικανή και η πρώτη γυναίκα που θα είναι επικεφαλής της Βιβλιοθήκης του Κογκρέσου. Μιας βιβλιοθήκης με ιστορία 216 ετών, από τις μεγαλύτερες του κόσμου, και μια πηγή γνώσης και πολιτισμού των ΗΠΑ. Η δρ Χέιντεν διορίστηκε από τον πρόεδρο των ΗΠΑ Μπαράκ Ομπάμα. Είναι η πρώτη νέα βιβλιοθηκάριος της Βιβλιοθήκης του Κογκρέσου από το 1987 και φέρνει μαζί της καινούργιες ιδέες για προσβασιμότητα, τεχνολογία αλλά τον ρόλο που θα πρέπει να διαδραματίζουν οι βιβλιοθήκες στην κοινωνία. Στόχος της είναι να ανοίξει διάπλατα τις πόρτες της Βιβλιοθήκης του Κογκρέσου στους Αμερικανούς. «Σε αυτό το σημείο της ιστορίας της, η αξία της βιβλιοθήκης ως ενός τόπου ακαδημαϊκής μελέτης δεν θα σβήσει. Αντιθέτως, θέλουμε να ενισχύσουμε αυτή την πραγματικότητα, αλλά αν η βιβλιοθήκη ανοίξει και ένα παράθυρο προς τα έξω, περισσότεροι άνθρωποι να συνειδητοποιήσουν ότι μπορούν να γίνουν ακαδημαϊκοί», τόνισε η δρ Χέιντεν.

Αυτό σημαίνει ότι θα ενισχυθεί η ψηφιακή πρόσβαση στη βιβλιοθήκη και θα συνδεθούν οι συλλογές της με τα προγράμματα των σχολείων σε ολόκληρη τη χώρα. Επίσης, η δρ Χέιντεν έχει δείξει ενδιαφέρον για ενίσχυση της συνεργασίας με τις δημόσιες και τις πανεπιστημιακές βιβλιοθήκες. Στην ομιλία της, μετά την ορκωμοσία της, η δρ Χέιντεν τόνισε ότι θα ήθελε οι Αμερικανοί να μπορούν να δουν περισσότερο υλικό όπως οι σημειώσεις και οι επιστολές της Ρόζας Παρκ, της Αφροαμερικανής που αγωνίστηκε για τα ατομικά δικαιώματα, τις οποίες η βιβλιοθήκη ψηφιοποίησε έτσι ώστε όλοι να έχουν σε αυτές πρόσβαση από τον ηλεκτρονικό τους υπολογιστή. Το έργο της δόκτορος Χέιντεν δεν θα είναι εύκολο. Οι έλεγχοι της βιβλιοθήκης έχουν αποκαλύψει αδυναμίες, μεταξύ των οποίων η σώρευση εκατομμυρίων βιβλίων που παραμένουν στις αποθήκες και κινδυνεύουν με καταστροφή αλλά και οι μεγάλες καθυστερήσεις στην ψηφιοποίηση των βιβλίων και άλλων εντύπων που διαθέτει η βιβλιοθήκη.

Πηγή: Η Καθημερινή


Παρασκευή, Μαρτίου 01, 2013

What the Library of Congress Plans to Do With All Your Tweets

…That’s a message I tweeted back in April 2010, when it was announced that the Library of Congress was planning to archive every publicly available tweet ever posted on the social network. Now, almost three years later, the Library’s Twitter archive is beginning to take shape, and there are clues as to what uses researchers will derive from all of our 140-character witticisms.
When the Library initially took on the Twitter archive in 2010, it was already a daunting 21 billion tweets filled with words, hashtags, geolocation info, and other metadata. Today the Library has access to more than 170 billion tweets or about 85 terabytes of data. With about half a billion tweets now flowing into the archive daily, the biggest immediate challenge is finding a way to make all this information coherent and usable.
“One of the things that makes this collection a little bit different is the velocity with which it’s growing,” says Gayle Osterberg, director of communications for the Library of Congress. “The computing capacity to search for an item or a series of items across billions and billions of tweets isn’t cost-effective at the present time for a public institution.”

Osterberg says the costs associated with the project, in terms of developing the infrastructure to house the tweets, is in the low tens of thousands of dollars. The tweets were offered as a free gift from Twitter, and are being transferred to the Library through a separate company, Gnip, at no cost. Each day tweets are automatically pulled in from Gnip, organized chronologically and scanned to ensure they’re not corrupted. Then the data are stored on two separate tapes which are housed in different parts of the Library for security reasons.
The Library has mostly figured out how to make the archive organized, but usability remains a challenge. A simple query of just the 2006-2010 tweets currently takes about 24 hours. Increasing search speeds to a reasonable level would require purchasing hundreds of servers, which the Library says is financially unfeasible right now. There’s no timetable for when the tweets might become accessible to researchers.
“The goal would be to be able to answer whatever query a researcher might have here at the Library in our reading room,” Osterberg says. “The balance is making the access both meaningful and cost-effective for the Library.”
While you can’t yet make a trip to Washington D.C. and have casual perusal of all the world’s tweets, the technology to do exactly that is readily available—for a cost. Gnip, the organization feeding the tweets to the Library, is a social media data company that has exclusive access to the Twitter “firehose,” the never-ending, comprehensive stream of all of our tweets. Companies such as IBM pay for Gnip’s services, which also include access to posts from other social networks like Facebook and Tumblr. The company also works with academics and public policy experts, the type of people likely to make use of a free, government-sponsored Twitter archive when it comes to fruition.
Through Gnip, researchers have already made extensive use of much of the Twitter archive. Sherry Emery, a senior scientist at the Institute for Health Research and Policy at the University of Illinois at Chicago, analyzes tweets about smoking to understand the role of media in influencing the habit. When the Center for Disease Control launched a graphic anti-smoking ad campaign last spring, Emery and her team were able to analyze every public tweet about smoking and understand how people were reacting to the commercials.

“We can’t use Twitter to look at whether they actually quit smoking,” Emery says. “But we can really get a better understanding of whether people embraced the message of the ad or not, which is an important intermediate step to making changes in their behavior.”
Even with the tools available to quickly search the Twitter archive, making sense of such a huge dataset can be a challenge. Emery’s group has amassed more than 50 million tweets about smoking since December 2011. “A big part of what we do is just cleaning the data to make sure that the tweets that we’re looking at are about smoking tobacco and not about smoking weed or smoking ribs or smoking hot girls,” she says.
Using computer software to assess human emotion in tweets is also tricky. For example, someone tweeting “This is scary!” about the CDC commercial featuring former smokers with artificial voice boxes might seem like a negative reaction to a computer program. In fact, it’s the desired effect for an ad aimed at curbing smoking. Human coders have to feed the computer between 500 and 1,000 sample tweets to help it properly understand how to organize responses to a research question.
Other uses for tweets have also emerged. Daniel Hodd, a business student at Fordham University, is studying the conversation surrounding 50 stocks on Twitter to see if investor sentiment correlates with stock price. Chris Cantey, a master’s student studying cartography at the University of Wisconsin, has already used the more limited Twitter API system to geographically map last month’s flu outbreak. He’s now using the full firehose to analyze how responses to Hurricane Sandy unfolded in real time.

All the researchers agree that Twitter is a powerful tool for sociological study. Soon, if the Library of Congress can make its database fully functional, it’ll also be an easily accessible one. And one day, long after we’ve all sent our final snarky tweet, our messages will live on.
“Social media gives even those among us who don’t have the time to pick up a pen every day an opportunity to be recorders of and witnesses to history,” says Osterberg. “Those perspectives will be incredibly valuable to researchers and authors and policy makers down the road who want to understand the times we’re living in today.”


Source: business.time.com

Τρίτη, Ιανουαρίου 08, 2013

Library of Congress has archive of tweets, but no plan for its public display


 
In the few minutes it will take you to read this story, some 3 million new tweets will have flitted across the publishing platform Twitter and ricocheted across the Internet. The Library of Congress is busy archiving the sprawling and frenetic Twitter canon — with some key exceptions — dating back to the site’s 2006 launch. That means saving for posterity more than 170 billion tweets and counting, with an average of more than 400 million new tweets sent each day, according to Twitter.

But in the two years since the library announced this unprecedented acquisition project, few details have emerged about how its unwieldy corpus of 140-character bursts will be made available to the public.
That’s because the library hasn’t figured it out yet.

“People expect fully indexed — if not online searchable — databases, and that’s very difficult to apply to massive digital databases in real time,” said Deputy Librarian of Congress Robert Dizard Jr. “The technology for archival access has to catch up with the technology that has allowed for content creation and distribution on a massive scale. Twitter is focused on creating and distributing content; that’s the model. Our focus is on collecting that data, archiving it, stabilizing it and providing access; a very different model.”

Colorado-based data company Gnip is managing the transfer of tweets to the archive, which is populated by a fully automated system that processes tweets from across the globe. Each archived tweet comes with more than 50 fields of metadata — where the tweet originated, how many times it was retweeted, who follows the account that posted the tweet and so on — although content from links, photos and videos attached to tweets are not included. For security’s sake, there are two copies of the complete collection.

But the library hasn’t started the daunting task of sorting or filtering its 133 terabytes of Twitter data, which it receives from Gnip in chronological bundles, in any meaningful way.

“It’s pretty raw,” Dizard said. “You often hear a reference to Twitter as a fire hose, that constant stream of tweets going around the world. What we have here is a large and growing lake. What we need is the technology that allows us to both understand and make useful that lake of information.”

For now, giving researchers access to the archive remains cost-prohibitive for the cash-strapped library, which has spent tens of thousands of dollars on the project so far, Dizard says. Like many federal agencies, the Library of Congress has been hit by budget cuts in recent years. Without a major overhaul to its computing infrastructure, it isn’t equipped to handle even the simplest queries.

“We know from the testing we’ve done with even small parts of the data that we are not going to be able to, on our own, provide really useful access at a cost that is reasonable for us,” Dizard said. “For even just the 2006 to 2010 [portion of the] archive, which is about 21 billion tweets, just to do one search could take 24 hours using our existing servers.”
Instead, the library is exploring whether it might be able to afford to pay a third party to provide public access to the archive. But for those who have immediate research interests — and many people have contacted the library, Dizard says — the wait is maddening.

Gnip President Chris Moody says he’s used to serving clients like major corporations and political campaigns that expect data right away.
“Milliseconds is not uncommon for expected latency from when the tweet happened to when someone would be able to get it and analyze it,” he said.

Even after questions of access are resolved, Moody says he expects centuries to pass before the full value of the Twitter archive can be realized.

“We’re very, very early,” Moody said. “We’re 1 percent of the way into what this data will mean.”

The eventual plan is to make the collection available only within the Library of Congress reading rooms. Requiring an in-person visit to search a database of material that originated online may seem incongruous, but Dizard says it’s a condition of the deal with Twitter, which gifted the archive, so that the library won’t be “competing with the commercial sector.”

There are other limitations. The library is not archiving tweets from those who opt for the strictest privacy settings, which allow Twitter users to approve or reject each potential follower. The library is also planning to scrub deleted tweets, meaning the public won’t have access to posts that were published but later removed. Dizard, citing privacy concerns, calls that decision “one of the more significant policy questions we face.”

In its terms of service, Twitter says that the default is “almost always to make the information you provide public for as long as you do not delete it from Twitter.”

Moody says it follows that deleted tweets are off-limits.

“Twitter’s terms of service are quite clear,” Moody said. “Any organization that accesses Twitter data through us, and this certainly applies to the Library of Congress as well, has to comply with the terms of service.”

The tension lies in the historical value of seeing what a person publishes, then erases. The sexually suggestive tweet that led to Rep. Anthony Weiner’s resignation is one of the splashiest examples of how deleted tweets can be significant, but even seemingly mundane deletions could carry weight with the passage of time.

The nonprofit Sunlight Foundation has a site, called Politwoops, which culls politicians’ deleted tweets in a searchable collection. Tom Lee, the director of the foundation’s Sunlight Labs, says he finds it bizarre that the library’s archive would exclude such tweets.

“You can’t make a TV appearance or press release or speech just disappear,” Lee said. “It’s not clear to me why someone should be allowed to remove something from the public record.”

A Twitter spokesman wouldn’t say whether the site has considered retroactively making public deleted tweets the way some government files are declassified after a certain number of years. But there are people within the Library of Congress who argue that deleted tweets ought to be part of the archive, Dizard says. The debate is unlike anything the library has had to consider in its 212-year existence.

“You could look at it strictly and say that anybody who puts a tweet up has published it,” Dizard said. “We have never received a collection that has ownership transferred through a click-through agreement, so that’s the difference. Most of our collections come with signed agreements or purchase. This is a different way of acquiring.”

The Twitter archive also signals a shift in how the library sees what kinds of acquisitions are possible. The library will continue amassing physical objects such as personal papers, books, maps and copyright registrations. But Dizard says the Library of Congress is also exploring how to acquire other more ephemeral trappings of the digital realm, such as Google-search histories, for example.

“Search-engine requests are a very, very good indication of what people are thinking,” Dizard said. “The acquisition of the Twitter archive better prepares us and encourages us to take other aspects of social media and digital content. That’s something we’ll have to do. . . . The acquisition of the Twitter archive is a start for us, not a test about whether we want to continue or not. This is really a critical part of the mission of the library.”
 



Μηχανή αναζήτησης ελληνικών ψηφιακών βιβλιοθηκών

Περί Βιβλίων & Βιβλιοθηκών