Woman And computer
Human And Computer
Showing posts with label Academia. Show all posts
Showing posts with label Academia. Show all posts

Plagiarism 2.0

Labels:

The German Defense Minister Karl-Theodor zu Guttenberg holds the title of Dr. jur. from the University in Bayreuth. He finished his thesis in 2006, at age 34, with more than 450 pages on the topic "Verfassung und Verfassungsvertrag: Konstitutionelle Entwicklungsstufen in den USA und der EU" (On the development of constitution in the USA and the EU). Guttenberg obtained the best possible grade, summa cum laude.

Two weeks ago, it turned out that big parts of his thesis were copied from other people's academic papers or newspaper articles. Since last week, one finds online a Wiki called GuttenPlag dedicated to collecting the copied paragraphs. The status is summarized in the below graphic (taken from mentioned Wiki):

Marked in black are pages on which plagiarized paragraphs have been found. Red are pages on which copies from several sources have been found. White means nothing has been found and blue is the table of contents and reference list that is not included in the search.

This eerily reminds me of a dissertation thesis I read last year. While the presented research was original, big parts of the text introducing the topic and explaining the relevance of the study were exact copies from other people's published review articles or research papers, including footnotes and references. The original work was cited in the text, but nowhere was it clearly marked the text was essentially an unauthorized reprint. Confronted with the evidence for his generous copying, the student first pointed out that he had cited the original papers. Yet, for a proper citation half of the thesis would have had to appear in quotation marks. Commenting on zu Guttenberg's "work," Volker Rieble, an expert on plagiarism summarized the core problem (as quoted in this Zeit article):

"Der Leser wird dar�ber get�uscht, dass ein bestimmter Absatz, ein bestimmtes Textst�ck, ein bestimmter Gedanke nicht vom Doktoranden zu Guttenberg, sondern von einem anderen stammt. Und das ist mit wissenschaftlichen Standards schlechterdings nicht vereinbar."

"The reader is deceived in knowing that a particular paragraph, a particular part of the text, a particular thought, did not come from doctoral candidate zu Guttenberg but from somebody else. This is not in accordance with scientific standards."


So you're not done with putting a citation somewhere, you have to make clear to the reader what is the extend of your borrowing. On further inquiry, the candidate whose thesis I had read - not a native English speaker - said with heartwarming honesty he had started writing the text but then found the other authors had said it so much better and clearer that the reader would benefit from using their words. The thesis was withdrawn and replaced prior to the defense. The candidate passed - as I said, his research was fine. Zu Guttenberg, whose copying work was only noticed after his defense, now has to await the University of Bayreuth's decision on whether he will be allowed to keep his title.

I know several examples where physicists, including myself on more than one occasion, have found paragraphs from their papers reappear in other people's papers. While the source was quoted somewhere in the text, the copied paragraphs were not marked as quotation. In all cases I know of, the people copying others' texts were not native English speakers.

Not a native English speaker myself, coming up with a well written motivation for a paper is a problem I can relate to. Otoh, at least I have an excuse for being grammatically challenged ;-) One should also note that some journals do offer editorial help with grammar and spelling. (Better read your proofs very, very carefully.) In any case, I'm bringing this up because already in a post some months ago, where I remarked upon the unreferenced spread of some of my pictures into other people's slide presentations, I was wondering if the possibility of copy-and-pasting is too much a temptation to resist or whether people just think nothing about it. The Times Higher Education for example recently reported that Chinese students admit to little or no idea about ethics:
"Research carried out by academics at Beihang University in Beijing found a startling lack of understanding of plagiarism and academic misconduct, with both students and staff admitting that they knew "very little" or "had no idea" about the norms of scientific ethics. [U]p to 10 per cent of the students surveyed said that they thought copying work directly from the internet should not be considered bad practice."

Even more depressingly, an increasing amount of college applicants seems to be lifting their "personal statements":
"[M]any applicants borrowed phrases from the same free website... In 234 applications to study medicine [from 50,000 applications to study medicine, dentistry and veterinary science at the universities of Oxford and Cambridge], candidates wrote that it was �burning a hole in my pyjamas at age eight� that sparked their passion for the subject."

So much about individualism.

The reason this depresses me is that these young people willingly give up the offered possibility of personalizing their application. The alternative is being reduced to numbers and, eventually, being assessed by some measure for success.

So, evidently, copy and pasting others' texts is becoming ever more common, and many people at least claim to not know it's unethical not to properly cite ones' sources. Where does this get us? I am wondering now if not time will come when a scientist can assemble parts of his paper from already published articles - a motivation from there, some literature review from there, summary of the method from there, of course marked as quotation - and just add the relevant new equations, tables, and figures. Does everybody really have to write the always same introduction in his own words (and then plagiarize himself in further publications)?

Sunny with scattered papers

Measuring SuccessSeed magazine has an interesting article On Science Transfer. It is about the measurement of scientific success by means of automatized metrics, a topic we have discussed several times on this blog, see eg my posts Science Metrics and Against Measure.

The mentioned article is interesting in that it focuses on measuring scientific activities that are not usually considered for academic purposes, those of communicating science and being relevant for science policies - that's what is meant with �science transfer.� To that end, commonly used measures based on citations are of limited use:

�If we want to know what scientific ideas are influencing decisions and policymaking in the public sphere or in disparate scientific fields, rather than simply the discipline in which an idea originated, citations are of less relevance [...] Writing in the popular press is equally unlikely to garner citations. Even trying to translate research into something more digestible by a lay audience within the academic publishing world is a dead end; editorial and other journalistic material is generally deemed �uncitable.�

Though it is by no means the only aspect of scientific culture responsible, the fixation on citations as a measure of scholarly impact has given scientists few reasons to communicate the value of their work to non-scientists.�
The article then discusses the possibility of more general measures of impact, based on usage, such as for example MESUR. I am skeptic that usage is an indicator for quality rather than for popularity. Some works arguably score a lot of hits and downloads exactly because they turn out to be utter nonsense.

But either way, I certainly welcome the attempt to take note of a scientist's impact on informing the public. A few days ago, Vivienne Raper had an interesting blogpost on Science Blogging and Tenure summarizing the pros and cons of blogging next to doing research. She reports an example from innovation-country Canada:
�Cell biologist Alexander Palazzo says his blog helped him secure an assistant professorship. "My department" -- the biochemistry department at the University of Toronto in Canada -- "told me part of the reason they hired me was because of stuff I'd written on my blog," he says. "It wasn't the main reason they hired me, but it helped."�

Another item on the topic of getting science closer to the public and the role of blogging: In the last 3 months or so I received about 5 emails from freelance writers with a record of science-themed articles, asking for a guest post. As you can see I said thanks but no thanks, but I find this an interesting development. It seems there's people for who blogs represent a useful medium to earn career credits.

But back to the Seed article: it is interesting for another reason. As we previously discussed, purely software generated measures can be unreliable, as is shown by the example of a whole university's high ranking going back to the number of publications of one of their researchers (who published several hundred papers in a journal of which he also happened to be editor in chief) and the example of how the h-index of a (not even existent) author can be pimped to that of an exceptional scientist. The Seed article takes note of this problem by acknowledging the need of human interpretation of data - a task for the "science meteorologist"
�Even if we erect massive databases filled with information on how scientific work is being used in real time, for the foreseeable future it seems inescapable that humans must provide oversight to derive actionable knowledge from the data. Modern weather forecasting provides an illustrative example: Copious real-time data on world weather patterns is available to anyone with a computer and an internet connection, but the vast majority of us rely on meteorologists to synthesize and analyze it to produce a daily forecast. Moreover, even more raw data and subsequent analysis are necessary to transform information about weather into knowledge about climate and how human activity has influenced it over the course of centuries.

Well-designed computer programs may be able to compile usage data on scientific discourse and publishing to generate real-time maps of scientific activity, but such maps can only inform our decision making, not replace it. A new skill set that makes use of such tools�a kind of �science meteorology��will be necessary to serve as a bridge between the academic and public spheres.�

Granted, they are concerned with measuring the impact of scientific work on policy decisions, but I couldn't help wondering what a science meteorologist would "forecast" from data of individual scientists. This candidate is sunny with scattered papers? Clear and cold with a student chill factor of zero K? Partly cloudy with a 10% chance of tenure?

The Seed article also touches on an issue I previously commented on here:
�The problem with evaluating all [scientists] with one fast and easy evaluation system is centralization and streamlining. The more people use the same system, the more likely it becomes everybody will do the same research with the same methods.�

Also Michael Nielsen recently wrote an excellent post on The Mismeasurement of Science making this point:
�I accept that metrics in some form are inevitable � after all [...] every granting or hiring committee is effectively using a metric every time they make a decision. My argument instead is essentially an argument against homogeneity in the evaluation of science: it�s not the use of metrics I�m objecting to, per se, rather it�s the idea that a relatively small number of metrics may become broadly influential. I shall argue that it�s much better if the system is very diverse, with all sorts of different ways being used to evaluate science.�

(Michael is btw writing a book titled �Reinventing Discovery,� about to be published this year. Something for your reading list.) In the Seed article now one finds a quotation from Johan Bollen, associate professor at Indiana University�s School of Informatics and Computing, who is the brain behind the MESUR project:
�If you have a bunch of different metrics, and they each embody different aspects of scholarly impact, I think that�s a much healthier system.�

We can agree on that. Then Bollen continues:
�People�s true value can be gleaned [...]�

Let's hope the day a scientist's �true value� is defined by a software will never come.

Summary:
  • Efforts are made to measure scientist's skills of communicating research to the public and policy makers. Useful for evaluating success, as defined by the measure, and for providing incentives. -- Good.

  • Measuring success by usage. -- Questionable.

  • Noting that data collection still needs human assessment. -- Good.

  • Diversifying in measures prevents streamlining and is thus welcome or, in other words, if you have to use metrics at least use them smartly. -- Indeed.

  • People's true value can be gleaned... -- Pooh.

  • Michael's book is almost done. -- Yeah!

Fun with the h-index

Labels:

The h-index is a widely used measure for a scientist's scientific productivity and impact somewhat more sophisticated than just the number of publications. The h-index is the greatest positive integer number h, such that the scientist has h papers each of which has been cited at least h times. If you're wondering how relevant the h-index is in practice, I have no way of telling in general. I know however that I've been in committees where the h-index evidently was an interesting point of reference for some of its members, and I have also been asked a few times what my h-index is. (Before you ask, according to SPIRES my h-index is either 14 or 16, depending on whether you count all or only published papers.) The absolute number isn't of much importance in most cases, it matters instead how you compare to others in your particular field - as Einstein taught us, everything is relative ;-)

Next time somebody asks for my h-index, I'll refer them to this hilarious paper by Cyril Labb� from the Laboratoire d'Informatique de Grenoble at the Universit� Joseph Fourier:

    "Ike Antkare, One of the Great Stars in the Scientific Firmament"
    22th newsletter of the International Society for Scientometrics and Informatrics (June 2010)
    PDF here

Labb� has created a fictional author, Ike Antkare, and pimped Ike's h-index to 94. For this, Labb� created 102 "publications" using a software resembling a dada-generator for computer science called Scigen, and a net of self-citations. Labb�'s paper contains an exact description of the procedure. His spoof works for tools that compute the h-index based on Google scholar's data; the best known is maybe Publish or Perish.

What lesson do we learn from that?

First, the Labb�'s method works mainly because he uses the h-index computed with a quite unreliable database, Google scholar, to which it is comparably easy to add "fake" papers. While for example the arXiv database also contains unpublished papers, it does have some amount of moderation which I doubt 102 dada-generated papers by the same author would get past. In addition, SPIRES offers the h-index for published papers only. (Considering however that I know more and more people - all tenured of course - who don't bother with journals, restricting to published papers only might in some cases give a very misleading result.)

Second, and maybe more importantly, I doubt that any committee that were faced with Ike's amazing h-index would be fooled, since it only takes a brief look at his publications to set the record straight.

Nevertheless, Labb�'s paper is a warning to not use automatically generated measures for scientific success without giving the so obtained results a look. Since the use of metrics in science for evaluation of departments and universities is becoming more and more common, it's an important message indeed, and an excellent example for how secondary criteria (high h-index) deviate from primary goals (good research).

For more on science metrics, see my most Science Metrics and Against Measure. For more on the dynamics of optimization in the academic system, and the mismatch between primary goals and secondary criteria, see The Marketplace of Ideas and We have only ourselves to judge each other.

Thanks to Christine for drawing my attention to this study.

 
Internet