Contents of the corpus ---------------------- The corpus consists of extracts from a number of texts: danish: A: Retsplejeloven (Chap 1) B: Karnorvs lovsamling (Chap 2) C: Gomard: Civilprocessen (Chap 4) D: Dansk dokumentsamling (Chap 5) spanish: A: Pietro-Castro: Tratado de derecho procesal civil (Chap 6) B: Ley de Enjuiciamiento Civil (Chap 7) C: Spansk dokumentsamling (Chap 8) english: A: The Rules of the Supreme Court Supreme Court act County Court Rules (Chap 9) B: Engelsk dokumentsamling (Chap 10) C: O'Hara & Hill: Civil Litigation (Chap 11) We have tried to build a corpus of similar texts in as much as for each language there is the law text itself, comments on the law, textbooks and different kinds of documents from 'real life'. It is supposed to be a representative text corpus in the area of civil law. We have chosen a specific area of civil law i.e. Civil procedure, courts of first instance. The corpus was contributed to the ECI project by: **************************************** * Steffen Leo Hansen * * Dept.of Computational Linguistics * * Dalgas Have 15, DK - 2000 F. * * ==================================== * * e-mail: slh/id@cbs.dk * * phone: + 45 38 15 31 27 * * fax: + 45 38 15 38 20 * **************************************** Structure of the ECI corpus --------------------------- The Civil Law corpus has been split into three subcorpora corresponding to the three languages which it contains viz. Danish (mda12), English (men12) and Spanish (msp12). Each subcorpus has been split into a number of files corresponding to the divisions given above, i.e. each .eci file contains all of the extracts from one of the book chapters given above. These divisions are marked in the corpus with: Each file contains a number of extracts taken from the book. These are marked by: The NNNNNN are identification numbers given to the extracts by the CBS. They are not neccessarily unique, ie more than one extract can have the same identifier as some long sections have been split into more than one extract. Each extract consists of a bibliographic header, the text extract itself, and a bibliographic trailer (which mostly repeats the information in the header). The texts are marked with: The text of the extracts is divided into paragraphs with

. The type 'cbs' indicates that these paragraphs were marked up in the original corpus from the Copenhagen Business School.

was used to mark the start of those texts which did not begin with a CBS paragraph. It may also sometimes contain notes marked with: The structure of a text extract is thus: ... header bibliographic entry ...

... text ...

...

... text of note ...

...
... trailer bibliographic entry ...
where there may be multiple

and within the . Corrections/Modifications to the corpus --------------------------------------- The corpus originally contained many occurances of the character ctl-U (\025). These have all been quietly converted to the character § (paragraph marker), as they usually appeared where this character would have been appropriate. And in fact \025 is the code for § in the IBM Code Page 850 character set. Each line in the original corpus ended with the "<" character. These have been removed. A number of obvious mistakes have been corrected with either .. or markup, with the original text as an attribute.