Arno Deceuninck & Basil Rommens
run.sh for running main.py with y.c
test.sh for running unittests
- Binary operations +, -, * and /
- Binary operations >, < and ==
- Unary operators + and -
- Brackets to overwrite the order of operations
- AST
- Binary operator %
- Comparison operators >=, <=, and !=
- Logical operators &&, ||, and !
- Types (char, float, int and pointer)
- Reserved words (const)
- Variables
- AST
- Identifier operations (++ and --)
- Conversions
- Comments (single line and multiline)
- Printf
- LLVM generation (niet mogelijk om pointers te gebruiken)
- Retaining comments in compilation process
- Comment after every instruction that contains the statement from the input code
- Reserved words: If
- Reserved words: Else
- Reserved words: While
- Reserved words: For
- Reserved words: Break
- Reserved words: Continue
- Scopes: unnamed scopes
- Scopes: Loops
- Scopes: Conditionals
- Reserved words: Switch
- Reserved words: Case
- Reserved words: Default
- Reserved words: return
- Reserved words: void
- Scopes: function
- Local and global variables
- Functions
- Do not generate code after return statement
- Do not generate code after break or continue
- Do not generate code for variables that are not used
- Do not generate code for conditionals that are always false
- Check if all functions end with return
- Arrays
- Import: printf(char *format, ...)
- Import: intf(const char *format, ...)
- Arrays: Multi-dimensional
- Arrays: assignments of complete arrays or array rows in case of multi-dimensional arrays
- Arrays: Dynamic arrays
Om llvm leesbaar te houden voor ons en het eenvoudig te houden in presentatie, maken we gebruik van verschillende middelen.
Voor de llvm generatie van variabelen zorgen we er voor dat er geen overlap gebeurt tussen de variabelen op de volgende manier.
Het eerste wat we doen is globale variabelen hun gewone naam laten houden (<var_name>). Andere variabelen worden anders behandeld.
Indien we deze tegenkomen nemen we de variabele naam en plakken we achter de variabele een . en dan nog de id van de node waar ze
in zitten (<var_name> + . + <node_id>). Het punt dient er voor dat variabelen niet met een getal kunnen overeenkomen als we dit op een of andere manier proberen te bereiken.
Het punt kan je niet in c gebruiken bij een variabele naam. Dan hebben we de automatisch gegenereerde variabelen die beginnen altijd
met een punt en worden gevolgd door een letter, dit kan een v of een t zijn. v wijst op een variabele die wordt later meestal gebruikt, terwijl t wijst
op een tijdelijke variabele, deze is altijd anders wanneer men die genereerd van een node. Na deze letter volgt een unieke id
bij een t is die bij elke generatie anders, maar bij een v is die gelijk aan de id van de node zo blijft deze variabele herbruikbaard (. + t|v + ).
Tot slot hebben we nog de strings die er gebruikt worden in printf en scanf. Deze zijn van de volgende structuur (. + str + . + <node_id>).
We zorgen er ook voor dat er bij de llvm generatie commentaren bij komen, dit zorgt er grotendeels voor dat we de snel kunnen terugvinden waar we eventueel een fout in hebben gemaakt. Ook worden code blokken aangeduid, let wel op dat dit niet enorm betrouwbaar is. Het enigste waar we in commentaren generatie in verschillen is dat deze voor de llvm statements gebeurd.
Om mips te genereren vertrekken we vanuit llvm. Zo kunnen we veel dubbel werk vermijden, zoals het bepalen van scopes, variabelen toekennen enz.
We slaan al de variabelen op in de stack. We zoeken eerst al de gebruikte variablen in de llvm code en bepalen dan hoe groot de stack moet zijn. Omdat er geen enkele variabele overlap heeft ook niet met functies, kunnen we de stackpointer relatief klein houden. Zo moeten we geen rekening houden met de variabelen die reeds gebruikt zijn, met de doorgegeven argumenten is er een speciale operatie nodig, net als al de variabelen in de functie, die moeten ook opgeslagen worden. Dit gebeurd met de stackpointer
Indien er een bewerking wordt gedaan op 2 integer getallen, wordt de uitkomst afgerond en omgezet naar een int
Constanten
Operatoren
Variabelen
Gereserveeerde types
Functies en argumentenlijst
Commentaren
Include
Default Nodes
Omdat het ridicuul zou zijn om elke mogelijke escape sequence van c te implementeren, hebben we er slechts eensubset van geimplementeerd. Deze subset is in de afbeelding hieronder afgebeeld.

In de tests folder vind je een bestand test.py. Als je dit runt (eventueel test.sh, runnen er ineens een paar testen. De uitleg waarvoor
welke test specifiek dient kan je terugvinden in het c bestand van de test zelf (in de folder tests/input). Om specifiek te weten waarop de testen werken, kan je een kijkje nemen in het tests/test.py.
The Benchmark tests have been changed from the given ones, because we don't support for example i++, so had to change this everywhere to i = i + 1.
- binaryOperations1.c
- binaryOperations2.c
- breakAndContinue.c
- comparisons1.c
- comparisons2.c
- dereferenceAssignment.c
- fibonacciRecursive.c
- floatToIntConversion.c
- for.c
- forwardDeclaration.c
- if.c
- ifElse.c
- intToFloatConversion.c
- modulo.c
- pointerArgument.c
- prime.c
- printf1.c
- printf2.c
- printf3.c
- scanf1.c
- scanf2.c (passes because manual input required)
- scoping.c
- unaryOperations.c
- variables1.c
- variables2.c
- variables3.c
- variables4.c
- variables5.c
- variables6.c
- variables7.c
- variables8.c
- while.c
- arrayAccessTypeMismatch2.c
- arrayAccessTypeMismatch.c
- arrayCompareError.c
- arraySizeTypeMismatch.c
- declarationDeclarationMismatch1.c
- declarationDeclarationMismatch2.c
- declarationDeclarationMismatch3.c
- declarationDefinitionMismatch1.c
- declarationDefinitionMismatch2.c
- declarationDefinitionMismatch3.c
- definitionInLocalScope.c
- dereferenceTypeMismatch1.c
- dereferenceTypeMismatch2.c
- functionCallargumentMismatch1.c
- functionCallargumentMismatch2.c
- functionCallargumentMismatch3.c
- functionCallargumentMismatch4.c
- functionRedefinition1.c
- functionRedefinition2.c
- functionRedefinition3.c
- incompatibleTypes1.c
- incompatibleTypes2.c
- incompatibleTypes3.c
- incompatibleTypes4.c
- incompatibleTypes5.c
- incompatibleTypes6.c
- incompatibleTypes7.c
- invalidIncludeError.c
- invalidLoopControlStatement.c
- invalidUnaryOperation.c
- mainNotFound.c
- parameterRedefinition1.c
- parameterRedefinition2.c
- parameterRedefinition3.c
- pointerOperationError.c
- returnOutsideFunction.c
- returnTypeMismatch.c
- undeclaredFunction.c
- undeclaredVariable1.c
- undeclaredVariable2.c
- undeclaredVariable3.c
- variableRedefinition1.c
- variableRedefinition2.c
- variableRedefinition3.c
- variableRedefinition4.c
- variableRedefinition5.c
- variableRedefinition6.c
###MIPS tests (28/32)
- binaryOperations1.c
- binaryOperations2.c
- breakAndContinue.c
- comparisons1.c
- comparisons2.c
- dereferenceAssignment.c
- fibonacciRecursive.c
- floatToIntConversion.c
- for.c
- forwardDeclaration.c
- if.c
- ifElse.c
- intToFloatConversion.c
- modulo.c
- pointerArgument.c
- prime.c
- printf1.c
- printf2.c
- printf3.c
- scanf1.c
- scanf2.c (passes because manual input required)
- scoping.c
- unaryOperations.c
- variables1.c
- variables2.c
- variables3.c
- variables4.c
- variables5.c
- variables6.c
- variables7.c
- variables8.c
- while.c
- Start with optional include and operation sequence
- operation sequence can be an operation, function, conditional statement or unnamed scope
- operation can be an assignment, return operation or arithmetic operation
- root is always statement sequence
- Statement sequence contains all statements as children
- operation is in the root, children are the operands
- first child of function contains arguments, second child the statement sequence
- conditional statements have two or three children: condition, if and optionally else
- recursive function, only continues folding if all children have folded successfully
- Folds using a lambda a class can have als attribute
- Tree of symbol tables -> testing in which scope we need to look
- There should be added that once a value is called in the scope it is created when running through the ast generation
- Also recursive
- output code is attribute of AST, each recursive function writes to this output
- function for getting a llvm variable of a node, which loads the variable into memory if required
- Functions of which the llvm code has the same structure have a get_llvm_template function, which gets filled in in the superclass (virtual functions)
- We should also note that autogenerated llvm variables are always preceded by a . instead of a letter, we do not use numbers because it vastly complicates the code
- Test function compares output and return code with clang compiled code
- Semantic error tests try to catch the error type that's expected and compares the string version of this error with the expected output
- Many tests fail because ++ isn't implemented (since it was optional) -> but we fixed it right now not in the zip though
- Arrays aren't working yet, but we are working on it, so these tests aren’t in the
- Function declaration c++ style also causes errors to fail: One function can be declared multiple times with different parameters
- comparisons are interpreted as float
- Definition in local scope not scope checked
- returnType mismatch not checked
- main not found not checked
- scan functions showed as succeeding but need to be tested manually :(
- We have chosen to translate llvm to mips
- this was easier to do because of not using the ast, and doing double work
- LLVM can be translated fast to mips, without too much hassle
- We have changed most of our c grammar because it was a mess as earlier said by Brent
- We’ve also made a new grammar for making a tree from the llvm code
- From the root (start_rule) you can go to a list of operations and after that your declarations
- An operation can be any of these,
- most of the grammar rules explain itself
- Then we have all the operations there are
- The grammar is just a generalization of the llvm code and is not that interesting to explain further
- The ast contains all the lines in llvm, it is directly a 1 to 1 translation
- Functions have a special representation in the ast
- In case of a definition these are the function arguments and the statement sequence
- In case of a function call we have the function arguments below it
- For a scanf and printf we do a separation of the string into the arguments in order to put the arguments passed through into the children like it is right there to be printed
- The ast is not deep at all so there are not a lot of things to take account for our llvm generation took all the hard work off our shoulders
- We use the same symbol table as before but we tweaked the code slightly in order to accomodate for llvm
- Now we have a separate dot generation for the symbol table
- Global variables are in the root of the symbol table
- There is also a global table available which contains all the variables
- It can be generated after all the symbol tables have been linked, and will be generated when executing the function merge
- For the code generation we use separate files in order to keep the ast class clean
- It uses a visitor which traverses in preorder traversal, it is the same as reading the llvm code from left to right
- We translate each operation directly from the llvm code to mips
- We do this by loading each variable into the 0 and 1 register of its kind
- So in the t or f register depending on the value it is
- And then we store the values in that memory location of like in llvm described
- You can see in the global variables that there are quite a lot this is because these are all the llvm variables
- For every function we have a building stackframe and a destructing stackframe, which speaks for itself
- The benchmark still fails on some tests but all are explainable
- There is still one test failing in semantic errors this is because we do not support -- and ++
- We also need to note that some of the benchmarks are modified because we do not support all the functionality, since some of it was optional
- There is also a massive improvement to the previous submission of our code because we have tweaked it such that the benchmarks will run
- This does not mean we have hardcoded all the benchmarks to succeed
- There is also the thing of scan tests not being tested, they need to be done manually
- Both our comparison tests do not work, because we have a special implementation
- This means when we compare two variables/constants we keep them in their type and do not convert them to decimal
- We do this in order to not mess around with types, it is too much work that is not necessary in our case
- Since our tests use a standard compiler we can not change the format tag since it will still fail because the standard compiler will give 0 for the output of a float when it is a decimal
- SPIM does behave weird because on mips it should work all right if tested by hand
- Also our float to int conversion do not work since they are optional
- A lot of tests for C to LLVM are fixed
